view article Article AX-Ray, Finding Causal-Leakage Defects in Two General-Purpose Public Models FINAL-Bench • Aug 14 • 13
view article Article Model Genome: Fingerprinting Whether an LLM Was Trained From Scratch or Derived mayafree • Aug 8 • 18
view article Article The Fast Gemma Challenge: our verified-SOTA recipe, in full FINAL-Bench • Aug 3 • 24
view article Article Aether-7B-5Attn: A 100% Open-Source Sovereign Foundation Model — and a Controlled Experiment in Heterogeneous Attention FINAL-Bench • Jul 19 • 21
Darwin Family: MRI-Trust-Weighted Evolutionary Merging for Training-Free Scaling of Language-Model Reasoning Paper • 2605.14386 • Published May 14 • 52
view article Article Training-Free Reasoning at 88.89% on GPQA Diamond: How Darwin Family Hit Frontier Scores Without a Single Gradient Step FINAL-Bench • May 15 • 18
view article Article 🏟️ Smol AI WorldCup: A 5-Axis Benchmark That Reveals What Small Language Models Can Really Do FINAL-Bench • Mar 10 • 38
view article Article FINAL Bench: The Real Bottleneck to AGI Is Self-Correction FINAL-Bench • Feb 21 • 20