- MERNIK β measure-first quantization protocol
- The two rings
- Columns (every claim carries all three or stays home)
- Slow-ring protocol: HumanEval (NeoHorse-1-9B reference implementation)
- Reference results (NeoHorse-1-9B, official BF16 98.17% @ undisclosed ctx)
- Reference results II β OxCoder-9B (finetune duel, same Qwen3.5 bones)
- Fast ring II: GPQA-recognition (non-standard, read this)
- Canon: MSE spreads, SMAPE sacrifices
- The budget law (confirmed 2026-09-14 after clean reruns)
- Method in brief (the queue being measured)
- Flags (full usage)
- Supported architectures
- Scales down (not just 9B flagships)
- Quant repos (verdicts live here)
- Incoming (logged when landed, win or lose)
- Lineage
- The two rings
MERNIK β measure-first quantization protocol
MERNIK ("the one who measures") is the evolution of the ASHQ1 battlefield zoo: fewer utility duels, more verdicts. The method (priority queue) is built; MERNIK is how we prove anything about it. Think of it as a finetune of ASHQ1: same base weights (queue, pins, tied groups), retrained objective (measure-first protocol, three-column verdicts) β and far more capabilities on top (slow capability ring, KLD-aware teachers, relief ceilings, the zoo bench, norms shields). But the legend is not forgotten: every scar in the ledger traces back to it.
The two rings
| Ring | Job | Cadence |
|---|---|---|
| Fast / allocator | per-decision signal for the queue (imatrix, KLD-damage sweeps) | every run |
| Slow / capability | scored finals, fixed seeds, once per build | per release |
Allocator signals never certify. Capability scores never steer. Mixing them is how the PPL disease happened.
Columns (every claim carries all three or stays home)
- PPL = canary. Cheap, lies by sharpening. Demoted everywhere.
- KLD-vs-ref = rank column for allocator decisions (Soulfate24 battery: fixed span, fixed chunks, Flash-Attention). Triage, not certificate: blind to long-ctx, instruction drift, tool format, rare modes.
- Tasks = verdict. Slow ring. Only column that can crown or kill a build.
House rules: metrics diverge β KLD wins the argument; metrics agree β trust the pair. Every retraction keeps its scar in the ledger (README-ZOO.md).
Silence is signal (audit 2026-09-13). Empty completions scale with weakness across all clean runs: Ox 4β16 β Neo 20β27 β Q2K 63 β S2 115. Weak models go silent rather than babble β incapacity, not harness bug (server deaths look different: connection errors, audited per-run). S2's 39 passes stand taller for the 115 silences beside them. Strength reads as decisiveness: the distillate answers almost always (Ox empties 4β5) and is right almost always. Whether a quant goes silent or babbles is set by the finetune, not the size: same Qwen3.5 bones, Neo hesitates (20β27), OxCoder commits (4β5). Compare finetunes, not just budgets.
Slow-ring protocol: HumanEval (NeoHorse-1-9B reference implementation)
Server (one model at a time, GPU is a strict queue):
llama-server -m MODEL.gguf --port 28082 -ngl 99 -c 8192 --jinja --log-disable
Sampling (NeoHorse/Ornith thinking models; peg-native template):
temperature 1.0, top_p 0.95, top_k 20, min_p 0.0,
presence_penalty 0.0, repetition_penalty 1.0, max_tokens 2048
presence_penalty 1.5 (vendor-reported recipe) breaks thinking templates
(500 peg-native format); 0.0 verified. Preflight /health before every
battery β never score 164 empties against a dead server (cost us one battery
once; lesson kept).
Runner: scripts/run_humaneval.py (env overrides HE_*, HUMANEVAL_OUT,
HUMANEVAL_SERVER). Eval: human_eval.evaluation.evaluate_functional_correctness,
pass@1, k=[1]. Chain: serve_and_run per model (kill server, serve, preflight,
run, log exit=), results to eval_results/humaneval_<tag>.jsonl{,_results.jsonl}.
Saturation rule: base at ~98% can't resolve 97 vs 96 (floor-effect noise, the mirror of the PPL disease). Capability duels need bases at 60β85%, or they only answer the collapse question (does the small build fall off the cliff?).
Reference results (NeoHorse-1-9B, official BF16 98.17% @ undisclosed ctx)
| Build | Size | PPL (ctx1024) | KLD vs Q8-proxy | HE pass@1 |
|---|---|---|---|---|
| MERNIK-6500-MSE | 6.83 GB | 7.7695 | 0.0453 | 85.98% (141/164) |
| MERNIK-6500-SMAPE | 6.83 GB | 7.8252 (~tie) | 0.0476 (~tie) | 82.32% (135/164) β ties stock Q6, loses to MSE |
| MERNIK-6500-TD-MSE | 6.83 GB | 7.7701 (=MSE) | 0.0452 (=MSE) | 82.32% β PPL/KLD-identical, capability β3.7pp: direction doesn't move distribution, moves code |
| MERNIK-6500-TDN-MSE (norms shield) | 6.83 GB | 7.7701 | β | 83.54% β shield recovers +1.2pp of 3.7: norms matter, rest is elsewhere |
| MERNIK-6500-TDN-SMAPE (norms shield) | 6.83 GB | β | β | 79.27% β bolted-on shield disrupts the greedy path (β3pp); shield must be native (base floors), not post-hoc. BU-SMAPE (native shield) holds 82.32. |
| MERNIK-6500-TD-SMAPE | 6.83 GB | 7.8184 | 0.0505 | 82.32% β F16-attention core holds Q6-level with 236 drowned; bet 75β80 lost, logged |
| Q6_K (stock) | 7.36 GB | 7.9419 | 0.0118 | 82.32% (135/164) |
| MERNIK-5100-SMAPE | 5.36 GB | 7.8012 | 0.0586 | 82.93% (136/164) β Q6-class code at β2.2 GB π |
| MERNIK-5100-MSE | 5.36 GB | 7.7962 | 0.0588 | 79.88% (131/164) β clean rerun; law holds by 3pp |
| Q4_K_M (stock) | 5.62 GB | 7.7824 | 0.0875 | 76.83% (126/164) β clean rerun; our pair beats stock both ways |
| Q2_K (stock) | 3.83 GB | 90.0284 π | 2.6354 π | 0.00% β stub completions, real collapse |
| MERNIK-3650-Q2 (SMAPE+q3) | 3.65 GB | 10.3586 | 0.4706 | 23.78% (39/164) β beats stock Q2_K (0%) and our MSE-5100 (18.3%): at this size distribution decides everything |
Cross-KLD between the quants: D(MERNIKβQ6) = 0.320, D(Q6βMERNIK) = 0.318 β symmetric, and large: the pie deviates structurally, as pies do.
Recipe (repeatable one-to-one): llama-perplexity -m MODEL -f wiki.test.raw -ngl 99 -c 1024 -n 64 -b 512 --seed 7, KLD via --kl-divergence --kl-divergence-base <REF.dat>, ref logits from --save-all-logits on the
Q8 build (no BF16 runs β the Q6 duel stands without it). HE per slow-ring
protocol above. Corpus: wikitext-2-raw wiki.test.raw.
The bet, resolved as designed: KLD ranks Q6_K higher (4Γ closer to Q8), tasks crown MERNIK (+3.7pp at β0.5 GB). Two readings stay on the card: (a) hybrids deviate from any flat reference by construction, so KLD punishes the pie for being a pie β expected, not a refutation of the rank column inside its jurisdiction (allocator decisions); (b) on the verdict column the rank column does not transfer β capability and distribution-fidelity diverge here, and any claim that KLD picks the better model answers to this row.
Reference results II β OxCoder-9B (finetune duel, same Qwen3.5 bones)
| Build | Size | PPL | HE pass@1 |
|---|---|---|---|
| Ox-MSE-6500 | 6.83 GB | 7.5125 | 90.24% (148/164) β distillate leads |
| Ox-SMAPE-5100 | 5.36 GB | 7.5670 | 88.41% (145/164) β small beats Neo big |
| Ox-Q6K (stock) | 7.36 GB | 7.6758 | 84.15% (138/164) β flat keeps up, allocation wins |
Full card: wepiqx/OxCoder-9B-MERNIK-GGUF.
Fast ring II: GPQA-recognition (non-standard, read this)
How we score (NOT the vendor method): no generation, no CoT. Each
Diamond question + shuffled choices goes in one prompt; we read P(letter)
off the first generated token's top-30 distribution and pick the max.
198 questions, ~3 min/model. Script: scripts/gpqa_duel.py.
What it measures: recognition decisiveness, not reasoning. Proof: OxCoder-6500 scores below NeoHorse-6500 here (49.5 vs 51.5) while beating it on HE (90.2 vs 86.0) β the thinker spreads first-token mass (it wants to reason first), the decisive model stabs the letter. Rank uses: cheap allocator signal only. Never compare these numbers to vendor generation+CoT scores (their 86.9 lives in another universe).
| Build | GPQA-rec | HE pass@1 |
|---|---|---|
| Neo-MSE-6500 | 51.5% | 85.98% |
| Neo-SMAPE-5100 | 48.0% | 82.93% |
| Ox-MSE-6500 | 49.5% | 90.24% |
| Ox-SMAPE-5100 | 49.0% | 88.41% |
Canon: MSE spreads, SMAPE sacrifices
- MSE (default): relative-blind absolute gain β tiers spread evenly across layers (Q5/Q6-heavy middle). Best PPL and KLD on Ornith-class models (Ornith-1.5 @6500: MSE 8.6341 vs SMAPE 8.7845; NeoHorse: 7.7695 vs Q6 7.9419, KLD 0.0453 vs pair 0.32).
- SMAPE: relative lens β junk layers dumped to the Q4 floor, kings pushed to Q8 penthouses (barbell). Loses distribution columns on dense 9B β and on code at big budgets (82.32 vs 85.98 @6500), but owns small budgets (82.93 vs 79.88 @5100, see budget law). The lens, not the size.
Distributions @6500 (dry-run, 427 tensors incl. 177 F16 norms/1D):
| Tier | BU-MSE (spread) | BU-SMAPE (barbell) | TD-MSE | TD-SMAPE (F16 core) |
|---|---|---|---|---|
| F16 | 177 (2 MiB) | 177 (2 MiB) | 147 (2 MiB) | 85 (949 MiB) |
| Q4_K | 48 (972 MiB) | 91 (1989 MiB) | 71 (972 MiB) | 236 (2368 MiB) |
| Q5_K | 28 (1994 MiB) | 3 (1345 MiB) | 28 (1983 MiB) | 7 (1411 MiB) |
| Q6_K | 66 (2114 MiB) | 7 (249 MiB) | 71 (2127 MiB) | 14 (486 MiB) |
| Q8_0 | 108 (1417 MiB) | 149 (2913 MiB) | 110 (1417 MiB) | 85 (1287 MiB) |
Same budget, opposite philosophies: MSE fills the middle (Q5+Q6 = 94
tensors, 4107 MiB), SMAPE hollows it (Q5+Q6 = 10) to double the floor and
buy 41 more Q8 penthouses. Top-importance ffn_down@31: Q6_K under MSE,
Q8_0 under SMAPE β the kings do fly higher.
Top-down twins: TD-MSE moved ~30 small tensors (norms!) from F16 down to Q4 β PPL didn't blink (7.7701 vs 7.7695), HE bled β3.7pp. Norms are free real estate for perplexity and load-bearing walls for code. TD-SMAPE kept a 949 MiB F16 core of pure attention (qkv layers 20β30 + 3) and drowned everything else (236 at Q4) β the attention-vs-FFN experiment, HE pending. Both statements are true; the column decides which one you see.
Verdict @6500: SMAPE HE 82.32% = stock Q6K to the digit, MSE holds 85.98%. The barbell taxes code exactly where it taxes distribution: 91 tensors on the Q4 floor cost β3.7pp of capability. Kings flying higher (149ΓQ8) did not pay for the drowned middle. (Bet log: predicted 81β83 β closed.)
The budget law (confirmed 2026-09-14 after clean reruns)
β οΈ Correction arc (scar kept). The 18.3% / 11.59% below were DEAD servers scored silently (121/136 errors, OOM kills). Fixed with a mid-run health watchdog (aborts LOUDLY). Clean reruns: MSE-5100 79.88%, Q4_K_M 76.83%, 0 errors. The "5 GB cliff" never existed β the law holds narrower and honest.
| Budget | MSE | SMAPE | Law |
|---|---|---|---|
| 6500 | PPL 7.7695 / HE 85.98% | PPL 7.8252 / HE 82.32% | spread wins big (+3.7pp) |
| 5100 | PPL 7.7962 / HE 79.88% | PPL 7.8012 / HE 82.93% | barbell wins small (+3pp) |
Same PPL both times (~tie). Symmetric crossover between 5100 and 6500 on 9B dense: SMAPE owns small budgets, MSE owns big ones. PPL never saw any of this.
Method in brief (the queue being measured)
Every tensor starts at a floor tier by class; a global max-heap drains the
budget best-first by Ξ£(importance) Γ ΞMSE / MiB. Tied groups (identical
imatrix energy) upgrade as one unit with summed importance. Structural pins β
MTP head β Q8_0, output/token_embd β Q5_K, MoE routers β F16 β sit outside
the budget. One knob: --size in MiB. --top-down mirrors it: everything
from F16, downgrade cheapest-loss-first, slack refilled by the same metric.
pip install -r requirements.txt
python main.py --model M.gguf --imatrix M.imatrix.gguf --size 6800 --run
Needs a stock llama.cpp build (llama-quantize, llama-perplexity,
llama-server; override via LLAMA_QUANTIZE / LLAMA_PPL / LLAMA_SERVER).
Code map: main.py (CLI), classifier.py (queue), constants.py (tiers,
pins, toxicity, arch presets), experimental.py (retired zoo utilities),
scripts/group_damage_sweep.py (teacher labels), scripts/run_humaneval.py
(slow ring), scripts/audit_tiers.py (per-layer tier map of any quant file),
scripts/build_normstd.py (top-down with native norms shield).
Flags (full usage)
# classify + dry-run (~1 sec) β always preview before burning 10 min
python main.py --model M.gguf --imatrix M.imatrix.gguf --size 6800
# actual quant
python main.py --model M.gguf --imatrix M.imatrix.gguf --size 6800 --run
# print the --tensor-type config only / show hard floors
python main.py --model M.gguf --imatrix M.imatrix.gguf --size 6800 --show-config
python main.py --show-floors
| Flag | What it does |
|---|---|
--model M.gguf |
source weights (BF16/F16) β required |
--imatrix I.gguf |
imatrix file; repeatable (--imatrix A --imatrix B --imatrix-method max|mean) |
--size MIB |
the only budget knob β target file size in MiB |
--output O.gguf |
output path (default: <model>-SHQ.gguf) |
--run |
execute llama-quantize; without it: dry-run only |
--allow-q3-or-lower |
CAN_Q3 types (ffn_gate/up/down, attn_output, ssm_out) may start at IQ2_XXS β wider spread, sub-4-bit risk |
--top-down |
reverse mode: everything from F16, downgrade cheapest-loss-first to fit (norms included) |
--pin-norms |
top-down native norms shield: norms/small tensors stay F16 outside the budget (post-hoc forcing disrupts the path β proven by TDN duels) |
--free-pins |
EXPERIMENTAL diagnostic: output/token_embd/MTP/routers fight for budget (validated AGAINST β pins stay) |
--utility NAME |
queue gain metric: mse (default, proven) / rmse / smape / logcosh / ssim / smape_ssim / smape_frag / retired: huber hybrid pw_ssim netdmg |
--relief labels.jsonl --relief-thr -0.5 |
measured sweet spots: groups scoring below thr get ceiling Q4, budget flows elsewhere |
--ssim-table / --frag-w / --ptable / --netpred / --netdmg-w |
data + weights for measured utilities (see zoo ledger for what survived) |
--verbose |
detailed per-tier output |
Fails fast: if even base floors exceed --size, the run aborts instead of
silently producing an oversized file. Estimator accuracy Β±1β3 MiB vs
llama-quantize dry-run.
# audit any quant file: histogram, biggest tensors, per-layer map
python scripts/audit_tiers.py --model M-6500.gguf
python scripts/audit_tiers.py --model M-6500.gguf --big 10
python scripts/audit_tiers.py --model M-6500.gguf --layer 31
Supported architectures
| Arch | Detection | Features |
|---|---|---|
qwen35 |
SSM + QKV | Hybrid attention, SSM layers, GQA, MTP support (whole head β Q8_0, outside budget) |
mellum2 |
MoE (exps tensors) |
Mixture of Experts, GQA, routers pinned F16 |
bailingmoe3 |
KDA+MLA + MoE | Ling-3.0 family: hybrid-linear attention, routed + shared experts, routers pinned F16 |
granite |
general.architecture |
Dense GQA, 40 layers, separate Q/K/V |
spark2_5 |
general.architecture |
Spark-X2.5 dense + hybrid sliding-window attention (1 full + 3 SWA), fused q_k_v_proj; PPL at -c 1024 |
gemma4 |
layer-scale norms | QAT support, Q4_K attention floor |
| llama (generic) | tensor names | MiniCPM5, NeoHorse and friends: dense GQA via standard names |
Detection prefers general.architecture metadata, falls back to tensor-name
heuristics. New archs plug in via ARCH_FEATURES in constants.py.
Scales down (not just 9B flagships)
The same queue, same --size knob, no retuning β from pocket budgets to
flagships (all PPL canary, ctx 1024, seed 7):
| Model | Budget | Result |
|---|---|---|
| Spark-X2.5-1.7B | 1000 MiB | zoo stand: full utility-duel ladder lives here |
| Spark-X2.5-4B | 4000 MiB | 31.25 vs Q8_0 30.86 (+0.39 at β100 MiB) |
| Ornith-1.5-9B-MTP (Qwen3.5) | 6500 MiB | 8.6541 + HE-verdict above |
| NeoHorse-1-9B | 6500 MiB | HE 85.98% > Q6_K 82.32% |
Small dense, MoE (Ling-3.0-tiny 8B), GDN hybrid, MTP heads β allocation holds across families. Per-family quirks (Spark pipeline bubble, MoE padding) are logged in the scar ledger, not hidden.
Quant repos (verdicts live here)
- NeoHorse-1-9B-MERNIK-GGUF β HE triple above
- Ornith-1.5-9B-MTP-ASHQ1-GGUF β KLD-validated GDN hybrid
- Spark-X2.5-1.7B-ASHQ1-GGUF β zoo stand + pin-validation duels
Incoming (logged when landed, win or lose)
- Q4_K_M stock duel (PPL β HE running): does our 5100 pair beat stock?
- Quality-36 vs MERNIK-6500 on HE β the fair fight, file requested
- KLD-damage teacher (MiniCPM tail on the forge) β Gini answer, MLP un-pause
- Relief ceilings on 9B (NeoHorse damage sweep β
--reliefduel vs MSE) - smape_frag for NeoHorse, gated on SMAPE-5100 beating stock + community demand
- KLD columns for the Ox trio (same Q8-proxy recipe)
- Gini v3 (Q5βQ2 drops): Q5βQ3 labels sit at noise floor (max 0.06) β bigger drops for a signal with teeth. Queued after the current tail.
Lineage
MERNIK is the ASHQ1 battlefield zoo grown up: same queue, plus a protocol.
Most features migrated from the original ASHQ1 and the ZOO β priority-queue
drain, tied groups with summed importance, MTP-head handling, embd/router
pins, top-down mode with slack refilling, sub-4-bit toxicity, MoE padding β
and there is more here: KLD-aware teacher sweeps, relief ceilings, the
experimental utility bench, the slow capability ring. Fully open code,
Apache-2.0 β take it, fork it, beat us with it (the ledger will record that too).
Built in the open with Soulfate24, whose
KLD battery ended our PPL theater β the scar ledger (README-ZOO.md) records
every round. Steel sharpens steel.
Model tree for wepiqx/MERNIK
Base model
wepiqx/ASHQ1