MERNIK β€” measure-first quantization protocol

MERNIK ("the one who measures") is the evolution of the ASHQ1 battlefield zoo: fewer utility duels, more verdicts. The method (priority queue) is built; MERNIK is how we prove anything about it. Think of it as a finetune of ASHQ1: same base weights (queue, pins, tied groups), retrained objective (measure-first protocol, three-column verdicts) β€” and far more capabilities on top (slow capability ring, KLD-aware teachers, relief ceilings, the zoo bench, norms shields). But the legend is not forgotten: every scar in the ledger traces back to it.

The two rings

Ring Job Cadence
Fast / allocator per-decision signal for the queue (imatrix, KLD-damage sweeps) every run
Slow / capability scored finals, fixed seeds, once per build per release

Allocator signals never certify. Capability scores never steer. Mixing them is how the PPL disease happened.

Columns (every claim carries all three or stays home)

  • PPL = canary. Cheap, lies by sharpening. Demoted everywhere.
  • KLD-vs-ref = rank column for allocator decisions (Soulfate24 battery: fixed span, fixed chunks, Flash-Attention). Triage, not certificate: blind to long-ctx, instruction drift, tool format, rare modes.
  • Tasks = verdict. Slow ring. Only column that can crown or kill a build.

House rules: metrics diverge β†’ KLD wins the argument; metrics agree β†’ trust the pair. Every retraction keeps its scar in the ledger (README-ZOO.md).

Silence is signal (audit 2026-09-13). Empty completions scale with weakness across all clean runs: Ox 4–16 β†’ Neo 20–27 β†’ Q2K 63 β†’ S2 115. Weak models go silent rather than babble β€” incapacity, not harness bug (server deaths look different: connection errors, audited per-run). S2's 39 passes stand taller for the 115 silences beside them. Strength reads as decisiveness: the distillate answers almost always (Ox empties 4–5) and is right almost always. Whether a quant goes silent or babbles is set by the finetune, not the size: same Qwen3.5 bones, Neo hesitates (20–27), OxCoder commits (4–5). Compare finetunes, not just budgets.

Slow-ring protocol: HumanEval (NeoHorse-1-9B reference implementation)

Server (one model at a time, GPU is a strict queue):

llama-server -m MODEL.gguf --port 28082 -ngl 99 -c 8192 --jinja --log-disable

Sampling (NeoHorse/Ornith thinking models; peg-native template):

temperature 1.0, top_p 0.95, top_k 20, min_p 0.0,
presence_penalty 0.0, repetition_penalty 1.0, max_tokens 2048

presence_penalty 1.5 (vendor-reported recipe) breaks thinking templates (500 peg-native format); 0.0 verified. Preflight /health before every battery β€” never score 164 empties against a dead server (cost us one battery once; lesson kept).

Runner: scripts/run_humaneval.py (env overrides HE_*, HUMANEVAL_OUT, HUMANEVAL_SERVER). Eval: human_eval.evaluation.evaluate_functional_correctness, pass@1, k=[1]. Chain: serve_and_run per model (kill server, serve, preflight, run, log exit=), results to eval_results/humaneval_<tag>.jsonl{,_results.jsonl}.

Saturation rule: base at ~98% can't resolve 97 vs 96 (floor-effect noise, the mirror of the PPL disease). Capability duels need bases at 60–85%, or they only answer the collapse question (does the small build fall off the cliff?).

Reference results (NeoHorse-1-9B, official BF16 98.17% @ undisclosed ctx)

Build Size PPL (ctx1024) KLD vs Q8-proxy HE pass@1
MERNIK-6500-MSE 6.83 GB 7.7695 0.0453 85.98% (141/164)
MERNIK-6500-SMAPE 6.83 GB 7.8252 (~tie) 0.0476 (~tie) 82.32% (135/164) β€” ties stock Q6, loses to MSE
MERNIK-6500-TD-MSE 6.83 GB 7.7701 (=MSE) 0.0452 (=MSE) 82.32% β€” PPL/KLD-identical, capability βˆ’3.7pp: direction doesn't move distribution, moves code
MERNIK-6500-TDN-MSE (norms shield) 6.83 GB 7.7701 β€” 83.54% β€” shield recovers +1.2pp of 3.7: norms matter, rest is elsewhere
MERNIK-6500-TDN-SMAPE (norms shield) 6.83 GB β€” β€” 79.27% β€” bolted-on shield disrupts the greedy path (βˆ’3pp); shield must be native (base floors), not post-hoc. BU-SMAPE (native shield) holds 82.32.
MERNIK-6500-TD-SMAPE 6.83 GB 7.8184 0.0505 82.32% β€” F16-attention core holds Q6-level with 236 drowned; bet 75–80 lost, logged
Q6_K (stock) 7.36 GB 7.9419 0.0118 82.32% (135/164)
MERNIK-5100-SMAPE 5.36 GB 7.8012 0.0586 82.93% (136/164) β€” Q6-class code at βˆ’2.2 GB πŸ‘‘
MERNIK-5100-MSE 5.36 GB 7.7962 0.0588 79.88% (131/164) β€” clean rerun; law holds by 3pp
Q4_K_M (stock) 5.62 GB 7.7824 0.0875 76.83% (126/164) β€” clean rerun; our pair beats stock both ways
Q2_K (stock) 3.83 GB 90.0284 πŸ’€ 2.6354 πŸ’€ 0.00% β€” stub completions, real collapse
MERNIK-3650-Q2 (SMAPE+q3) 3.65 GB 10.3586 0.4706 23.78% (39/164) β€” beats stock Q2_K (0%) and our MSE-5100 (18.3%): at this size distribution decides everything

Cross-KLD between the quants: D(MERNIKβ€–Q6) = 0.320, D(Q6β€–MERNIK) = 0.318 β€” symmetric, and large: the pie deviates structurally, as pies do.

Recipe (repeatable one-to-one): llama-perplexity -m MODEL -f wiki.test.raw -ngl 99 -c 1024 -n 64 -b 512 --seed 7, KLD via --kl-divergence --kl-divergence-base <REF.dat>, ref logits from --save-all-logits on the Q8 build (no BF16 runs β€” the Q6 duel stands without it). HE per slow-ring protocol above. Corpus: wikitext-2-raw wiki.test.raw.

The bet, resolved as designed: KLD ranks Q6_K higher (4Γ— closer to Q8), tasks crown MERNIK (+3.7pp at βˆ’0.5 GB). Two readings stay on the card: (a) hybrids deviate from any flat reference by construction, so KLD punishes the pie for being a pie β€” expected, not a refutation of the rank column inside its jurisdiction (allocator decisions); (b) on the verdict column the rank column does not transfer β€” capability and distribution-fidelity diverge here, and any claim that KLD picks the better model answers to this row.

Reference results II β€” OxCoder-9B (finetune duel, same Qwen3.5 bones)

Build Size PPL HE pass@1
Ox-MSE-6500 6.83 GB 7.5125 90.24% (148/164) β€” distillate leads
Ox-SMAPE-5100 5.36 GB 7.5670 88.41% (145/164) β€” small beats Neo big
Ox-Q6K (stock) 7.36 GB 7.6758 84.15% (138/164) β€” flat keeps up, allocation wins

Full card: wepiqx/OxCoder-9B-MERNIK-GGUF.

Fast ring II: GPQA-recognition (non-standard, read this)

How we score (NOT the vendor method): no generation, no CoT. Each Diamond question + shuffled choices goes in one prompt; we read P(letter) off the first generated token's top-30 distribution and pick the max. 198 questions, ~3 min/model. Script: scripts/gpqa_duel.py.

What it measures: recognition decisiveness, not reasoning. Proof: OxCoder-6500 scores below NeoHorse-6500 here (49.5 vs 51.5) while beating it on HE (90.2 vs 86.0) β€” the thinker spreads first-token mass (it wants to reason first), the decisive model stabs the letter. Rank uses: cheap allocator signal only. Never compare these numbers to vendor generation+CoT scores (their 86.9 lives in another universe).

Build GPQA-rec HE pass@1
Neo-MSE-6500 51.5% 85.98%
Neo-SMAPE-5100 48.0% 82.93%
Ox-MSE-6500 49.5% 90.24%
Ox-SMAPE-5100 49.0% 88.41%

Canon: MSE spreads, SMAPE sacrifices

  • MSE (default): relative-blind absolute gain β€” tiers spread evenly across layers (Q5/Q6-heavy middle). Best PPL and KLD on Ornith-class models (Ornith-1.5 @6500: MSE 8.6341 vs SMAPE 8.7845; NeoHorse: 7.7695 vs Q6 7.9419, KLD 0.0453 vs pair 0.32).
  • SMAPE: relative lens β€” junk layers dumped to the Q4 floor, kings pushed to Q8 penthouses (barbell). Loses distribution columns on dense 9B β€” and on code at big budgets (82.32 vs 85.98 @6500), but owns small budgets (82.93 vs 79.88 @5100, see budget law). The lens, not the size.

Distributions @6500 (dry-run, 427 tensors incl. 177 F16 norms/1D):

Tier BU-MSE (spread) BU-SMAPE (barbell) TD-MSE TD-SMAPE (F16 core)
F16 177 (2 MiB) 177 (2 MiB) 147 (2 MiB) 85 (949 MiB)
Q4_K 48 (972 MiB) 91 (1989 MiB) 71 (972 MiB) 236 (2368 MiB)
Q5_K 28 (1994 MiB) 3 (1345 MiB) 28 (1983 MiB) 7 (1411 MiB)
Q6_K 66 (2114 MiB) 7 (249 MiB) 71 (2127 MiB) 14 (486 MiB)
Q8_0 108 (1417 MiB) 149 (2913 MiB) 110 (1417 MiB) 85 (1287 MiB)

Same budget, opposite philosophies: MSE fills the middle (Q5+Q6 = 94 tensors, 4107 MiB), SMAPE hollows it (Q5+Q6 = 10) to double the floor and buy 41 more Q8 penthouses. Top-importance ffn_down@31: Q6_K under MSE, Q8_0 under SMAPE β€” the kings do fly higher.

Top-down twins: TD-MSE moved ~30 small tensors (norms!) from F16 down to Q4 β€” PPL didn't blink (7.7701 vs 7.7695), HE bled βˆ’3.7pp. Norms are free real estate for perplexity and load-bearing walls for code. TD-SMAPE kept a 949 MiB F16 core of pure attention (qkv layers 20–30 + 3) and drowned everything else (236 at Q4) β€” the attention-vs-FFN experiment, HE pending. Both statements are true; the column decides which one you see.

Verdict @6500: SMAPE HE 82.32% = stock Q6K to the digit, MSE holds 85.98%. The barbell taxes code exactly where it taxes distribution: 91 tensors on the Q4 floor cost βˆ’3.7pp of capability. Kings flying higher (149Γ—Q8) did not pay for the drowned middle. (Bet log: predicted 81–83 β€” closed.)

The budget law (confirmed 2026-09-14 after clean reruns)

⚠️ Correction arc (scar kept). The 18.3% / 11.59% below were DEAD servers scored silently (121/136 errors, OOM kills). Fixed with a mid-run health watchdog (aborts LOUDLY). Clean reruns: MSE-5100 79.88%, Q4_K_M 76.83%, 0 errors. The "5 GB cliff" never existed β€” the law holds narrower and honest.

Budget MSE SMAPE Law
6500 PPL 7.7695 / HE 85.98% PPL 7.8252 / HE 82.32% spread wins big (+3.7pp)
5100 PPL 7.7962 / HE 79.88% PPL 7.8012 / HE 82.93% barbell wins small (+3pp)

Same PPL both times (~tie). Symmetric crossover between 5100 and 6500 on 9B dense: SMAPE owns small budgets, MSE owns big ones. PPL never saw any of this.

Method in brief (the queue being measured)

Every tensor starts at a floor tier by class; a global max-heap drains the budget best-first by Ξ£(importance) Γ— Ξ”MSE / MiB. Tied groups (identical imatrix energy) upgrade as one unit with summed importance. Structural pins β€” MTP head β†’ Q8_0, output/token_embd β†’ Q5_K, MoE routers β†’ F16 β€” sit outside the budget. One knob: --size in MiB. --top-down mirrors it: everything from F16, downgrade cheapest-loss-first, slack refilled by the same metric.

pip install -r requirements.txt
python main.py --model M.gguf --imatrix M.imatrix.gguf --size 6800 --run

Needs a stock llama.cpp build (llama-quantize, llama-perplexity, llama-server; override via LLAMA_QUANTIZE / LLAMA_PPL / LLAMA_SERVER). Code map: main.py (CLI), classifier.py (queue), constants.py (tiers, pins, toxicity, arch presets), experimental.py (retired zoo utilities), scripts/group_damage_sweep.py (teacher labels), scripts/run_humaneval.py (slow ring), scripts/audit_tiers.py (per-layer tier map of any quant file), scripts/build_normstd.py (top-down with native norms shield).

Flags (full usage)

# classify + dry-run (~1 sec) β€” always preview before burning 10 min
python main.py --model M.gguf --imatrix M.imatrix.gguf --size 6800

# actual quant
python main.py --model M.gguf --imatrix M.imatrix.gguf --size 6800 --run

# print the --tensor-type config only / show hard floors
python main.py --model M.gguf --imatrix M.imatrix.gguf --size 6800 --show-config
python main.py --show-floors
Flag What it does
--model M.gguf source weights (BF16/F16) β€” required
--imatrix I.gguf imatrix file; repeatable (--imatrix A --imatrix B --imatrix-method max|mean)
--size MIB the only budget knob β€” target file size in MiB
--output O.gguf output path (default: <model>-SHQ.gguf)
--run execute llama-quantize; without it: dry-run only
--allow-q3-or-lower CAN_Q3 types (ffn_gate/up/down, attn_output, ssm_out) may start at IQ2_XXS β€” wider spread, sub-4-bit risk
--top-down reverse mode: everything from F16, downgrade cheapest-loss-first to fit (norms included)
--pin-norms top-down native norms shield: norms/small tensors stay F16 outside the budget (post-hoc forcing disrupts the path β€” proven by TDN duels)
--free-pins EXPERIMENTAL diagnostic: output/token_embd/MTP/routers fight for budget (validated AGAINST β€” pins stay)
--utility NAME queue gain metric: mse (default, proven) / rmse / smape / logcosh / ssim / smape_ssim / smape_frag / retired: huber hybrid pw_ssim netdmg
--relief labels.jsonl --relief-thr -0.5 measured sweet spots: groups scoring below thr get ceiling Q4, budget flows elsewhere
--ssim-table / --frag-w / --ptable / --netpred / --netdmg-w data + weights for measured utilities (see zoo ledger for what survived)
--verbose detailed per-tier output

Fails fast: if even base floors exceed --size, the run aborts instead of silently producing an oversized file. Estimator accuracy Β±1–3 MiB vs llama-quantize dry-run.

# audit any quant file: histogram, biggest tensors, per-layer map
python scripts/audit_tiers.py --model M-6500.gguf
python scripts/audit_tiers.py --model M-6500.gguf --big 10
python scripts/audit_tiers.py --model M-6500.gguf --layer 31

Supported architectures

Arch Detection Features
qwen35 SSM + QKV Hybrid attention, SSM layers, GQA, MTP support (whole head β†’ Q8_0, outside budget)
mellum2 MoE (exps tensors) Mixture of Experts, GQA, routers pinned F16
bailingmoe3 KDA+MLA + MoE Ling-3.0 family: hybrid-linear attention, routed + shared experts, routers pinned F16
granite general.architecture Dense GQA, 40 layers, separate Q/K/V
spark2_5 general.architecture Spark-X2.5 dense + hybrid sliding-window attention (1 full + 3 SWA), fused q_k_v_proj; PPL at -c 1024
gemma4 layer-scale norms QAT support, Q4_K attention floor
llama (generic) tensor names MiniCPM5, NeoHorse and friends: dense GQA via standard names

Detection prefers general.architecture metadata, falls back to tensor-name heuristics. New archs plug in via ARCH_FEATURES in constants.py.

Scales down (not just 9B flagships)

The same queue, same --size knob, no retuning β€” from pocket budgets to flagships (all PPL canary, ctx 1024, seed 7):

Model Budget Result
Spark-X2.5-1.7B 1000 MiB zoo stand: full utility-duel ladder lives here
Spark-X2.5-4B 4000 MiB 31.25 vs Q8_0 30.86 (+0.39 at βˆ’100 MiB)
Ornith-1.5-9B-MTP (Qwen3.5) 6500 MiB 8.6541 + HE-verdict above
NeoHorse-1-9B 6500 MiB HE 85.98% > Q6_K 82.32%

Small dense, MoE (Ling-3.0-tiny 8B), GDN hybrid, MTP heads β€” allocation holds across families. Per-family quirks (Spark pipeline bubble, MoE padding) are logged in the scar ledger, not hidden.

Quant repos (verdicts live here)

Incoming (logged when landed, win or lose)

  • Q4_K_M stock duel (PPL β†’ HE running): does our 5100 pair beat stock?
  • Quality-36 vs MERNIK-6500 on HE β€” the fair fight, file requested
  • KLD-damage teacher (MiniCPM tail on the forge) β†’ Gini answer, MLP un-pause
  • Relief ceilings on 9B (NeoHorse damage sweep β†’ --relief duel vs MSE)
  • smape_frag for NeoHorse, gated on SMAPE-5100 beating stock + community demand
  • KLD columns for the Ox trio (same Q8-proxy recipe)
  • Gini v3 (Q5β†’Q2 drops): Q5β†’Q3 labels sit at noise floor (max 0.06) β€” bigger drops for a signal with teeth. Queued after the current tail.

Lineage

MERNIK is the ASHQ1 battlefield zoo grown up: same queue, plus a protocol. Most features migrated from the original ASHQ1 and the ZOO β€” priority-queue drain, tied groups with summed importance, MTP-head handling, embd/router pins, top-down mode with slack refilling, sub-4-bit toxicity, MoE padding β€” and there is more here: KLD-aware teacher sweeps, relief ceilings, the experimental utility bench, the slow capability ring. Fully open code, Apache-2.0 β€” take it, fork it, beat us with it (the ledger will record that too). Built in the open with Soulfate24, whose KLD battery ended our PPL theater β€” the scar ledger (README-ZOO.md) records every round. Steel sharpens steel.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for wepiqx/MERNIK

Base model

wepiqx/ASHQ1
Finetuned
(2)
this model

Collection including wepiqx/MERNIK