Qwythos-9B-v2-NVFP4
NVFP4 (W4A4) build of empero-ai/Qwythos-9B-v2 for vLLM on NVIDIA Blackwell — the native 4-bit serving path nobody ships for this checkpoint. Empero already publishes the MTP-GGUF for llama.cpp; this is the piece that was missing: a compressed-tensors NVFP4 artifact that runs the model's GEMMs directly on the Blackwell FP4 tensor cores under vLLM.
Quantized by protoLabs. Base model, its capabilities, and its uncensored research posture are Empero AI's — see their card.
What this is
- NVFP4 W4A4 on the 128 transformer linears (attention + MLP). The hybrid DeltaNet (
linear_attn) layers, the vision tower,lm_head, and the MTP head are kept BF16 — DeltaNet corrupts under 4-bit, and the vision path stays lossless. - 11.2 GB on disk (vs 19.3 GB BF16 source; the BF16-preserved vision tower is most of the remainder).
- MTP sidecar included (
model-mtp.safetensors) — see the status note below. - Inherits the base wholesale: 1M-token context (YaRN), Qwen3.5 multimodal stack, FTPO looping-fix (greedy loop rate 6.7%→0%), uncensored for research/red-team/bio-chem-clinical work.
Speed
Single RTX PRO 6000 Blackwell (sm120), vLLM 0.24.0, NVFP4 --linear-backend marlin, BF16 KV:
| Regime | tok/s | TPOT |
|---|---|---|
| single-stream (real text) | 138 | — |
| chat 1k/1k · C1 | 152 | 6.5 ms |
| chat 1k/1k · C8 | 1051 | 7.1 ms |
| chat 1k/1k · C32 | 2523 | 11.4 ms |
(Client-side vllm bench serve, random dataset, fixed seed. --ignore-eos. Single trial.)
Serving (vLLM ≥ 0.24.0, Blackwell / sm120)
vllm serve protoLabsAI/Qwythos-9B-v2-NVFP4 \
--reasoning-parser qwen3 \
--enable-auto-tool-choice --tool-call-parser qwen3_xml \
--linear-backend marlin \
--max-model-len 65536 \
--trust-remote-code
sm120 requires the FlashInfer/CUDA-13 env (or you hit the "no CUDA arch for major 12" crash):
VLLM_USE_FLASHINFER_SAMPLER=0, CUDA_HOME=<cu13>, FLASHINFER_CUDA_ARCH_LIST=12.0f,
NVCC_APPEND_FLAGS=-DCCCL_DISABLE_CTK_COMPATIBILITY_CHECK. --linear-backend marlin is the
proven NVFP4 kernel on sm120 (fastest at the single-stream / low-batch regime that matters here).
Recipe
llm-compressor NVFP4, calibrated (512 samples, seq 2048). Ignore list: lm_head, re:.*visual.*,
re:.*linear_attn.*, re:.*mtp.*. VL keys canonicalized post-quant. Reproducible from the BF16 source.
Quality gate — NVFP4 vs BF16 base (paired)
Every suite run against both the NVFP4 build and the BF16 source, same prompts/harness, on this rig. Deterministic suites are judge-free (solver-verified reasoning, execution-graded code); claw uses an independent LLM grader. Single trial (coding ×3).
| Suite | NVFP4 | BF16 base | Δ |
|---|---|---|---|
| Function-calling | 89% (48/54) | 94% (51/54) | −5 pp |
| Reasoning-v2 (solver-verified) | 0.759 | 0.789 | −0.03 |
| Coding (exec-graded, ×3) | 0.518 | 0.553 | −0.03 |
| Claw (agentic, 10 tasks) | 0.651 | 0.614 | +0.04 |
| Coherence @ depth 4K–60K | clean | — | needle ✓, no repetition/degradation |
Near-parity. The only real cost is a few function-calling cases (single-trial, 54-task suite); reasoning and coding are within noise, agentic is flat-to-better, and there is no quant-rot at long context (verified to 60K). Expected NVFP4 behaviour for a 9B.
Status (2026-07-11)
- ✅ Serves + coherent output on vLLM 0.24.0/sm120; quality gate passed (table above).
- ⚠️ MTP sidecar is included but does not accelerate vLLM decoding for this checkpoint (0% draft acceptance). We isolated the cause: the sidecar is byte-identical in layout to a head that drafts at ~67% on the same NVFP4+MTP stack for another Qwen3.5-9B, so this is the head's weights, not our packaging or the runtime — the restored MTP head doesn't predict this FTPO-tuned checkpoint's distribution. The sidecar ships for completeness/llama.cpp; for speculative decoding today use Empero's GGUF-MTP. A distilled-head fix may follow.
Requests
Need a different size or format (GGUF, other quant, longer-context config)? Open a Community discussion — we usually ship within 48h.
License
apache-2.0, inherited from the base. Uncensored — this is Empero's research/red-team model; the quant changes numerics, not alignment. Use responsibly and within the base model's terms.
- Downloads last month
- 9,336
Model tree for protoLabsAI/Qwythos-9B-v2-NVFP4
Base model
Qwen/Qwen3.5-9B-Base