GLM-5.3-U
GLM-5.3 Vision NVFP4 / ARVQ 8+8 hybrid.
This is the complete unrotated PV-tuned 8+8 ARVQ cold tier (all 75 MoE layers), with hot experts, language nonexpert weights and MTP from the selected NVFP4 donor. Vision tower/projector tensors come from the default base hybrid. Cold weights were fitted directly from the selected GLM-5.3 source weights; AQLM-decoded weights were not used as the fitting source.
Two shared FP4 codebooks (256+256 entries, dimension 8), 16-bit index pairs, E4M3 scales per 128 weights and a float32 global scale: 2.0625 bpw plus codebooks. The format is rvq256_256x8, version 2: 512 codewords and 64 uint32 index words per tile. Requires the matching 8+8 loader/kernel; the earlier 8+7 loader is incompatible. Four residual FP4 activation planes are the intended serving setting.
Initialization uses activation Hessians, LDLQ feedback and scale refitting. Each layer receives up to 200 Adam PV steps jointly tuning gate/up and down codebooks/scales against routing-weighted cold outputs from original FP8 weights. Indices stay frozen; FP4 codebooks and E4M3 scales are quantized in each forward. PV arithmetic: fp4_planes4_fp16_boundaries_v1. Activation-aware mode emulates four FP4 activation planes and FP16 SwiGLU boundaries using FP32 GEMMs. It does not reproduce SM120 MMA accumulation or TP reduction bit-for-bit. PV partitions: disjoint problem-hash partitions, captured separately for training, validation and audit. Validation selects the best state, including initialization; early stopping follows three checks without meaningful improvement after at least 80 steps. The audit is checked once after selection; regression falls back to initialization without retrying on the audit. Retained weights are reloaded for parity/nonregression checks. These are calibration quality checks, not independent task benchmarks; consult calibration_corpus.json for prompt-source overlap and capture limitations.
Cold source: /tmp/glm53-unc-fp8. NVFP4 donor (hot experts, language nonexpert weights
and MTP): /tmp/glm53-unc-nvfp4. Vision source: jarrelscy/GLM-5.3-Vision-NVFP4-AQLM-hybrid-1m.
Calibration uses text activations captured from the selected NVFP4 teacher.
Corpus provenance: {"counts": {"audit": {"problems": 104, "tokens": 322277}, "train": {"problems": 856, "tokens": 3048596}, "validation": {"problems": 89, "tokens": 345875}}, "limitations": "Text-only PV capture, 2048-token windows; image traces saved for subsequent MM capture. Earlier calibration used same prompt source datasets.", "partitions": "disjoint prompt hashes; separately captured", "reasoning_rollouts": true, "source": "/tmp/glm53-unc-traces"}.
Allocation: historical r3 blend recomputed; historical r3 recipe with 75% text REAP and
25% multimodal salience. See cold_assignment.json for allocation provenance.
Original GLM reasoning rollouts were deferred. See pv_report.json and
build_provenance.json for the actual build inputs and per-layer output errors.
Packed-index roundtrips, sampled decoded-weight parity, checkpoint structure, scales and hashes are checked. SM120 execution, full-model perplexity and multimodal task accuracy have not been validated by this A100 fitting campaign. This checkpoint remains experimental. No serving fork was edited by the fitter.
- Downloads last month
- 40