GAP model archive

This repository backs up the retained local models used for GAP reference generation and quantization evaluation. It contains multiple model families and checkpoints. Select an individual checkpoint and its projector when loading a model.

The directory layout preserves the original model-family/provider paths. The two Kimi checkpoints keep their separate directories, conversion records, and original Hugging Face source files. Split GGUF files, projectors, and draft heads are retained with their model files.

Files and provenance

The manifest records file sizes, SHA256 values, checksum evidence, and pinned source identities. SHA256SUMS covers the backed-up source files. Files produced by local conversion retain their conversion records; a pinned base checkpoint does not imply that the converted GGUF was downloaded from that source repository.

Some source identities are established by exact full-file SHA256 agreement with a pinned upstream file. This establishes byte identity without reconstructing the original download history. Verified identical Hub files may be copied server-side into this repository; other files are uploaded from local storage. Both transfer paths are checked against the same final file manifest.

Qualification scope

Candidate disposition is the current qualification record. It supersedes older smoke-test summaries retained in conversion directories.

Candidate usability is evaluated with the current llama.cpp CUDA 13.2.51 build on an NVIDIA RTX PRO 6000. The upstream model-support boundary is commit 6a1a922d269908a29cbd4b49c27e6a8e7fd10fae. The sparse numerical screen uses three fixed text windows and 765 next-token targets per candidate; flagged numerical results receive additional fixed-window confirmation. Reference NLL columns and runtime settings are checked against same-checkpoint controls.

A passing sparse screen means that the tested files did not show the screened severe failures. It does not certify full-corpus quality, all image/context combinations, or a statistical SNR claim. Projectors, draft heads, and original source tensors are auxiliary artifacts, not separately qualified standalone candidates.

Confirmed candidates unusable in this CUDA environment are excluded from this archive. When their full-file SHA matches upstream, the recorded reason is a runtime-specific numerical failure, not a corrupt download.

Gemma metadata compatibility is scoped to the current no-tool benchmark: token IDs, decoded pieces, full-corpus tokenization, and rendered benchmark prompts must agree. EOS metadata differences do not by themselves establish numerical failure. The applicable compatibility setting and numerical outcome are recorded with the candidate disposition.

Licenses and attribution

The files retain the licenses and terms of their respective model authors and providers. This archive does not assign one replacement license to all included models. Consult the pinned source repositories in the manifest and the original model/conversion records for attribution and applicable model terms.

Rebuildable virtual environments, caches, Git internals, and agent working directories are excluded from the backup.

Downloads last month
42,820
GGUF
Model size
31B params
Architecture
gemma4
Hardware compatibility
Log In to add your hardware

3-bit

4-bit

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collection including elichen-skymizer/GAP-models