Instructions to use IsValorum/KAT-Coder-V2.5-Dev-APEX-I-MiniPlus-V2-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use IsValorum/KAT-Coder-V2.5-Dev-APEX-I-MiniPlus-V2-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf IsValorum/KAT-Coder-V2.5-Dev-APEX-I-MiniPlus-V2-GGUF # Run inference directly in the terminal: llama cli -hf IsValorum/KAT-Coder-V2.5-Dev-APEX-I-MiniPlus-V2-GGUF
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf IsValorum/KAT-Coder-V2.5-Dev-APEX-I-MiniPlus-V2-GGUF # Run inference directly in the terminal: llama cli -hf IsValorum/KAT-Coder-V2.5-Dev-APEX-I-MiniPlus-V2-GGUF
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf IsValorum/KAT-Coder-V2.5-Dev-APEX-I-MiniPlus-V2-GGUF # Run inference directly in the terminal: ./llama-cli -hf IsValorum/KAT-Coder-V2.5-Dev-APEX-I-MiniPlus-V2-GGUF
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf IsValorum/KAT-Coder-V2.5-Dev-APEX-I-MiniPlus-V2-GGUF # Run inference directly in the terminal: ./build/bin/llama-cli -hf IsValorum/KAT-Coder-V2.5-Dev-APEX-I-MiniPlus-V2-GGUF
Use Docker
docker model run hf.co/IsValorum/KAT-Coder-V2.5-Dev-APEX-I-MiniPlus-V2-GGUF
- LM Studio
- Jan
- vLLM
How to use IsValorum/KAT-Coder-V2.5-Dev-APEX-I-MiniPlus-V2-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "IsValorum/KAT-Coder-V2.5-Dev-APEX-I-MiniPlus-V2-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "IsValorum/KAT-Coder-V2.5-Dev-APEX-I-MiniPlus-V2-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/IsValorum/KAT-Coder-V2.5-Dev-APEX-I-MiniPlus-V2-GGUF
- Ollama
How to use IsValorum/KAT-Coder-V2.5-Dev-APEX-I-MiniPlus-V2-GGUF with Ollama:
ollama run hf.co/IsValorum/KAT-Coder-V2.5-Dev-APEX-I-MiniPlus-V2-GGUF
- Unsloth Desktop
- Pi
How to use IsValorum/KAT-Coder-V2.5-Dev-APEX-I-MiniPlus-V2-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf IsValorum/KAT-Coder-V2.5-Dev-APEX-I-MiniPlus-V2-GGUF
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "IsValorum/KAT-Coder-V2.5-Dev-APEX-I-MiniPlus-V2-GGUF" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use IsValorum/KAT-Coder-V2.5-Dev-APEX-I-MiniPlus-V2-GGUF with Docker Model Runner:
docker model run hf.co/IsValorum/KAT-Coder-V2.5-Dev-APEX-I-MiniPlus-V2-GGUF
- Lemonade
How to use IsValorum/KAT-Coder-V2.5-Dev-APEX-I-MiniPlus-V2-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull IsValorum/KAT-Coder-V2.5-Dev-APEX-I-MiniPlus-V2-GGUF
Run and chat with the model
lemonade run user.KAT-Coder-V2.5-Dev-APEX-I-MiniPlus-V2-GGUF-{{QUANT_TAG}}List all available models
lemonade list
- Hermes Agent
How to use IsValorum/KAT-Coder-V2.5-Dev-APEX-I-MiniPlus-V2-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf IsValorum/KAT-Coder-V2.5-Dev-APEX-I-MiniPlus-V2-GGUF
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default IsValorum/KAT-Coder-V2.5-Dev-APEX-I-MiniPlus-V2-GGUF
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use IsValorum/KAT-Coder-V2.5-Dev-APEX-I-MiniPlus-V2-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf IsValorum/KAT-Coder-V2.5-Dev-APEX-I-MiniPlus-V2-GGUF
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "IsValorum/KAT-Coder-V2.5-Dev-APEX-I-MiniPlus-V2-GGUF" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
ARCHITECTURE SELECTION GUIDE — MINIPLUS V2 & V2.1 EDITIONS
This repository hosts the MiniPlus V2 edition of KAT-Coder-V2.5-Dev. Our releases are precision-engineered for specific hardware budgets and memory topologies. V2 is NOT obsolete or "worse"; each edition serves distinct inference requirements:
- MiniPlus V2 (High Theoretical Layer Protection): On paper, V2 provides extra protective envelopes on edge layers (10 layers in
IQ3_S+IQ4_NLshared experts +Q8_0attention gates). However, in practical inference benchmarks—even across extreme long-context windows exceeding +160K tokens—there is virtually NO perceptible difference in quality or reasoning compared to V2.1.- MiniPlus V2.1 (System RAM Streaming Specialist with Deep Context): Specially prepared to run totally or partially in system RAM (DDR4/DDR5) across large codebase contexts (up to 256k tokens). By replacing non-linear codebooks with linear
Q3_Kedge experts, keepingQ8_0attention gates, and upgrading shared foundation experts toQ5_Kacross all 40 layers, it completely eliminates AVX2 CPU dequantization stalls (+24 to 28+ tok/s streaming). Depending on your processor and memory bandwidth (DDR4/DDR5), streaming generation in system RAM can be almost as fast as having everything in VRAM, while supporting deep context reserving GPU VRAM for the codebase KV cache while model weights stream from system RAM. It provides this massive RAM streaming acceleration for only ~100 MB more, which is completely negligible in system RAM.Which one should you choose?
- If you offload 100% into GPU VRAM (24GB+ VRAM,
-ngl 99): Both V2 and V2.1 run blistering fast on GPU tensor cores with virtually identical top-tier intelligence. V2 is an exceptional build for full VRAM offload.- If you run with most/all layers in system RAM (DDR4/DDR5): V2.1 is strongly recommended to eliminate CPU AVX2 lookup latency and achieve peak streaming speeds.
Both editions are handcrafted and vastly outperform flat 3-bit quants and generic community APEX-I-Mini releases. To explore or download the V2.1 edition of KAT-Coder-V2.5-Dev optimized for system RAM streaming, visit: IsValorum/KAT-Coder-V2.5-Dev-APEX-I-MiniPlus-V2.1-GGUF
DO NOT CONFUSE APEX-I-MINIPLUS WITH GENERIC COMMUNITY APEX-I-MINI!
Regardless of release version (whether V1, V2, or V2.1), NEVER confuse handcrafted APEX-I-MiniPlus builds with generic community APEX-I-Mini releases:
- Generic Community APEX-I-Mini: Uniformly compresses all core MoE experts down to aggressive 2-bit
IQ2_S(dropping below the critical quality floor), leaves the sensitive token output head unarmored at 3-bitQ3_K_M, and compresses attention projections down toQ3_K. In deep reasoning models, this triggers severe perplexity spikes, syntax errors, and broken code brackets.- Handcrafted APEX-I-MiniPlus (All Editions by IsValorum): Every single MiniPlus release—from V1 and V2 to V2.1—is a custom tensor-by-tensor architecture that preserves uncompressed
F32router gates, armors the token output head in high-precisionQ6_K, safeguards attention gates inQ8_0, and keeps core reasoning experts at or above calibrated 3-bit (IQ3_XXS/IQ3_S). Even our earlier builds vastly outperform generic community APEX recipes and flat 3-bit quants.
Quick Navigation Index
- Model Files & Technical Specifications
- Comparative Quantization Analysis (vs. Flat Quants & Generic APEX)
- Everyday Laptop Benchmarks (DDR4 / DDR5 RAM)
- The 24GB Miracle: Full 256K Context Runs In VRAM!
- Hardware Throughput Projections (RTX 30 / 40 / 50)
- The Speed vs. Precision Trade-off
- Surgical Tensor Quantization Map
- Recommended Configuration & Setup
Model Files & Technical Specifications
| File Name | File Size | Memory Footprint | BPW | Description |
|---|---|---|---|---|
KAT-Coder-V2.5-Dev.APEX-I-MiniPlus-V2.gguf |
14.64 GB (13.64 GiB) |
13.64 GiB |
3.38 BPW | Core agentic code synthesis, syntax verification, refactoring & logic |
- Base Architecture:
Qwen3_5MoeForConditionalGeneration(40 layers, 256 fine-grained micro-experts with intermediate dimension 512, 8 active per token). - Active Parameters: approx. 3.2B active parameters per token (delivering small-model throughput with 35B-scale reasoning).
- Quantization Profile: Armored boundary layers (
IQ3_S/IQ4_NL), deep core expert compression (IQ3_XXS+imatrix), uncompressed router gates (F32), and high-precision syntax output head (Q6_K). - Memory Footprint: Ultracompact 13.64 GiB footprint engineered specifically to avoid Out-Of-Memory (OOM) crashes on 16GB and 24GB VRAM hardware.
Comparative Quantization Analysis (vs. Flat Quants & Generic APEX)
Also, don't confuse APEX-I-MiniPlus-V2 with a generic baseline APEX-I-Mini. Traditional APEX-I-Mini drops core experts aggressively to 2-bit IQ2_S and leaves output.weight at 3-bit Q3_K_M, which creates a noticeable perplexity hit on complex reasoning tasks. V2 was specifically re-engineered to avoid that quality floor (keeping core experts at calibrated IQ3_XXS, output in Q6_K, shared expert in non-linear IQ4_NL, and routers in F32).
To put the numbers in perspective: this cuts nearly 2 GB off a flat 3-bit quant (approx. 15.6 GB), and weighs only about approx. 1 GB more than a generic APEX-I-Mini (approx. 12.5 GB). For that single extra gigabyte of VRAM, you get a massive jump in reasoning and syntactic stability.
Take a look at the tensor-by-tensor comparison table below to inspect the exact architectural differences and see why this specific allocation is optimal. That's specifically what this was built for:
| Architectural Component | Generic Automated Quants (Flat Q3_K_S / IQ3_S) |
Generic APEX-I-Mini (Baseline Recipe) | Our Handcrafted APEX-I-MiniPlus-V2 (IsValorum) | Perceived Quality & Real-World Impact |
|---|---|---|---|---|
Output Head (output.weight) |
Flat IQ3_S / Q3_K_S (approx. 3.44 BPW) |
Inherits base type Q3_K_M (approx. 3.44 BPW unarmored) |
Q6_K (approx. 6.56 BPW uncompromised) |
Eliminates Syntax & Vocabulary Hallucinations: Low-bit output heads cause tokenizer classification noise, breaking code indentation, brackets ({}, []), math symbols, and domain terms. Q6_K preserves near-FP16 output classification. |
Expert Routers (ffn_gate_inp.weight) |
Blindly quantized to 3-bit / unoptimized | Inherits base type Q3_K_M (approx. 3.44 BPW compressed) |
F32 uncompressed (32.0 BPW, 2 MB/layer) |
Zero Router Drift: In micro-expert models, even minuscule quantization errors in router logits misdirect tokens to wrong experts. Retaining uncompressed F32 guarantees 100% routing fidelity with virtually zero memory overhead (approx. 80 MB total). |
Attention & Language (attn_output, attn_qkv) |
Flat IQ3_S / Q3_K_S |
Q3_K on 34 middle layers (L3–36), Q4_K on 6 edge layers |
Q6_K for attn_output, IQ3_S for attn_qkv |
Contextual Retrieval Precision: Generic APEX reduces attention and language projections to Q3_K across 85% of layers. Our V2 build protects attention output in high-precision Q6_K and uses calibrated non-linear IQ3_S, ensuring flawless needle-in-a-haystack retrieval across deep 128k–256k context windows. |
Attention Gates (attn_gate.weight) |
Blindly compressed to 3-bit | Compressed to Q3_K (middle) / Q4_K (edges) |
Q8_0 (8.50 BPW) |
Attention Head Stability: Attention gates modulate query-key routing across hybrid attention layers. Keeping them in 8-bit prevents attention crosstalk and hallucination over long contexts. |
Shared Foundation Expert (ffn_*_shexp) |
Flat IQ3_S / Q3_K_S (3.44 BPW) |
Linear Q4_K (middle) / Q5_K (edges) |
IQ4_NL (4.50 BPW non-linear codebook) |
Foundational Knowledge Armor: The shared expert executes for 100% of tokens. In 256 micro-expert models, IQ4_NL non-linear codebooks preserve heavy-tailed outlier representations far better than standard linear quantization. |
| Core MoE Layers (Middle: 10–29) | Flat IQ3_S / Q3_K_S (uniform bit-rate across all layers) |
Aggressive IQ2_S (2.50 BPW) |
IQ3_XXS (3.06 BPW) + calibrated imatrix |
Above the Quality Threshold: Generic 2-bit IQ2_S baselines drop below the critical quality floor for 35B MoEs, resulting in perplexity spikes on reasoning tasks. Our IQ3_XXS with imatrix achieves deep compression (272 MiB → 98 MiB per block) without sacrificing logic. |
| Edge MoE Layers (Layers 0–9 & 30–39) | Flat IQ3_S / Q3_K_S (no layer-wise gradient) |
Q3_K (limited to first/last 5 layers only: L0–4, L35–39) |
IQ3_S (expanded to 10 input & 10 output layers) |
Protected Ingestion & Synthesis: Half of the model's layers (10 at input, 10 at output) form a non-linear armored envelope, preventing prompt misunderstanding and token degeneration across 256 micro-experts. |
| Normalization & Biases | Often degraded | Standard | F32 uncompressed |
Numerical Stability: Prevents cumulative floating-point underflow/overflow across deep 40-layer computation. |
Everyday Laptop Benchmarks (DDR4 / DDR5 RAM)
Estimated Projections on Consumer Hardware
You do not need an expensive workstation to run a cutting-edge 35B Mixture-of-Experts coding model. Estimated throughput projections on a standard consumer laptop (Intel Core i5 / AMD Ryzen, 4GB/6GB Laptop GPU, 32GB DDR4/DDR5 RAM):
- GPU VRAM Allocation: Uses only approx. 3.8 GB VRAM (fits effortlessly on budget 4GB/6GB laptop GPUs such as RTX 3050, 4050, or 2060).
- System Memory Offload: Standard 32GB system RAM accommodates the remaining layers.
- Estimated Document / Code Ingestion (Prefill): 300 to 450+ tokens/second sustained across long prompt files.
- Estimated Streaming Generation: 20 to 24+ tokens/second sustained output across system RAM!
Pro Tip for Consumer Laptop Users: Because the bulk of the model runs from system memory in partial offload mode, standard autoregressive generation streams seamlessly at 20 to 24+ tokens/second across everyday DDR4/DDR5 memory buses, perfectly sufficient for real-time IDE pair programming!
The 24GB Miracle: Full 256K Context Runs In VRAM!
For developers running 24GB GPUs (RTX 3090, RTX 4090, or professional workstations), standard community 3-bit or 4-bit quants weigh 15.8 to 19.5 GiB in weights alone. When combined with KV cache and compute buffers for large codebases, they trigger immediate CUDA Out-Of-Memory crashes.
KAT-Coder APEX-I-MiniPlus-V2 fits massive contexts entirely within 24GB VRAM:
| Context Length | Model Weights (Est.) | KV Cache (q8_0, 4 slots) | Compute Buffers | Total GPU VRAM (Est.) | Hardware Feasibility |
|---|---|---|---|---|---|
| 32,768 (32k) | 13.64 GiB |
0.58 GiB |
1.80 GiB |
16.02 GiB |
Full offload on 24GB; partial on 16GB |
| 65,536 (64k) | 13.64 GiB |
0.92 GiB |
1.95 GiB |
16.51 GiB |
Effortless fit on 24GB GPUs |
| 131,072 (128k) | 13.64 GiB |
1.58 GiB |
2.22 GiB |
17.44 GiB |
Effortless fit on 24GB GPUs |
| 262,144 (256k) | 13.64 GiB |
2.92 GiB |
2.80 GiB |
19.36 GiB |
FULL 256K CODE REPO IN VRAM! |
Note: Projections estimate approx. 4.64 GiB of headroom remaining on 24GB cards for system display buffers and tooling.
Hardware Throughput Projections (RTX 30 / 40 / 50)
When running with full GPU offload (-ngl 99), KAT-Coder's fine-grained MoE architecture (approx. 3.2B active parameters) unlocks extraordinary generation throughput:
| Hardware Target | Offload Mode | Generation Speed (Est.) | Prompt Prefill Speed (Est.) | Engineering Highlights |
|---|---|---|---|---|
| NVIDIA RTX 5080 / 5090 (Blackwell) | Full GPU (-ngl 99) |
110 – 135+ tok/s | 2,500 – 3,600+ tok/s | Blistering throughput on next-gen memory bandwidth |
| NVIDIA RTX 4090 (24GB GDDR6X) | Full GPU (-ngl 99) |
80 – 105+ tok/s | 1,800 – 2,600+ tok/s | Near-instantaneous code completion & refactoring |
| NVIDIA RTX 3090 (24GB GDDR6) | Full GPU (-ngl 99) |
65 – 80+ tok/s | 1,400 – 2,000+ tok/s | Full 256k repository context in dedicated VRAM |
| NVIDIA RTX 4080 / 5070 (16GB) | Partial offload (approx. 30 layers) | 35 – 45+ tok/s | 800 – 1,200+ tok/s | High-efficiency local coding assistant |
| Consumer Laptop (4GB GPU + 32GB RAM) | Hybrid Offload | 20 – 24+ tok/s | 300 – 450+ tok/s | Smooth streaming from system DDR4/DDR5 RAM |
- Downloads last month
- 919
We're not able to determine the quantization variants.
Model tree for IsValorum/KAT-Coder-V2.5-Dev-APEX-I-MiniPlus-V2-GGUF
Base model
Kwaipilot/KAT-Coder-V2.5-Dev