Model Card for Orca Sonar Document Classifier Edge

Edge/quantized ONNX variant of patronus-studio/orca-sonar-document-classifier. Private / internal.

Quantized with the Patronus RunPod ONNX method (runpod/quantize_single_heads.py): quantize_dynamic for INT8 linears + MatMulNBitsQuantizer (4-bit, block 128) for the embedding. The parent FP32/FP16 model is unchanged; this repo only adds the small variants.

7-class document topic classifier (mmBERT-small / ModernBERT). Metrics measured on the held-out document-classifier test set; FP32 reproduces the parent card (macro-F1 0.978).

Variants (measured, essentially lossless)

Variante Größe macro-F1 Δ vs FP32 Accuracy
fp16 282 MB 0.9783 +0.0000 0.9783
int8 142 MB 0.9793 +0.0011 0.9792
int8_int4_embeddings 96 MB 0.9790 +0.0007 0.9792

Recommended: onnx/int8_int4_embeddings/model.onnx — ~5.5× smaller than FP32 at ~0 quality loss.

Files

  • onnx/fp16/model.onnx, onnx/int8/model.onnx, onnx/int8_int4_embeddings/model.onnx
  • config.json, tokenizer.json, tokenizer_config.json (from the parent)
  • metrics/quant_bench.json — the full FP32-vs-quant benchmark

Usage

import onnxruntime as ort, numpy as np
from transformers import AutoTokenizer
tok = AutoTokenizer.from_pretrained("patronus-studio/orca-sonar-document-classifier-edge")
sess = ort.InferenceSession("onnx/int8_int4_embeddings/model.onnx", providers=["CPUExecutionProvider"])
enc = tok("your text", truncation=True, max_length=256, padding="max_length", return_tensors="np")
logits = sess.run(["logits"], {"input_ids": enc["input_ids"].astype(np.int64),
                               "attention_mask": enc["attention_mask"].astype(np.int64)})[0]

License

This model is released under the Apache License 2.0. A copy of the license is included as LICENSE in this repository.

Patronus Ark

This model is built to run inside Patronus Ark, Patronus' open-source on-device AI-security scanning library (L1 native rules → L2 NTDB cascade → L3 transformer). Ark is not publicly released yet — a repository link will be added here at launch.


🛡️ Patronus Protect

Brought to you by Patronus Protect — a local AI firewall that secures every AI interaction (prompts, tools, documents) before it reaches your models. Try it for free at patronus.studio.

Downloads last month
24
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for patronus-studio/orca-sonar-document-classifier-edge

Quantized
(1)
this model

Collection including patronus-studio/orca-sonar-document-classifier-edge