--- license: apache-2.0 language: - en - ru tags: - text-classification - feature-extraction - vocabulary-filtering - denoising - multilingual - quantized - int8 - openvino - tensorrt pipeline_tag: text-classification base_model: - intfloat/multilingual-e5-small --- # code-daemon-denoise-v1 A **bilingual (EN + RU) word filter**: given one word form, it answers whether that word is a **meaningful technical term** worth keeping in a search vocabulary, or **ballast** to drop. It is deliberately small and one-purpose. A frozen `multilingual-e5-small` encoder produces a 384-dim vector, and a **single trained affine** turns that vector into `P(keep)`. No fine-tuning of the encoder, no classification head with its own weights to load — the entire learned decision is 384 numbers and a bias, shipped as a 6 KB JSON file. That buys throughput: **~800 words/sec on a CPU core**, **~17 800/sec on a laptop GPU**. ```python emb = session.run(None, {"input_ids": ids, "attention_mask": mask})[0] # [B, 384] pooled + L2 p_keep = 1 / (1 + np.exp(-(emb @ w + b))) # the whole classifier ``` --- ## 1. What it is for Vocabulary hygiene. Harvest every word form out of a codebase — identifiers, doc prose, comments, commit messages — and most of what you get is not worth indexing: inflected function words, chopped identifier fragments, transliteration noise, boilerplate. Keeping them inflates a search vocabulary and dilutes term statistics; dropping them by frequency alone throws away rare-but-real technical terms, which are exactly the ones worth searching for. This model makes that call per word, in both English and Russian, at a rate that keeps up with a full-repository scan. **Suited to** - Filtering a harvested vocabulary before indexing. - Any per-token keep/drop decision over short, single-word inputs. - Mixed EN/RU corpora — including Cyrillic identifiers and comments. **Not suited to** - Sentences or phrases. Inputs are single word forms; the sequence budget is 40 tokens. - Languages outside Latin/Cyrillic scripts — the vocabulary was pruned to those on purpose. - Domain term-vs-stopword calls outside software; the label set is technical-corpus flavoured. --- ## 2. Architecture | | | |---|---| | Encoder | [`intfloat/multilingual-e5-small`](https://huggingface.co/intfloat/multilingual-e5-small) — XLM-RoBERTa, **frozen, unchanged** | | Embedding dim | 384, mean-pooled and L2-normalised **inside the graph** | | Vocabulary | **142k** pieces, pruned from 250k by character class (Latin + Cyrillic + punctuation) | | Classifier | one affine: `P(keep) = sigmoid(w·e + b)`, `w ∈ ℝ³⁸⁴` | | Sequence | 40 tokens, batch 64 | | Inputs | `input_ids`, `attention_mask` | | Output | `[batch, 384]` — pooled, normalised, ready for the dot product | ### Two decisions that make it small **The encoder is frozen.** The head is a logistic regression fitted on top of fixed embeddings, then folded — its `StandardScaler` and the LR coefficients are multiplied out into a single `(w, b)` pair. There is no scikit-learn at inference, and no second model to keep in sync: the decision boundary is a dot product you can apply in any language. **The vocabulary is pruned by script.** Cutting the 250k multilingual SentencePiece table to the Latin + Cyrillic + punctuation pieces removes ~43% of the rows, and the embedding table is most of this model's weight. The pruned-vocab id remap is baked into the graph as a Gather at the input, so callers still feed ordinary SentencePiece ids and never see the mapping. INT8 weights drop from ~121 MB to **~76 MB** — lossless for the two languages it targets, because nothing outside those scripts was reachable anyway. ### The `"vocab: "` prefix Words are embedded with a fixed `"vocab: "` prefix. The head was trained on prefixed embeddings, so **reproduce the prefix** for standalone use or the decision boundary will not line up. --- ## 3. How it was made 1. **Encoder** — export the frozen mE5-small to ONNX with mean-pooling and L2-norm fused into the graph, prune the embedding table to the kept character classes, and PTQ-quantize to INT8 (NNCF). 2. **Head** — embed a bilingual labelled word set (English: WordNet / BNC mid-frequency lemmas; Russian: Taiga / OpenCorpora / Nerus mid-Zipf) plus per-language hand-checked gold, fit `LogisticRegression(class_weight="balanced")`, then fold the scaler and the LR into one affine. `strip_threshold` (default **0.95**) sets where you cut. It is high on purpose: dropping a real technical term is the expensive error, keeping a bit of ballast is not. --- ## 4. Speed Measured on one laptop: Intel Core Ultra 9 275HX / NVIDIA RTX 5060 Laptop, batch 64 × seq 40. | lane | per batch | throughput | per word | |---|--:|--:|--:| | **TensorRT FP16, RTX 5060 Laptop** | **3.60 ms** | **17 790 words/s** | 0.056 ms | | **OpenVINO INT8, CPU** | **80.4 ms** | **796 words/s** | 1.26 ms | | ONNX Runtime FP32, CPU | 188 ms | 341 words/s | 2.93 ms | The INT8 CPU lane is the intended default — 800 words/sec is enough to filter a repository's whole harvested vocabulary in seconds without touching a GPU, and it is 2.3× the unquantized ONNX path. The GPU lane exists for hosts that have spare VRAM anyway. --- ## 5. Standalone use ```python import json, numpy as np, onnxruntime as ort, sentencepiece as spm sp = spm.SentencePieceProcessor(model_file="sentencepiece.bpe.model") sess = ort.InferenceSession("model.onnx", providers=["CPUExecutionProvider"]) head = json.load(open("denoise_head.json")) # {dim, w[384], b, strip_threshold} w, b, thr = np.array(head["w"], np.float32), head["b"], head["strip_threshold"] def p_keep(words, max_len=40): toks = [[2, *sp.encode("vocab: " + x)[: max_len - 2], 3] for x in words] # bos … eos L = max(len(t) for t in toks) ids = np.array([t + [0] * (L - len(t)) for t in toks], dtype=np.int64) # pad = 0 mask = (ids != 0).astype(np.int64) emb = sess.run(None, {"input_ids": ids, "attention_mask": mask})[0] # [B, 384] return 1.0 / (1.0 + np.exp(-(emb @ w + b))) scores = p_keep(["mutex", "tensorrt", "пожалуйста", "asdfgh"]) keep = scores >= thr ``` The ONNX bakes in the fairseq `+1` id offset and the pruned-vocab remap, so feed raw SentencePiece ids — do not remap them yourself. --- ## 6. Evaluation On a frozen held-out bilingual word set, at `strip_threshold = 0.95`: | metric | value | |---|--:| | SAFE (keep) F1 | **0.79** | | BALLAST (drop) F1 | **0.84** | | Strip precision | **0.88** | The INT8 vocab-pruned build scores the same as the full-vocabulary FP build (F1 0.79 vs 0.79) at 38% of the size — the pruning removes rows the two target languages never reach, so there is nothing to lose by it. Strip precision is the number to watch if you tune the threshold: it says how often a word the model drops really was ballast. --- ## 7. What is in this repo - **OpenVINO INT8** — `code-daemon-denoise-v1-s_ov2026.2_{cpu,igpu_lnl}_int8_b64_s40.{xml,bin}` — the default lane (CPU) and an Intel iGPU build. - **OpenVINO INT4, NPU** — `code-daemon-denoise-v1-s_ov2026.2_npu_int4_b16_s40.{xml,bin}` — weight-only INT4 at batch 16 for Intel NPUs. - **TensorRT FP16** — `code-daemon-denoise-v1-s_{win_x64,linux_x64}_trt11.0_sm_120.engine`. - **TVM Vulkan** — `code-daemon-denoise-v1_{win_x64,linux_x64}_tvm0.25_vulkan.{dll,so}` — GPU fallback for non-NVIDIA hardware. - **Head** — `denoise_head.json`. **Required**: the ONNX alone emits embeddings, not a decision. - **Tokenizer** — `sentencepiece.bpe.model`, `tokenizer_config.json`. - **ONNX** — `model.onnx`, FP32, pruned, with mean-pool + L2-norm + id-remap fused. The build source for every engine above and the path for standalone `onnxruntime` use. --- ## 8. License & attribution The encoder weights are **[`intfloat/multilingual-e5-small`](https://huggingface.co/intfloat/multilingual-e5-small)** (Apache-2.0), redistributed here in compiled form **unchanged**; this repository is therefore released under **Apache-2.0**. The trained head and the build/quantization tooling are original. Backbone: XLM-RoBERTa. Not legal advice. Used by the [UltraCode](https://github.com/faxenoff/ultracode) code assistant, though nothing about the model is specific to it.