MauBERT quantisers
The quantisation models used to produce the pseudo-labels of
the MauBERT paper (ACL 2026). Each of them maps the frame-level features
of one layer of coml/maubert-feat,
coml/maubert-phone or the HuBERT base model to discrete units, and was
fitted on 10 hours of one language.
The fine-tuned models themselves are not released: use maubert to extract the
features of a pre-trained model, then one of these quantisers to transcribe them into units.
Layout
<model>/<dataset>/<language>/<n_clusters>cl.<ext>
manifest.tsv lists every file with its base model, layer, quantisation method, dataset, language and number of
clusters.
| model | base model | layer | method | file |
|---|---|---|---|---|
fr-feat-simple |
MauBERT-feat | feat_proj |
feat-freq |
.pt |
fr-kmeans-feature |
MauBERT-feat | feat_proj |
kmeans |
.km |
fr-kmeans-hubert_9 |
MauBERT-feat | HuBERT layer 10 | kmeans |
.km |
fr-phone-ft |
MauBERT-feat | phone_proj |
phone-ft |
.ft |
kmeans-hubert_11 |
HuBERT base | HuBERT layer 12 | kmeans |
.km |
pr-kmeans-phone |
MauBERT-phone | phone_proj |
kmeans |
.km |
pr-kmeans-ups_proj |
MauBERT-phone | ups_proj |
kmeans |
.km |
pr-phone-ft |
MauBERT-phone | phone_proj |
phone-ft |
.ft |
pr-phone-simple |
MauBERT-phone | phone_proj |
phone-simple |
.pt |
The .km files are joblib pickles of a scikit-learn MiniBatchKMeans, the .ft files are torch state dicts of
the adapted phone-prediction layer, and the other .pt files are torch tensors.
Datasets and languages
| dataset | languages |
|---|---|
discophon |
cmn, deu, eng, eus, fra, jpn, swa, tam, tha, tur, ukr, wol |
dev-languages |
sw, ta, th, tr, uk |
zrc2017 (ZeroSpeech 2017 test) |
english, french, german (LANG1), mandarin, wolof (LANG2) |
casual-languages |
english, french |
Every quantiser comes in a 100-cluster variant (370 for pr-phone-simple, which uses the full phone inventory of
MauBERT-phone). The discophon languages additionally have a variant with as many clusters as the language has phones:
cmn 43, deu 42, eng 40, eus 30, fra 35, jpn 43, swa 30, tam 30, tha 41, tur 28, ukr 36,
wol 40.
Usage
Extract the features of the corresponding layer with maubert, then transcribe
them with the quantiser:
hf download coml/maubert-quantisers fr-kmeans-hubert_9/discophon/deu/100cl.km --local-dir quantisers
python -m maubert.features MANIFEST_FILE FEATURE_DIR \
--type maubert-feat --checkpoint coml/maubert-feat --layer 9
python -m maubert.transcribe FEATURE_DIR \
quantisers/fr-kmeans-hubert_9/discophon/deu/100cl.km units.jsonl kmeans
Citing
@inproceedings{ortiztandazo-etal-2026-maubert,
title = "{M}au{BERT}: Universal Phonetic Inductive Biases for Few-Shot Acoustic Units Discovery",
author = "Ortiz Tandazo, Angelo and
Khentout, Manel and
Benchekroun, Youssef and
Hueber, Thomas and
Dupoux, Emmanuel",
editor = "Liakata, Maria and
Moreira, Viviane P. and
Zhang, Jiajun and
Jurgens, David",
booktitle = "Proceedings of the 64th Annual Meeting of the {A}ssociation for {C}omputational {L}inguistics (Volume 1: Long Papers)",
month = jul,
year = "2026",
address = "San Diego, California, United States",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2026.acl-long.24/",
doi = "10.18653/v1/2026.acl-long.24",
pages = "568--585",
ISBN = "979-8-89176-390-6",
}
Model tree for coml/maubert-quantisers
Base model
coml/maubert-feat