You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

Tigre HuBERT

A self-supervised speech representation model for Tigre (ISO 639-3: tig), a Semitic language spoken primarily in Eritrea. This is, to our knowledge, one of the first publicly available speech models for Tigre.

Status: this is a pretrained representation model, not a fine-tuned task model. It hasn't yet been benchmarked on a downstream task (e.g. ASR word/character error rate). Validation loss during pretraining improved consistently and has plateaued (see Training below), which is a good sign for representation quality, but the honest way to know how well it performs on a specific task is to fine-tune it for that task and measure the result directly. If you do that, we'd love to hear how it goes.

What this model is

HuBERT (Hidden-Unit BERT) learns speech representations by predicting cluster assignments of masked audio segments, without needing any transcribed text. That makes it well-suited to low-resource languages like Tigre, where labeled speech-to-text data is scarce but raw audio is more attainable.

The output of this model is not text β€” it's a sequence of learned embeddings per audio frame. To get something task-specific (e.g. speech recognition, speaker identification, language identification), you fine-tune a small task head on top of these representations.

How it was trained

  • Data: ~500 hours of Tigre speech audio.
  • Recipe: the standard two-iteration HuBERT pretraining procedure (fairseq's implementation):
    • Iteration 1: pseudo-labels from k-means (k=100) on raw MFCC features β€” a weak but label-free starting signal.
    • Iteration 2: the iteration-1 model's own layer-6 hidden features are re-clustered (k=500) into new, better pseudo-labels, and a fresh model is trained from scratch against those. This is the standard recipe's main quality jump, since the model now learns from speech-aware clusters instead of raw acoustic ones.
  • Architecture: HuBERT-base (~90M parameters, 12 transformer layers).
  • Note on scale: the original HuBERT paper validated this exact 2-iteration recipe on 960 hours of English speech, and only used a 3rd iteration at a much larger scale (60,000+ hours). At ~500 hours, this model is trained at roughly half the data scale the base recipe was designed for β€” a reasonable and appropriately-sized recipe for this amount of data, but worth keeping in mind when setting quality expectations relative to large, high-resource-language HuBERT models.

Iteration-2 training curve

Validation loss over iteration-2 pretraining

Validation loss (the masked-unit prediction objective, not a downstream task metric) over iteration-2 pretraining, from epoch 60 to the point training was stopped once improvement plateaued (under 1% change over the last few checkpoints). This shows the pretraining objective converging smoothly β€” it does not measure accuracy on any specific downstream task.

Basic usage

This model outputs frame-level embeddings, not text. Example: extracting features from a 16kHz mono audio clip.

import torch
from fairseq import checkpoint_utils

# Download checkpoint_best.pt from this repo first
models, cfg, task = checkpoint_utils.load_model_ensemble_and_task(["checkpoint_best.pt"])
model = models[0].eval()

# waveform: torch.FloatTensor of shape [1, num_samples], 16kHz, mono
with torch.inference_mode():
    features, _ = model.extract_features(
        source=waveform,
        padding_mask=None,
        mask=False,       # inference, not the masked-training objective
        output_layer=None # final transformer layer
    )

# features: [1, num_frames, hidden_dim] -- one embedding per ~20ms of audio

fairseq isn't a well-maintained PyPI package β€” install it from source:

git clone https://github.com/facebookresearch/fairseq.git
cd fairseq && pip install --editable .

What this is useful for

  • Fine-tuning for Tigre ASR β€” attach a CTC head and fine-tune on a (even fairly small) labeled Tigre speech-to-text dataset.
  • Speaker or language identification β€” the embeddings can feed a lightweight classifier for tasks that don't need text transcription at all.
  • A starting point, not an endpoint β€” as a foundation model for further Tigre speech research, in a language with very little existing tooling.

Limitations

  • Not evaluated on any downstream task yet β€” treat performance claims with appropriate skepticism until you've tested it on your own task/data.
  • Trained on ~500 hours, notably less than the 960 hours the base recipe was designed around β€” representation quality may reflect that.
  • No fine-tuned ASR head is included in this repo; this model produces embeddings, not transcriptions, out of the box.

Citation

If you use this model, please cite the original HuBERT paper:

@article{hsu2021hubert,
  title={HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units},
  author={Hsu, Wei-Ning and Bolte, Benjamin and Tsai, Yao-Hung Hubert and Lakhotia, Kushal and Salakhutdinov, Ruslan and Mohamed, Abdelrahman},
  journal={IEEE/ACM Transactions on Audio, Speech, and Language Processing},
  year={2021}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Paper for BeitTigreAI/tigre-hubert