Align
Accurate word timestamps for any transcript.
Word-timestamp refinement for Apple's SpeechAnalyzer pipeline.
- SDKs, install and examples: https://github.com/Desert-Ant-Labs/desert-ant-core/blob/main/docs/models/align.md
Corrects the word-level timings that Apple's SpeechTranscriber and SpeechAnalyzer
return, without replacing them. Align observes the same audio the analyzer already
receives, runs a small Core ML cascade on the CPU and Neural Engine, and returns the
familiar result surface with tightened audioTimeRange values. The models are tiny
(560KB compiled Core ML) and refine a typical result in a few milliseconds
on device.
Apple:
"world"2.61-3.04s ➜ Align:"world"2.57-2.98s
Try it
| Platforms | iOS, macOS, tvOS, visionOS |
| Weights | v1.0.0 |
Install
Swift (requirements)
.package(url: "https://github.com/Desert-Ant-Labs/desert-ant-core.git", from: "3.1.0")
Then add the Align product to your target.
Files
| File | Format | Size | Contents |
|---|---|---|---|
align_coarse.mlmodelc |
Compiled Core ML (FP16) | 285KB | Coarse stage |
align_fine.mlmodelc |
Compiled Core ML (FP16) | 274KB | Fine stage |
mel_filters.bin |
Float32 filter bank | 40KB | Log-mel filter bank the runtime frontend needs |
calibrator.bin |
Gradient-boosted trees | 70KB | Correction calibrator |
refiner_config.json |
JSON | tiny | Runtime config |
The compiled .mlmodelc stages, mel_filters.bin, calibrator.bin, and refiner_config.json
are exactly what the Swift SDK downloads.
Inputs and outputs
- Input: mono audio plus Apple's recognized words with their proposed start/end times.
- Output: the same words with corrected start/end times, or Apple's original time when a correction is not structurally safe.
Accuracy
Measured on v1.0.0 over held-out recordings.
All nine languages
| Condition | Apple raw error | Align error | Reduction |
|---|---|---|---|
| Clean | 124.2ms | 43.9ms | 65% |
| Noisy | 88.3ms | 33.4ms | 62% |
Macro-averaged over the nine languages, so a language with more test data cannot carry the figure on its own.
Public benchmark, English
A 500-clip sample of each official LibriSpeech test-clean and test-other split.
| Engine | Split | Raw | Refined | Reduction | Within 50ms |
|---|---|---|---|---|---|
| Apple SpeechAnalyzer | test-clean | 106.4ms | 20.2ms | 81% | 37% to 95% |
| Apple SpeechAnalyzer | test-other | 111.6ms | 24.8ms | 78% | 35% to 92% |
For editing work the p90 matters more than the mean. Large errors are what a viewer notices when a caption slips or a clip cuts mid-word.
| Split | Raw p90 | Refined p90 |
|---|---|---|
| test-clean | 230.7ms | 33.0ms, about one frame of 30fps video |
Against hand-corrected boundaries
258 word boundaries across 10 recordings, corrected by hand against the waveform rather than by another aligner. The only figure here not measured against machine references.
| System | Error | Within 50ms |
|---|---|---|
| Raw Whisper | 100.8ms | 43% |
| WhisperX | 53.5ms | 67% |
| Align | 45.0ms | 76% |
Languages
English, Spanish, French, Italian, Portuguese, German, Japanese, Korean, and Chinese. A locale outside this set is passed through unchanged.
Limitations
- References are machine forced-alignment estimates, not human annotations, so the figures show a large, consistent reduction of Apple's timing error rather than sample-accurate ground truth.
- A learned correction is not guaranteed to improve every boundary; the structural fallback keeps Apple's timestamp when a correction looks unsafe but cannot catch every plausible-looking error.
- Japanese, Korean, and Chinese were the weakest languages before v1.0.0. They now improve their proposals by 33%, 55%, and 51%.
- Spoken numbers are the weakest remaining case. On a small sample, refinement moved digit boundaries further from the reference than leaving them alone, so treat them as unimproved until a larger sample settles it.
License
Desert Ant Labs Source-Available License. Free for most apps, and a commercial license is required at scale. Full terms are at the link. Licensing: licensing@desertant.com.
Citation
@software{align_2026,
title = {Align: Word-timestamp refinement for Apple's SpeechAnalyzer pipeline},
author = {Desert Ant Labs},
year = {2026},
url = {https://huggingface.co/desert-ant-labs/align},
}
© 2026 Desert Ant Labs · https://desertant.com