Audio-to-Audio
MLX
Safetensors
deepfilternet3
speech-enhancement
denoising
apple-silicon
deepfilternet
Instructions to use aufklarer/DeepFilterNet3-MLX with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use aufklarer/DeepFilterNet3-MLX with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] huggingface-cli download --local-dir DeepFilterNet3-MLX aufklarer/DeepFilterNet3-MLX
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
File size: 5,674 Bytes
8e4ece5 8f9ba6a 8e4ece5 8f9ba6a 8e4ece5 8f9ba6a 8e4ece5 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 | ---
license: apache-2.0
tags:
- speech-enhancement
- denoising
- mlx
- apple-silicon
- deepfilternet
- safetensors
base_model: Rikorose/DeepFilterNet3
library_name: mlx
pipeline_tag: audio-to-audio
---
# DeepFilterNet3 β MLX
Real-time speech enhancement for Apple Silicon. Removes background noise from
speech audio. FP32 MLX weights converted from the official DeepFilterNet3
checkpoint, with BatchNorm fused into the convolutions.
- **2.1M params**, fp32 safetensors, **8.1 MB**
- 48 kHz native, 10 ms frames, 2-frame (20 ms) lookahead
- Network-only export: STFT, ERB feature extraction, and deep-filter
application live in the runtime (`auxiliary.npz` ships the exact DSP
constants)
## Model
| Detail | Value |
|--------|-------|
| Architecture | DeepFilterNet3 (encoder + ERB decoder + deep-filter decoder) |
| Parameters | 2,131,824 (BatchNorm fused, GRU biases folded) |
| Precision | float32 |
| Sample rate | 48 kHz (FFT 960, hop 480) |
| ERB bands / DF bins | 32 / 96 |
| Deep-filter order | 5, lookahead 2 frames |
## Files
| File | Size | Description |
|------|------|-------------|
| `model.safetensors` | 8.1 MB | fp32 network weights in MLX layouts |
| `auxiliary.npz` | 126 KB | ERB filterbank + inverse, Vorbis window, normalization init states |
| `config.json` | 1 KB | Model and DSP hyperparameters |
| `dfn3_mlx.py` | 9 KB | Pure-MLX reference implementation of the network |
## Quality
30 VoiceBank-DEMAND test clips (16 kHz mirror, resampled to 48 kHz for the
model); the neural network forward runs on MLX while STFT / ERB / deep-filter
post-processing stay in the upstream PyTorch pipeline β the same methodology
used for the [CoreML variant](https://huggingface.co/aufklarer/DeepFilterNet3-CoreML).
| Backend | PESQ | STOI | SI-SDR |
|---------|------|------|--------|
| Noisy input | 2.205 | 0.932 | 9.28 |
| PyTorch FP32 (reference) | 2.900 | 0.947 | 18.19 |
| **MLX, CPU stream** | **2.900** | **0.947** | **18.19** |
| **MLX, GPU stream** | **2.902** | **0.947** | **18.18** |
Network outputs match PyTorch within 9e-7 on the MLX CPU stream β
numerically exact for fp32. The GPU stream uses Metal fast-math
transcendentals (max output delta ~4e-3) with no measurable metric impact.
## Latency (Apple M5 Pro, network forward only, Python MLX)
| Duration | GPU | RTF | CPU | RTF |
|----------|-----|-----|-----|-----|
| 5 s | 0.14 s | 0.028 | 0.18 s | 0.036 |
| 10 s | 0.29 s | 0.029 | 0.40 s | 0.039 |
| 20 s | 0.54 s | 0.027 | 0.79 s | 0.040 |
## Usage
```python
import sys
import mlx.core as mx
from huggingface_hub import snapshot_download
model_dir = snapshot_download("aufklarer/DeepFilterNet3-MLX")
sys.path.append(model_dir)
from dfn3_mlx import DFN3MLX
model = DFN3MLX(model_dir)
# Normalized features from your DSP front-end (constants in auxiliary.npz):
feat_erb = mx.zeros((1, 100, 32, 1)) # ERB features, dB-scaled + mean-normalized
feat_spec = mx.zeros((1, 100, 96, 2)) # complex spectrum (real, imag), unit-normalized
erb_mask, df_coefs, lsnr = model(feat_erb, feat_spec)
```
Full audio-in/audio-out enhancement additionally needs the DeepFilterNet DSP
path (STFT β features β ERB mask + deep filtering β iSTFT), documented in the
[upstream repository](https://github.com/Rikorose/DeepFilterNet).
On Apple devices, [speech-swift](https://github.com/soniqo/speech-swift)
runs this model directly via its MLX engine (or use the
[CoreML variant](https://huggingface.co/aufklarer/DeepFilterNet3-CoreML) for
the Neural Engine):
```swift
import SpeechEnhancement
let enhancer = try await SpeechEnhancer.fromPretrained(engine: .mlx)
let clean = try enhancer.enhance(audio: noisyAudio, sampleRate: 48000)
```
CLI:
```bash
swift run speech denoise noisy.wav --engine mlx
```
## Tensor contract
Channels-last `[B, T, F, C]` β time is the conv H axis, frequency the W axis.
The 2-frame lookahead shift is applied to the inputs inside the model.
| Direction | Name | Shape | Notes |
|-----------|------|-------|-------|
| input | `feat_erb` | `[B, T, 32, 1]` | ERB features, dB-scaled, exp-mean-normalized |
| input | `feat_spec` | `[B, T, 96, 2]` | complex spectrum features, unit-normalized |
| output | `erb_mask` | `[B, T, 32, 1]` | sigmoid ERB gain mask |
| output | `df_coefs` | `[B, 5, T, 96, 2]` | deep-filter coefficients (order, real/imag) |
| output | `lsnr` | `[B, T, 1]` | local SNR estimate, dB in [-15, 35] |
## Weight conventions
- **Conv2d**: `[O, kH, kW, I/groups]` (PyTorch `[O, I/g, kH, kW]` transposed);
causal time padding, symmetric frequency padding
- **Depthwise ConvTranspose2d** (`erb_dec.convt{1,2}.dwt`): `[C, kH, kW, 1]`,
true transposed-conv kernels (not pre-flipped); k=(1,3), freq stride 2,
padding 1, output padding 1
- **GRU** (`mlx.nn.GRU` convention, gate order r/z/n): `Wx = weight_ih`,
`Wh = weight_hh`, `b = bias_ih + [bias_hh_r; bias_hh_z; 0]`,
`bhn = bias_hh_n`; run with an explicit zero initial hidden state
- **Grouped linear**: `[groups, in/groups, out/groups]`, applied as
`einsum("btgi,gih->btgh")` then flattened
- BatchNorm is fused into the preceding conv; the unused `df_fc_a` head is
dropped
## Source
- Base model: [Rikorose/DeepFilterNet](https://github.com/Rikorose/DeepFilterNet)
v0.5.6 checkpoint (Apache-2.0 / MIT)
- Paper: [DeepFilterNet3](https://arxiv.org/abs/2305.08227)
## License
- Model weights: Apache-2.0 / MIT dual license
- MLX conversion: Apache-2.0
## Links
- [speech-swift](https://github.com/soniqo/speech-swift) β Apple SDK
- [Docs](https://soniqo.audio/getting-started) β install and CLI docs
- [soniqo.audio](https://soniqo.audio) β website
- [blog](https://soniqo.audio/blog) β blog
|