Files
FluidAudio/benchmarks.md
Alexandmiro 0f7493bdac feat: Support Parakeet-TDT-CTC-110M hybrid model (#433)
## Summary
Adds support for NVIDIA's Parakeet-TDT-CTC-110M hybrid model with fused
preprocessor+encoder architecture.

Based on the work by @JarbasAl in #383.

## Key Changes

### Model Architecture
- **Fused preprocessor+encoder**: No separate Encoder.mlmodelc file
- **Smaller dimensions**: encoderHidden=512, vocabSize=1024, single LSTM
layer
- **Array-format vocabulary**: vocab.json instead of dict format
- **BlankId**: 1024 (same as v2)

### Code Modifications
- **AsrModels**: Optional encoder support, fused frontend loading, array
vocab handling
- **AsrManager**: Version-aware decoder state shapes, fused frontend
availability checking
- **AsrTranscription**: Skip encoder step when preprocessor output is
fused
- **TdtDecoderState**: Parameterized LSTM layer count
- **TdtDecoderV3**: Use config.encoderHiddenSize instead of
auto-detection
- **EncoderFrameView**: Accept explicit hidden size parameter
- **TranscribeCommand**: New `--model-version tdt-ctc-110m` and
`--model-dir` flags
- **ModelNames**: parakeetTdtCtc110m repo reference

### CLI Usage
```bash
swift run fluidaudiocli transcribe audio.wav --model-version tdt-ctc-110m
swift run fluidaudiocli transcribe audio.wav --model-version tdt-ctc-110m --model-dir /path/to/custom/models
```

## Testing
- [ ] iOS compatibility testing (per concerns in #383)
- [ ] Benchmark performance documentation
- [ ] Verify fused model behavior on both macOS and iOS

## Related
- Closes #383
- Model repo:
[FluidInference/parakeet-tdt-ctc-110m-coreml](https://huggingface.co/FluidInference/parakeet-tdt-ctc-110m-coreml)

<img width="642" height="1389" alt="IMG_5033"
src="https://github.com/user-attachments/assets/a9105cf7-552b-4573-acfb-2a089bf52820"
/><!-- devin-review-badge-begin -->

---

<a href="https://app.devin.ai/review/fluidinference/fluidaudio/pull/433"
target="_blank">
  <picture>
<source media="(prefers-color-scheme: dark)"
srcset="https://static.devin.ai/assets/gh-open-in-devin-review-dark.svg?v=1">
<img
src="https://static.devin.ai/assets/gh-open-in-devin-review-light.svg?v=1"
alt="Open with Devin">
  </picture>
</a>
<!-- devin-review-badge-end -->

---------

Co-authored-by: miro <jarbasai@mailfence.com>
2026-03-26 15:21:01 -04:00

3.0 KiB

Parakeet TDT-CTC-110M Benchmark Results

LibriSpeech test-clean (Full Dataset)

Metric Value
Files processed 2,620
Average WER 3.01%
Median WER 0.0%
Average CER 1.09%
Audio duration 19,452.5s (~5.4 hours)
Processing time 201.5s (~3.4 minutes)
Overall RTFx 96.5x
Median RTFx 86.4x

Configuration

  • Model: Parakeet TDT-CTC-110M (CoreML)
  • Architecture: Hybrid TDT-CTC with fused preprocessor+encoder
  • Platform: Apple Silicon (M2)
  • Date: March 26, 2026

Key Features

  • 96.5x real-time factor - 1 hour of audio transcribes in 37 seconds
  • 3.01% WER - Competitive accuracy on LibriSpeech test-clean
  • 0% median WER - Most files transcribed perfectly
  • iOS compatible - Runs on iPhone with full CoreML optimization
  • Stateless processing - No encoder state carryover needed

Running the Benchmark

# Build release
swift build -c release

# Run full benchmark (auto-downloads dataset and models)
.build/release/fluidaudiocli asr-benchmark --subset test-clean --model-version tdt-ctc-110m

# Run with limited files
.build/release/fluidaudiocli asr-benchmark --subset test-clean --model-version tdt-ctc-110m --max-files 100

# Process single file
.build/release/fluidaudiocli asr-benchmark --single-file 1089-134686-0000 --model-version tdt-ctc-110m

Notes

  • TDT (Token-and-Duration Transducer) decoder with CTC-constrained beam search
  • Fused preprocessor+encoder reduces model load time and memory usage
  • Models available at: FluidInference/parakeet-tdt-ctc-110m-coreml
  • iOS test app validates on-device performance with LibriSpeech ground truth

Nemotron Speech Streaming 0.6B Benchmark Results

LibriSpeech test-clean (Full Dataset)

Metric Value
Files processed 2,620
Total words 53,120
Total errors 1,334
WER 2.51%
Audio duration 19,452.5s (~5.4 hours)
Processing time 3,393.7s (~56.6 minutes)
RTFx 5.7x
Peak memory 1.452 GB

Configuration

  • Model: Nemotron Speech Streaming 0.6B (CoreML)
  • Encoder variant: int8
  • Platform: Apple Silicon (M4 Pro)
  • Date: January 15, 2026

Running the Benchmark

# Build release
swift build -c release

# Run full benchmark (auto-downloads dataset and models)
.build/release/fluidaudiocli nemotron-benchmark --subset test-clean

# Run with limited files
.build/release/fluidaudiocli nemotron-benchmark --subset test-clean --max-files 100

# Use float32 encoder variant
.build/release/fluidaudiocli nemotron-benchmark --encoder float32 --max-files 50

Notes