Commit Graph
2 Commits
Author SHA1 Message Date
Alexandmiro 0f7493bdac feat: Support Parakeet-TDT-CTC-110M hybrid model (#433)
## Summary
Adds support for NVIDIA's Parakeet-TDT-CTC-110M hybrid model with fused
preprocessor+encoder architecture.

Based on the work by @JarbasAl in #383.

## Key Changes

### Model Architecture
- **Fused preprocessor+encoder**: No separate Encoder.mlmodelc file
- **Smaller dimensions**: encoderHidden=512, vocabSize=1024, single LSTM
layer
- **Array-format vocabulary**: vocab.json instead of dict format
- **BlankId**: 1024 (same as v2)

### Code Modifications
- **AsrModels**: Optional encoder support, fused frontend loading, array
vocab handling
- **AsrManager**: Version-aware decoder state shapes, fused frontend
availability checking
- **AsrTranscription**: Skip encoder step when preprocessor output is
fused
- **TdtDecoderState**: Parameterized LSTM layer count
- **TdtDecoderV3**: Use config.encoderHiddenSize instead of
auto-detection
- **EncoderFrameView**: Accept explicit hidden size parameter
- **TranscribeCommand**: New `--model-version tdt-ctc-110m` and
`--model-dir` flags
- **ModelNames**: parakeetTdtCtc110m repo reference

### CLI Usage
```bash
swift run fluidaudiocli transcribe audio.wav --model-version tdt-ctc-110m
swift run fluidaudiocli transcribe audio.wav --model-version tdt-ctc-110m --model-dir /path/to/custom/models
```

## Testing
- [ ] iOS compatibility testing (per concerns in #383)
- [ ] Benchmark performance documentation
- [ ] Verify fused model behavior on both macOS and iOS

## Related
- Closes #383
- Model repo:
[FluidInference/parakeet-tdt-ctc-110m-coreml](https://huggingface.co/FluidInference/parakeet-tdt-ctc-110m-coreml)

<img width="642" height="1389" alt="IMG_5033"
src="https://github.com/user-attachments/assets/a9105cf7-552b-4573-acfb-2a089bf52820"
/><!-- devin-review-badge-begin -->

---

<a href="https://app.devin.ai/review/fluidinference/fluidaudio/pull/433"
target="_blank">
  <picture>
<source media="(prefers-color-scheme: dark)"
srcset="https://static.devin.ai/assets/gh-open-in-devin-review-dark.svg?v=1">
<img
src="https://static.devin.ai/assets/gh-open-in-devin-review-light.svg?v=1"
alt="Open with Devin">
  </picture>
</a>
<!-- devin-review-badge-end -->

---------

Co-authored-by: miro <jarbasai@mailfence.com>
2026-03-26 15:21:01 -04:00
Alex 88527fc329 feat(nemotron): add Nemotron Speech Streaming 0.6B with vDSP optimization (#432)
## Summary

Add streaming ASR support for NVIDIA's Nemotron Speech Streaming 0.6B
model converted to CoreML, with Accelerate framework optimization.

This PR addresses issue #389 by implementing
`NemotronStreamingAsrManager` for RNNT streaming inference.

**Key features:**
- True streaming with 560ms chunks and encoder cache
- Support for multiple chunk sizes: 80ms, 160ms, 560ms, 1120ms
- Int8 quantized encoder (default, 4x smaller than float32)
- **vDSP_maxvi optimization** for argmax operation (3.2% RTFx
improvement)
- CLI command `nemotron-benchmark` for LibriSpeech evaluation

## Performance

Benchmark on LibriSpeech test-clean (100 files, Apple M2):

| Metric | Value |
|--------|-------|
| **WER** | 2.12% |
| **RTFx** | 6.4x (real-time factor) |
| **Processing Time** | 141.3s (for 901.1s audio) |
| **Peak Memory** | 4.4 GB |

### Optimization Impact

Applied vDSP_maxvi from Accelerate framework for argmax operation:
- **2.2% faster** processing (144.5s → 141.3s)
- **3.2% RTFx improvement** (6.2x → 6.4x)
- Micro-benchmark shows 590x speedup for argmax itself
- See benchmark analysis: `/tmp/nemotron_benchmark_results.md`

## Implementation Details

**Architecture:**
1. **Preprocessor** — audio `[1, N]` → mel spectrogram `[1, 128, 56]`
2. **Encoder** (int8, with cache) — mel + cache → encoded features + new
cache
3. **Decoder + Joint** — RNNT greedy decode with vDSP-optimized argmax
4. **Tokenizer** — 1024-token vocab

**Model variants:**
- `nemotronStreaming80` — 80ms chunks (lowest latency)
- `nemotronStreaming160` — 160ms chunks
- `nemotronStreaming560` — 560ms chunks (default, best accuracy)
- `nemotronStreaming1120` — 1120ms chunks (highest throughput)

## Resolves

Closes #389

## Test Plan

- [x] Run `nemotron-benchmark --max-files 100` on LibriSpeech test-clean
- [x] Verify vDSP optimization maintains accuracy (WER unchanged)
- [x] Benchmark baseline vs optimized (2.2% speedup confirmed)
- [x] Test multi-variant support (80ms, 160ms, 560ms, 1120ms)
- [ ] Full LibriSpeech test-clean (2620 files) - optional

## Usage

```bash
# Run benchmark (default: 560ms variant, int8 encoder)
fluidaudiocli nemotron-benchmark --max-files 100

# Test different chunk sizes
fluidaudiocli nemotron-benchmark --chunk-size 160ms --max-files 10
fluidaudiocli nemotron-benchmark --chunk-size 1120ms --max-files 10
```

## Credits

- Original implementation: @Alex-Wengg
- vDSP optimization inspired by [Muesli
app](https://github.com/pHequals7/muesli) (@pHequals7)
- Issue reported by: @pHequals7 (#389)

🤖 Generated with [Claude Code](https://claude.com/claude-code)
<!-- devin-review-badge-begin -->

---

<a href="https://app.devin.ai/review/fluidinference/fluidaudio/pull/432"
target="_blank">
  <picture>
<source media="(prefers-color-scheme: dark)"
srcset="https://static.devin.ai/assets/gh-open-in-devin-review-dark.svg?v=1">
<img
src="https://static.devin.ai/assets/gh-open-in-devin-review-light.svg?v=1"
alt="Open with Devin">
  </picture>
</a>
<!-- devin-review-badge-end -->
2026-03-26 09:59:09 -04:00