Commit Graph
11 Commits
Author SHA1 Message Date
Alex 1dfe1dbd37 Refine model descriptions in Models.md (#496)
Updated descriptions for various models, clarifying features and
performance metrics. Enhanced details for TDT, streaming, custom
vocabulary, VAD, diarization, and TTS models.

### Why is this change needed?
<!-- Explain the motivation for this change. What problem does it solve?
-->
2026-04-07 19:38:09 -04:00
Alex 2593f55415 Add Japanese ASR support with JSUT and Common Voice datasets (#478)
## Summary

Adds comprehensive Japanese ASR support to FluidAudio with benchmark
datasets and CLI commands.

## Changes

### Core Japanese ASR Support
- **CtcJaManager.swift** - Japanese CTC transcription manager
(actor-based)
- **CtcJaModels.swift** - Japanese model loading and management
- **ModelNames.swift** - Added Japanese model registry (`parakeetCtcJa`,
`CTCJa` enum)
- **AsrModels.swift** - Added `.ctcJa` model version (3,072 vocab, 1,024
hidden, blank_id=3072)
- **AsrManager.swift** - Added `.ctcJa` case with error directing to
`CtcJaManager`

### CLI Commands
- **JapaneseAsrBenchmark.swift** (459 lines) - New `ja-benchmark`
command
  - JSUT basic5000 dataset support
  - Mozilla Common Voice (MCV) test set support
  - Auto-download capability
  - CER (Character Error Rate) evaluation
- **DownloadCommand.swift** - Added JSUT and MCV Japanese dataset
downloads
- **TranscribeCommand.swift** - Added `.ctcJa` model version support
- **AsrBenchmark.swift** - Added `.ctcJa` switch case

### Dataset Support
- **JapaneseDatasetDownloader.swift** (387 lines) - Dataset download and
parsing
  - JSUT basic5000 (5,000 sentences, clean studio recordings)
  - Mozilla Common Voice Japanese test split
  - Efficient streaming downloads
  - Metadata extraction and validation

## Usage

### CLI Commands
```bash
# Benchmark on JSUT basic5000 (100 samples)
swift run fluidaudiocli ja-benchmark --dataset jsut --samples 100

# Benchmark on Common Voice test (500 samples, auto-download)
swift run fluidaudiocli ja-benchmark --dataset cv-test --samples 500 --auto-download

# Download datasets
swift run fluidaudiocli download --dataset jsut
swift run fluidaudiocli download --dataset cv-ja-test
```

### Swift API
```swift
// Load and use Japanese CTC transcription
let manager = try await CtcJaManager.load()
let text = try manager.transcribe(audioURL: japaneseAudioFile)
```

## Model Info
- **Repo**: `FluidInference/parakeet-ctc-0.6b-ja-coreml`
- **Architecture**: 600M parameter CTC-only
- **Vocabulary**: 3,072 Japanese SentencePiece tokens + 1 blank (id:
3072)
- **Encoder**: 1,024 hidden size
- **Expected CER**: 6.5% on JSUT basic5000, 13.3% on MCV 16.1 test

## Testing
- ✅ Builds successfully (`swift build`)
- ✅ Model loading integration tested
- ✅ CLI commands compile and link correctly
- ⏳ Runtime benchmark testing pending (requires model download)

## Related
- Mobius PR #39: Japanese CTC CoreML conversion
(https://github.com/FluidInference/mobius/pull/39)

🤖 Generated with Claude Code
<!-- devin-review-badge-begin -->

---

<a href="https://app.devin.ai/review/fluidinference/fluidaudio/pull/478"
target="_blank">
  <picture>
<source media="(prefers-color-scheme: dark)"
srcset="https://static.devin.ai/assets/gh-open-in-devin-review-dark.svg?v=1">
<img
src="https://static.devin.ai/assets/gh-open-in-devin-review-light.svg?v=1"
alt="Open with Devin">
  </picture>
</a>
<!-- devin-review-badge-end -->

---------
2026-04-04 12:57:32 -04:00
Alex e418cbca7d Mark KittenTTS and Qwen3-TTS as not supported (#437)
## Summary

- Add KittenTTS to the "Evaluated Models (Not Supported)" section
- Update section title from "Not Shipped" to "Not Supported" for clarity
- Clarify these models are not maintained or recommended for use

## References

- KittenTTS: #409
- Qwen3-TTS: #290

## Changes

- Updated `Documentation/Models.md` to list KittenTTS alongside
Qwen3-TTS in the unsupported models section
- Changed section heading to "Evaluated Models (Not Supported)" to be
more explicit
<!-- devin-review-badge-begin -->

---

<a href="https://app.devin.ai/review/fluidinference/fluidaudio/pull/437"
target="_blank">
  <picture>
<source media="(prefers-color-scheme: dark)"
srcset="https://static.devin.ai/assets/gh-open-in-devin-review-dark.svg?v=1">
<img
src="https://static.devin.ai/assets/gh-open-in-devin-review-light.svg?v=1"
alt="Open with Devin">
  </picture>
</a>
<!-- devin-review-badge-end -->
2026-03-26 17:52:10 -04:00
Alexandmiro 0f7493bdac feat: Support Parakeet-TDT-CTC-110M hybrid model (#433)
## Summary
Adds support for NVIDIA's Parakeet-TDT-CTC-110M hybrid model with fused
preprocessor+encoder architecture.

Based on the work by @JarbasAl in #383.

## Key Changes

### Model Architecture
- **Fused preprocessor+encoder**: No separate Encoder.mlmodelc file
- **Smaller dimensions**: encoderHidden=512, vocabSize=1024, single LSTM
layer
- **Array-format vocabulary**: vocab.json instead of dict format
- **BlankId**: 1024 (same as v2)

### Code Modifications
- **AsrModels**: Optional encoder support, fused frontend loading, array
vocab handling
- **AsrManager**: Version-aware decoder state shapes, fused frontend
availability checking
- **AsrTranscription**: Skip encoder step when preprocessor output is
fused
- **TdtDecoderState**: Parameterized LSTM layer count
- **TdtDecoderV3**: Use config.encoderHiddenSize instead of
auto-detection
- **EncoderFrameView**: Accept explicit hidden size parameter
- **TranscribeCommand**: New `--model-version tdt-ctc-110m` and
`--model-dir` flags
- **ModelNames**: parakeetTdtCtc110m repo reference

### CLI Usage
```bash
swift run fluidaudiocli transcribe audio.wav --model-version tdt-ctc-110m
swift run fluidaudiocli transcribe audio.wav --model-version tdt-ctc-110m --model-dir /path/to/custom/models
```

## Testing
- [ ] iOS compatibility testing (per concerns in #383)
- [ ] Benchmark performance documentation
- [ ] Verify fused model behavior on both macOS and iOS

## Related
- Closes #383
- Model repo:
[FluidInference/parakeet-tdt-ctc-110m-coreml](https://huggingface.co/FluidInference/parakeet-tdt-ctc-110m-coreml)

<img width="642" height="1389" alt="IMG_5033"
src="https://github.com/user-attachments/assets/a9105cf7-552b-4573-acfb-2a089bf52820"
/><!-- devin-review-badge-begin -->

---

<a href="https://app.devin.ai/review/fluidinference/fluidaudio/pull/433"
target="_blank">
  <picture>
<source media="(prefers-color-scheme: dark)"
srcset="https://static.devin.ai/assets/gh-open-in-devin-review-dark.svg?v=1">
<img
src="https://static.devin.ai/assets/gh-open-in-devin-review-light.svg?v=1"
alt="Open with Devin">
  </picture>
</a>
<!-- devin-review-badge-end -->

---------

Co-authored-by: miro <jarbasai@mailfence.com>
2026-03-26 15:21:01 -04:00
Alex 88527fc329 feat(nemotron): add Nemotron Speech Streaming 0.6B with vDSP optimization (#432)
## Summary

Add streaming ASR support for NVIDIA's Nemotron Speech Streaming 0.6B
model converted to CoreML, with Accelerate framework optimization.

This PR addresses issue #389 by implementing
`NemotronStreamingAsrManager` for RNNT streaming inference.

**Key features:**
- True streaming with 560ms chunks and encoder cache
- Support for multiple chunk sizes: 80ms, 160ms, 560ms, 1120ms
- Int8 quantized encoder (default, 4x smaller than float32)
- **vDSP_maxvi optimization** for argmax operation (3.2% RTFx
improvement)
- CLI command `nemotron-benchmark` for LibriSpeech evaluation

## Performance

Benchmark on LibriSpeech test-clean (100 files, Apple M2):

| Metric | Value |
|--------|-------|
| **WER** | 2.12% |
| **RTFx** | 6.4x (real-time factor) |
| **Processing Time** | 141.3s (for 901.1s audio) |
| **Peak Memory** | 4.4 GB |

### Optimization Impact

Applied vDSP_maxvi from Accelerate framework for argmax operation:
- **2.2% faster** processing (144.5s → 141.3s)
- **3.2% RTFx improvement** (6.2x → 6.4x)
- Micro-benchmark shows 590x speedup for argmax itself
- See benchmark analysis: `/tmp/nemotron_benchmark_results.md`

## Implementation Details

**Architecture:**
1. **Preprocessor** — audio `[1, N]` → mel spectrogram `[1, 128, 56]`
2. **Encoder** (int8, with cache) — mel + cache → encoded features + new
cache
3. **Decoder + Joint** — RNNT greedy decode with vDSP-optimized argmax
4. **Tokenizer** — 1024-token vocab

**Model variants:**
- `nemotronStreaming80` — 80ms chunks (lowest latency)
- `nemotronStreaming160` — 160ms chunks
- `nemotronStreaming560` — 560ms chunks (default, best accuracy)
- `nemotronStreaming1120` — 1120ms chunks (highest throughput)

## Resolves

Closes #389

## Test Plan

- [x] Run `nemotron-benchmark --max-files 100` on LibriSpeech test-clean
- [x] Verify vDSP optimization maintains accuracy (WER unchanged)
- [x] Benchmark baseline vs optimized (2.2% speedup confirmed)
- [x] Test multi-variant support (80ms, 160ms, 560ms, 1120ms)
- [ ] Full LibriSpeech test-clean (2620 files) - optional

## Usage

```bash
# Run benchmark (default: 560ms variant, int8 encoder)
fluidaudiocli nemotron-benchmark --max-files 100

# Test different chunk sizes
fluidaudiocli nemotron-benchmark --chunk-size 160ms --max-files 10
fluidaudiocli nemotron-benchmark --chunk-size 1120ms --max-files 10
```

## Credits

- Original implementation: @Alex-Wengg
- vDSP optimization inspired by [Muesli
app](https://github.com/pHequals7/muesli) (@pHequals7)
- Issue reported by: @pHequals7 (#389)

🤖 Generated with [Claude Code](https://claude.com/claude-code)
<!-- devin-review-badge-begin -->

---

<a href="https://app.devin.ai/review/fluidinference/fluidaudio/pull/432"
target="_blank">
  <picture>
<source media="(prefers-color-scheme: dark)"
srcset="https://static.devin.ai/assets/gh-open-in-devin-review-dark.svg?v=1">
<img
src="https://static.devin.ai/assets/gh-open-in-devin-review-light.svg?v=1"
alt="Open with Devin">
  </picture>
</a>
<!-- devin-review-badge-end -->
2026-03-26 09:59:09 -04:00
Benjamin Lee ba17ebc600 LS-EEND Diarizer (#376)
---

## Add LS-EEND speaker diarization

Sortformer handles up to 4 speakers and works best at 16 kHz in noisy
environments. That leaves a gap for phone calls, large meetings, and
recordings with unknown conditions. LS-EEND fills it: up to 10 speakers
(variant-dependent), trained on telephone, meeting, and in-the-wild
corpora, operating at 8 kHz.

This PR adds LS-EEND as a first-class diarizer alongside Sortformer —
same `Diarizer` protocol, same CLI patterns, same post-processing
pipeline.

### Why these changes are needed

**Unified timeline** — `SortformerTimeline` was Sortformer-specific and
couldn't be shared. LS-EEND needs the same post-processing (threshold,
median filter, onset/offset padding, min-duration filtering, finalized
vs tentative segments). `DiarizerTimeline` replaces `SortformerTimeline`
with a shared implementation that both models use, eliminating
duplicated logic.

**LS-EEND diarizer** — The model was partially wired up but missing a
clean public API, proper `Diarizer` protocol conformance, and
integration with `DiarizerTimeline`. This completes the implementation:
offline file processing with automatic resampling, streaming with
committed + speculative preview frames, and session-level control via
`LSEENDStreamingSession`.

**CLI** — Without `lseend` and `lseend-benchmark`, the model can't be
used or evaluated outside of Swift code. The benchmark also validates
that DER matches the paper's reported numbers before shipping to users.

**AMI ground truth fallback** — `lseend-benchmark --variant ami`
silently produced no results because the benchmark looked for RTTM files
that don't exist in the standard dataset layout. Added the same
`AMIParser` XML annotation fallback that the Sortformer benchmark uses.

**Tests** — `LSEENDRuntimeTests` runs the inference engine, streaming
session, and feature extractor against known-good outputs to catch
regressions in the CoreML pipeline.

**Documentation** — LS-EEND has a substantially different API surface
than Sortformer (five source files, streaming session layer, matrix
type, full evaluation namespace, per-variant speaker caps). Documents
the entire public API and provides a variant selection guide.

### Changes

**`DiarizerTimeline.swift`** (new) — Unified post-processing timeline
shared by both Sortformer and LS-EEND. Replaces
`SortformerTimeline.swift` (deleted). `SortformerDiarizerPipeline`
updated to use it.

**`LSEENDDiarizer.swift`** — `Diarizer` protocol conformance; offline
(`processComplete(audioFileURL:)`) and streaming (`addAudio` / `process`
/ `finalizeSession`) APIs; thread-safe via `NSLock`.

**`LSEENDInference.swift`** — `LSEENDInferenceEngine` (offline,
streaming, simulation) and `LSEENDStreamingSession` (stateful,
frame-in-frame-out with committed + preview outputs).

**`LSEENDFeatureExtraction.swift`** — `LSEENDOfflineFeatureExtractor`
and `LSEENDStreamingFeatureExtractor`; log-mel cumulative mean
normalization and splice-and-subsample.

**`LSEENDEvaluation.swift`** — DER computation with collar masking and
optimal speaker assignment (Hungarian); RTTM parsing and writing.

**`LSEENDCommand.swift`**, **`LSEENDBenchmark.swift`** — CLI commands
`lseend` and `lseend-benchmark`, with the same post-processing flags as
the Sortformer equivalents.

**`LSEENDRuntimeTests.swift`** — Integration tests for offline
inference, streaming, session behavior, and feature extraction.

**`Documentation/Diarization/LSEEND.md`** — Full public API reference
and variant selection guide (`.ami` → 4 speakers, `.callhome` → 7,
`.dihard2`/`.dihard3` → 10; DER numbers from the paper).

All tasks from the previous session are complete:

1. **Merge conflict** in `LSEENDRuntimeProbeSupport.swift` — resolved
using the async approach, merged `claude/nice-brattain` into `ls-eend`
2. **RTTM not found bug** in `lseend-benchmark` — fixed with AMI XML
annotation fallback, `public init` on `LSEENDRTTMEntry`, async
`processMeeting`
3. **Documentation** — `Documentation/Diarization/LSEEND.md` with full
public API reference, correct speaker counts (AMI→4, CALLHOME→7,
DIHARD2/3→10)
4. **PR description** — written in chat covering the full `ls-eend`
branch scope

Everything is committed to the `ls-eend` branch at
`/Users/benjaminlee/Documents/FluidAudio`. Let me know what you'd like
to work on next.
<!-- devin-review-badge-begin -->

---

<a href="https://app.devin.ai/review/fluidinference/fluidaudio/pull/376"
target="_blank">
  <picture>
<source media="(prefers-color-scheme: dark)"
srcset="https://static.devin.ai/assets/gh-open-in-devin-review-dark.svg?v=1">
<img
src="https://static.devin.ai/assets/gh-open-in-devin-review-light.svg?v=1"
alt="Open with Devin">
  </picture>
</a>
<!-- devin-review-badge-end -->

---------
2026-03-17 18:03:54 -04:00
Alex 1afc00c0a6 docs: add Qwen3-TTS and Qwen3-ForcedAligner to evaluated models (#387)
## Summary

- Add Qwen3-TTS and Qwen3-ForcedAligner-0.6B to the "Evaluated Models
(Not Shipped)" section in Documentation/Models.md
- Links to FluidAudio PRs, mobius conversion PRs, and HuggingFace repos

## Test plan

- [ ] Verify markdown renders correctly
<!-- devin-review-badge-begin -->

---

<a href="https://app.devin.ai/review/fluidinference/fluidaudio/pull/387"
target="_blank">
  <picture>
<source media="(prefers-color-scheme: dark)"
srcset="https://static.devin.ai/assets/gh-open-in-devin-review-dark.svg?v=1">
<img
src="https://static.devin.ai/assets/gh-open-in-devin-review-light.svg?v=1"
alt="Open with Devin">
  </picture>
</a>
<!-- devin-review-badge-end -->
2026-03-16 15:29:50 -04:00
Alex 597dcebec0 docs: update Kokoro TTS docs — not deprecated, add known issues (#360)
## Summary
- Remove language in Models.md that framed PocketTTS as an "upgrade" /
"improvement over Kokoro" — they're two backends with different
tradeoffs
- Update Kokoro description to reflect the CoreML G2P model (no longer
uses espeak)
- Add Known Issues section to Kokoro.md documenting sibilance in
high-pitched `af_*` voices (see
[mobius#23](https://github.com/FluidInference/mobius/issues/23));
low-tone voices are unaffected

## Test plan
- [ ] Documentation renders correctly on GitHub
<!-- devin-review-badge-begin -->

---

<a href="https://app.devin.ai/review/fluidinference/fluidaudio/pull/360"
target="_blank">
  <picture>
<source media="(prefers-color-scheme: dark)"
srcset="https://static.devin.ai/assets/gh-open-in-devin-review-dark.svg?v=1">
<img
src="https://static.devin.ai/assets/gh-open-in-devin-review-light.svg?v=1"
alt="Open with Devin">
  </picture>
</a>
<!-- devin-review-badge-end -->
2026-03-11 12:46:58 -04:00
Alex 050296998d docs: add Evaluated Models section for models tried but not shipped (#285)
## Summary

Add documentation for models we converted and tested but haven't shipped
yet — provides transparency on what we've explored.

### ASR Models Documented
- **Nemotron Speech Streaming 0.6B** — In development, strong
performance (1.99% WER)
- **Qwen3-ASR-0.6B** — In development, encoder-decoder architecture
- **Canary 1B v2** — Too large for mobile
- **Hybrid CTC-TDT 110M** — Superseded by separate CTC models

### TTS Models Documented
- **Kokoro Chinese (MLX)** — Pure Swift MLX inference, experimental
- **Kitten TTS** — Superseded by PocketTTS

## Test plan
- [x] Documentation renders correctly in GitHub
2026-02-03 00:46:58 -05:00
Alex 9fcdf2f32c feat: add PocketTTS backend for lightweight text-to-speech (#273)
## Summary
- Add PocketTTS as a new TTS backend — flow-matching language model with
autoregressive streaming synthesis
- Pure Swift implementation using 4 CoreML models (cond_step,
flowlm_step, flow_decoder, mimi_decoder)
- iOS 17 compatible — no `scaled_dot_product_attention` ops (avoids BNNS
crash)
- Add audio post-processor with de-esser for reducing sibilant harshness

## Test plan
- [x] Short sentence: WER 0, 3.44s audio
- [x] Long sentence: WER 0, 6.64s audio
- [x] Fresh HuggingFace download works end-to-end
- [x] iOS build succeeds (`xcodebuild -destination
'generic/platform=iOS'`)
- [x] macOS build succeeds (`swift build -c release`)
2026-02-02 23:49:47 -05:00
Alex 5d9176eb35 docs: organize Documentation folder structure and fix stale content (#280)
## Summary

Closes #274

- **Reorganized docs into subdirectories**: moved diarization docs →
`Diarization/`, custom vocabulary docs → `ASR/`, eSpeak docs → `TTS/`
- **Added `Models.md`**: comprehensive guide to all CoreML model
pipelines (ASR, VAD, Diarization, TTS) with architecture, performance,
and source links
- **Fixed stale content across 5 files**: removed nonexistent
`compareSpeakers` from API.md, fixed year in Benchmarks.md, fixed
`TtSManager` casing in SSML.md, updated cross-references in
SpeakerManager.md, removed dead MCP link from README
- **Rewrote README.md index** with correct paths for all reorganized
docs
- **Added `.gitignore` patterns** for benchmark artifact files
(`*benchmark*.json`, `.sortformer_progress*.json`)

## Files changed

| Change | File |
|--------|------|
| Moved | `SpeakerDiarization.md` → `Diarization/GettingStarted.md` |
| Moved | `SpeakerManager.md` → `Diarization/SpeakerManager.md` |
| Moved | `Sortformer.md` → `Diarization/Sortformer.md` |
| Moved | `DIARIZATION_INVESTIGATION_REPORT.md` →
`Diarization/InvestigationReport.md` |
| Moved | `CtcCustomVocabulary.md` → `ASR/CustomVocabulary.md` |
| Moved | `CustomPronunciationDictionary.md` →
`ASR/CustomPronunciation.md` |
| Moved | `EspeakFramework.md` → `TTS/EspeakFramework.md` |
| New | `Models.md` — all CoreML model pipelines |
| Updated | `README.md` — new index with correct paths |
| Fixed | `API.md` — removed nonexistent `compareSpeakers` |
| Fixed | `Benchmarks.md` — year 2024 → 2025 |
| Fixed | `TTS/SSML.md` — `TtsManager` → `TtSManager` |
| Fixed | `Diarization/SpeakerManager.md` — cross-references |
| Updated | `.gitignore` — benchmark artifact patterns |

## Test plan

- [x] No code changes — documentation only
- [ ] Verify all internal doc links resolve correctly
- [ ] Review Models.md content for accuracy
2026-01-30 13:42:33 -05:00