mirror of
https://github.com/FluidInference/FluidAudio.git
synced 2026-06-11 20:24:36 +00:00
docs/update-documentation
11
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
1dfe1dbd37 |
Refine model descriptions in Models.md (#496)
Updated descriptions for various models, clarifying features and performance metrics. Enhanced details for TDT, streaming, custom vocabulary, VAD, diarization, and TTS models. ### Why is this change needed? <!-- Explain the motivation for this change. What problem does it solve? --> |
||
|
|
2593f55415 |
Add Japanese ASR support with JSUT and Common Voice datasets (#478)
## Summary Adds comprehensive Japanese ASR support to FluidAudio with benchmark datasets and CLI commands. ## Changes ### Core Japanese ASR Support - **CtcJaManager.swift** - Japanese CTC transcription manager (actor-based) - **CtcJaModels.swift** - Japanese model loading and management - **ModelNames.swift** - Added Japanese model registry (`parakeetCtcJa`, `CTCJa` enum) - **AsrModels.swift** - Added `.ctcJa` model version (3,072 vocab, 1,024 hidden, blank_id=3072) - **AsrManager.swift** - Added `.ctcJa` case with error directing to `CtcJaManager` ### CLI Commands - **JapaneseAsrBenchmark.swift** (459 lines) - New `ja-benchmark` command - JSUT basic5000 dataset support - Mozilla Common Voice (MCV) test set support - Auto-download capability - CER (Character Error Rate) evaluation - **DownloadCommand.swift** - Added JSUT and MCV Japanese dataset downloads - **TranscribeCommand.swift** - Added `.ctcJa` model version support - **AsrBenchmark.swift** - Added `.ctcJa` switch case ### Dataset Support - **JapaneseDatasetDownloader.swift** (387 lines) - Dataset download and parsing - JSUT basic5000 (5,000 sentences, clean studio recordings) - Mozilla Common Voice Japanese test split - Efficient streaming downloads - Metadata extraction and validation ## Usage ### CLI Commands ```bash # Benchmark on JSUT basic5000 (100 samples) swift run fluidaudiocli ja-benchmark --dataset jsut --samples 100 # Benchmark on Common Voice test (500 samples, auto-download) swift run fluidaudiocli ja-benchmark --dataset cv-test --samples 500 --auto-download # Download datasets swift run fluidaudiocli download --dataset jsut swift run fluidaudiocli download --dataset cv-ja-test ``` ### Swift API ```swift // Load and use Japanese CTC transcription let manager = try await CtcJaManager.load() let text = try manager.transcribe(audioURL: japaneseAudioFile) ``` ## Model Info - **Repo**: `FluidInference/parakeet-ctc-0.6b-ja-coreml` - **Architecture**: 600M parameter CTC-only - **Vocabulary**: 3,072 Japanese SentencePiece tokens + 1 blank (id: 3072) - **Encoder**: 1,024 hidden size - **Expected CER**: 6.5% on JSUT basic5000, 13.3% on MCV 16.1 test ## Testing - ✅ Builds successfully (`swift build`) - ✅ Model loading integration tested - ✅ CLI commands compile and link correctly - ⏳ Runtime benchmark testing pending (requires model download) ## Related - Mobius PR #39: Japanese CTC CoreML conversion (https://github.com/FluidInference/mobius/pull/39) 🤖 Generated with Claude Code <!-- devin-review-badge-begin --> --- <a href="https://app.devin.ai/review/fluidinference/fluidaudio/pull/478" target="_blank"> <picture> <source media="(prefers-color-scheme: dark)" srcset="https://static.devin.ai/assets/gh-open-in-devin-review-dark.svg?v=1"> <img src="https://static.devin.ai/assets/gh-open-in-devin-review-light.svg?v=1" alt="Open with Devin"> </picture> </a> <!-- devin-review-badge-end --> --------- |
||
|
|
e418cbca7d |
Mark KittenTTS and Qwen3-TTS as not supported (#437)
## Summary - Add KittenTTS to the "Evaluated Models (Not Supported)" section - Update section title from "Not Shipped" to "Not Supported" for clarity - Clarify these models are not maintained or recommended for use ## References - KittenTTS: #409 - Qwen3-TTS: #290 ## Changes - Updated `Documentation/Models.md` to list KittenTTS alongside Qwen3-TTS in the unsupported models section - Changed section heading to "Evaluated Models (Not Supported)" to be more explicit <!-- devin-review-badge-begin --> --- <a href="https://app.devin.ai/review/fluidinference/fluidaudio/pull/437" target="_blank"> <picture> <source media="(prefers-color-scheme: dark)" srcset="https://static.devin.ai/assets/gh-open-in-devin-review-dark.svg?v=1"> <img src="https://static.devin.ai/assets/gh-open-in-devin-review-light.svg?v=1" alt="Open with Devin"> </picture> </a> <!-- devin-review-badge-end --> |
||
|
|
0f7493bdac |
feat: Support Parakeet-TDT-CTC-110M hybrid model (#433)
## Summary Adds support for NVIDIA's Parakeet-TDT-CTC-110M hybrid model with fused preprocessor+encoder architecture. Based on the work by @JarbasAl in #383. ## Key Changes ### Model Architecture - **Fused preprocessor+encoder**: No separate Encoder.mlmodelc file - **Smaller dimensions**: encoderHidden=512, vocabSize=1024, single LSTM layer - **Array-format vocabulary**: vocab.json instead of dict format - **BlankId**: 1024 (same as v2) ### Code Modifications - **AsrModels**: Optional encoder support, fused frontend loading, array vocab handling - **AsrManager**: Version-aware decoder state shapes, fused frontend availability checking - **AsrTranscription**: Skip encoder step when preprocessor output is fused - **TdtDecoderState**: Parameterized LSTM layer count - **TdtDecoderV3**: Use config.encoderHiddenSize instead of auto-detection - **EncoderFrameView**: Accept explicit hidden size parameter - **TranscribeCommand**: New `--model-version tdt-ctc-110m` and `--model-dir` flags - **ModelNames**: parakeetTdtCtc110m repo reference ### CLI Usage ```bash swift run fluidaudiocli transcribe audio.wav --model-version tdt-ctc-110m swift run fluidaudiocli transcribe audio.wav --model-version tdt-ctc-110m --model-dir /path/to/custom/models ``` ## Testing - [ ] iOS compatibility testing (per concerns in #383) - [ ] Benchmark performance documentation - [ ] Verify fused model behavior on both macOS and iOS ## Related - Closes #383 - Model repo: [FluidInference/parakeet-tdt-ctc-110m-coreml](https://huggingface.co/FluidInference/parakeet-tdt-ctc-110m-coreml) <img width="642" height="1389" alt="IMG_5033" src="https://github.com/user-attachments/assets/a9105cf7-552b-4573-acfb-2a089bf52820" /><!-- devin-review-badge-begin --> --- <a href="https://app.devin.ai/review/fluidinference/fluidaudio/pull/433" target="_blank"> <picture> <source media="(prefers-color-scheme: dark)" srcset="https://static.devin.ai/assets/gh-open-in-devin-review-dark.svg?v=1"> <img src="https://static.devin.ai/assets/gh-open-in-devin-review-light.svg?v=1" alt="Open with Devin"> </picture> </a> <!-- devin-review-badge-end --> --------- Co-authored-by: miro <jarbasai@mailfence.com> |
||
|
|
88527fc329 |
feat(nemotron): add Nemotron Speech Streaming 0.6B with vDSP optimization (#432)
## Summary Add streaming ASR support for NVIDIA's Nemotron Speech Streaming 0.6B model converted to CoreML, with Accelerate framework optimization. This PR addresses issue #389 by implementing `NemotronStreamingAsrManager` for RNNT streaming inference. **Key features:** - True streaming with 560ms chunks and encoder cache - Support for multiple chunk sizes: 80ms, 160ms, 560ms, 1120ms - Int8 quantized encoder (default, 4x smaller than float32) - **vDSP_maxvi optimization** for argmax operation (3.2% RTFx improvement) - CLI command `nemotron-benchmark` for LibriSpeech evaluation ## Performance Benchmark on LibriSpeech test-clean (100 files, Apple M2): | Metric | Value | |--------|-------| | **WER** | 2.12% | | **RTFx** | 6.4x (real-time factor) | | **Processing Time** | 141.3s (for 901.1s audio) | | **Peak Memory** | 4.4 GB | ### Optimization Impact Applied vDSP_maxvi from Accelerate framework for argmax operation: - **2.2% faster** processing (144.5s → 141.3s) - **3.2% RTFx improvement** (6.2x → 6.4x) - Micro-benchmark shows 590x speedup for argmax itself - See benchmark analysis: `/tmp/nemotron_benchmark_results.md` ## Implementation Details **Architecture:** 1. **Preprocessor** — audio `[1, N]` → mel spectrogram `[1, 128, 56]` 2. **Encoder** (int8, with cache) — mel + cache → encoded features + new cache 3. **Decoder + Joint** — RNNT greedy decode with vDSP-optimized argmax 4. **Tokenizer** — 1024-token vocab **Model variants:** - `nemotronStreaming80` — 80ms chunks (lowest latency) - `nemotronStreaming160` — 160ms chunks - `nemotronStreaming560` — 560ms chunks (default, best accuracy) - `nemotronStreaming1120` — 1120ms chunks (highest throughput) ## Resolves Closes #389 ## Test Plan - [x] Run `nemotron-benchmark --max-files 100` on LibriSpeech test-clean - [x] Verify vDSP optimization maintains accuracy (WER unchanged) - [x] Benchmark baseline vs optimized (2.2% speedup confirmed) - [x] Test multi-variant support (80ms, 160ms, 560ms, 1120ms) - [ ] Full LibriSpeech test-clean (2620 files) - optional ## Usage ```bash # Run benchmark (default: 560ms variant, int8 encoder) fluidaudiocli nemotron-benchmark --max-files 100 # Test different chunk sizes fluidaudiocli nemotron-benchmark --chunk-size 160ms --max-files 10 fluidaudiocli nemotron-benchmark --chunk-size 1120ms --max-files 10 ``` ## Credits - Original implementation: @Alex-Wengg - vDSP optimization inspired by [Muesli app](https://github.com/pHequals7/muesli) (@pHequals7) - Issue reported by: @pHequals7 (#389) 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- devin-review-badge-begin --> --- <a href="https://app.devin.ai/review/fluidinference/fluidaudio/pull/432" target="_blank"> <picture> <source media="(prefers-color-scheme: dark)" srcset="https://static.devin.ai/assets/gh-open-in-devin-review-dark.svg?v=1"> <img src="https://static.devin.ai/assets/gh-open-in-devin-review-light.svg?v=1" alt="Open with Devin"> </picture> </a> <!-- devin-review-badge-end --> |
||
|
|
ba17ebc600 |
LS-EEND Diarizer (#376)
--- ## Add LS-EEND speaker diarization Sortformer handles up to 4 speakers and works best at 16 kHz in noisy environments. That leaves a gap for phone calls, large meetings, and recordings with unknown conditions. LS-EEND fills it: up to 10 speakers (variant-dependent), trained on telephone, meeting, and in-the-wild corpora, operating at 8 kHz. This PR adds LS-EEND as a first-class diarizer alongside Sortformer — same `Diarizer` protocol, same CLI patterns, same post-processing pipeline. ### Why these changes are needed **Unified timeline** — `SortformerTimeline` was Sortformer-specific and couldn't be shared. LS-EEND needs the same post-processing (threshold, median filter, onset/offset padding, min-duration filtering, finalized vs tentative segments). `DiarizerTimeline` replaces `SortformerTimeline` with a shared implementation that both models use, eliminating duplicated logic. **LS-EEND diarizer** — The model was partially wired up but missing a clean public API, proper `Diarizer` protocol conformance, and integration with `DiarizerTimeline`. This completes the implementation: offline file processing with automatic resampling, streaming with committed + speculative preview frames, and session-level control via `LSEENDStreamingSession`. **CLI** — Without `lseend` and `lseend-benchmark`, the model can't be used or evaluated outside of Swift code. The benchmark also validates that DER matches the paper's reported numbers before shipping to users. **AMI ground truth fallback** — `lseend-benchmark --variant ami` silently produced no results because the benchmark looked for RTTM files that don't exist in the standard dataset layout. Added the same `AMIParser` XML annotation fallback that the Sortformer benchmark uses. **Tests** — `LSEENDRuntimeTests` runs the inference engine, streaming session, and feature extractor against known-good outputs to catch regressions in the CoreML pipeline. **Documentation** — LS-EEND has a substantially different API surface than Sortformer (five source files, streaming session layer, matrix type, full evaluation namespace, per-variant speaker caps). Documents the entire public API and provides a variant selection guide. ### Changes **`DiarizerTimeline.swift`** (new) — Unified post-processing timeline shared by both Sortformer and LS-EEND. Replaces `SortformerTimeline.swift` (deleted). `SortformerDiarizerPipeline` updated to use it. **`LSEENDDiarizer.swift`** — `Diarizer` protocol conformance; offline (`processComplete(audioFileURL:)`) and streaming (`addAudio` / `process` / `finalizeSession`) APIs; thread-safe via `NSLock`. **`LSEENDInference.swift`** — `LSEENDInferenceEngine` (offline, streaming, simulation) and `LSEENDStreamingSession` (stateful, frame-in-frame-out with committed + preview outputs). **`LSEENDFeatureExtraction.swift`** — `LSEENDOfflineFeatureExtractor` and `LSEENDStreamingFeatureExtractor`; log-mel cumulative mean normalization and splice-and-subsample. **`LSEENDEvaluation.swift`** — DER computation with collar masking and optimal speaker assignment (Hungarian); RTTM parsing and writing. **`LSEENDCommand.swift`**, **`LSEENDBenchmark.swift`** — CLI commands `lseend` and `lseend-benchmark`, with the same post-processing flags as the Sortformer equivalents. **`LSEENDRuntimeTests.swift`** — Integration tests for offline inference, streaming, session behavior, and feature extraction. **`Documentation/Diarization/LSEEND.md`** — Full public API reference and variant selection guide (`.ami` → 4 speakers, `.callhome` → 7, `.dihard2`/`.dihard3` → 10; DER numbers from the paper). All tasks from the previous session are complete: 1. **Merge conflict** in `LSEENDRuntimeProbeSupport.swift` — resolved using the async approach, merged `claude/nice-brattain` into `ls-eend` 2. **RTTM not found bug** in `lseend-benchmark` — fixed with AMI XML annotation fallback, `public init` on `LSEENDRTTMEntry`, async `processMeeting` 3. **Documentation** — `Documentation/Diarization/LSEEND.md` with full public API reference, correct speaker counts (AMI→4, CALLHOME→7, DIHARD2/3→10) 4. **PR description** — written in chat covering the full `ls-eend` branch scope Everything is committed to the `ls-eend` branch at `/Users/benjaminlee/Documents/FluidAudio`. Let me know what you'd like to work on next. <!-- devin-review-badge-begin --> --- <a href="https://app.devin.ai/review/fluidinference/fluidaudio/pull/376" target="_blank"> <picture> <source media="(prefers-color-scheme: dark)" srcset="https://static.devin.ai/assets/gh-open-in-devin-review-dark.svg?v=1"> <img src="https://static.devin.ai/assets/gh-open-in-devin-review-light.svg?v=1" alt="Open with Devin"> </picture> </a> <!-- devin-review-badge-end --> --------- |
||
|
|
1afc00c0a6 |
docs: add Qwen3-TTS and Qwen3-ForcedAligner to evaluated models (#387)
## Summary - Add Qwen3-TTS and Qwen3-ForcedAligner-0.6B to the "Evaluated Models (Not Shipped)" section in Documentation/Models.md - Links to FluidAudio PRs, mobius conversion PRs, and HuggingFace repos ## Test plan - [ ] Verify markdown renders correctly <!-- devin-review-badge-begin --> --- <a href="https://app.devin.ai/review/fluidinference/fluidaudio/pull/387" target="_blank"> <picture> <source media="(prefers-color-scheme: dark)" srcset="https://static.devin.ai/assets/gh-open-in-devin-review-dark.svg?v=1"> <img src="https://static.devin.ai/assets/gh-open-in-devin-review-light.svg?v=1" alt="Open with Devin"> </picture> </a> <!-- devin-review-badge-end --> |
||
|
|
597dcebec0 |
docs: update Kokoro TTS docs — not deprecated, add known issues (#360)
## Summary - Remove language in Models.md that framed PocketTTS as an "upgrade" / "improvement over Kokoro" — they're two backends with different tradeoffs - Update Kokoro description to reflect the CoreML G2P model (no longer uses espeak) - Add Known Issues section to Kokoro.md documenting sibilance in high-pitched `af_*` voices (see [mobius#23](https://github.com/FluidInference/mobius/issues/23)); low-tone voices are unaffected ## Test plan - [ ] Documentation renders correctly on GitHub <!-- devin-review-badge-begin --> --- <a href="https://app.devin.ai/review/fluidinference/fluidaudio/pull/360" target="_blank"> <picture> <source media="(prefers-color-scheme: dark)" srcset="https://static.devin.ai/assets/gh-open-in-devin-review-dark.svg?v=1"> <img src="https://static.devin.ai/assets/gh-open-in-devin-review-light.svg?v=1" alt="Open with Devin"> </picture> </a> <!-- devin-review-badge-end --> |
||
|
|
050296998d |
docs: add Evaluated Models section for models tried but not shipped (#285)
## Summary Add documentation for models we converted and tested but haven't shipped yet — provides transparency on what we've explored. ### ASR Models Documented - **Nemotron Speech Streaming 0.6B** — In development, strong performance (1.99% WER) - **Qwen3-ASR-0.6B** — In development, encoder-decoder architecture - **Canary 1B v2** — Too large for mobile - **Hybrid CTC-TDT 110M** — Superseded by separate CTC models ### TTS Models Documented - **Kokoro Chinese (MLX)** — Pure Swift MLX inference, experimental - **Kitten TTS** — Superseded by PocketTTS ## Test plan - [x] Documentation renders correctly in GitHub |
||
|
|
9fcdf2f32c |
feat: add PocketTTS backend for lightweight text-to-speech (#273)
## Summary - Add PocketTTS as a new TTS backend — flow-matching language model with autoregressive streaming synthesis - Pure Swift implementation using 4 CoreML models (cond_step, flowlm_step, flow_decoder, mimi_decoder) - iOS 17 compatible — no `scaled_dot_product_attention` ops (avoids BNNS crash) - Add audio post-processor with de-esser for reducing sibilant harshness ## Test plan - [x] Short sentence: WER 0, 3.44s audio - [x] Long sentence: WER 0, 6.64s audio - [x] Fresh HuggingFace download works end-to-end - [x] iOS build succeeds (`xcodebuild -destination 'generic/platform=iOS'`) - [x] macOS build succeeds (`swift build -c release`) |
||
|
|
5d9176eb35 |
docs: organize Documentation folder structure and fix stale content (#280)
## Summary Closes #274 - **Reorganized docs into subdirectories**: moved diarization docs → `Diarization/`, custom vocabulary docs → `ASR/`, eSpeak docs → `TTS/` - **Added `Models.md`**: comprehensive guide to all CoreML model pipelines (ASR, VAD, Diarization, TTS) with architecture, performance, and source links - **Fixed stale content across 5 files**: removed nonexistent `compareSpeakers` from API.md, fixed year in Benchmarks.md, fixed `TtSManager` casing in SSML.md, updated cross-references in SpeakerManager.md, removed dead MCP link from README - **Rewrote README.md index** with correct paths for all reorganized docs - **Added `.gitignore` patterns** for benchmark artifact files (`*benchmark*.json`, `.sortformer_progress*.json`) ## Files changed | Change | File | |--------|------| | Moved | `SpeakerDiarization.md` → `Diarization/GettingStarted.md` | | Moved | `SpeakerManager.md` → `Diarization/SpeakerManager.md` | | Moved | `Sortformer.md` → `Diarization/Sortformer.md` | | Moved | `DIARIZATION_INVESTIGATION_REPORT.md` → `Diarization/InvestigationReport.md` | | Moved | `CtcCustomVocabulary.md` → `ASR/CustomVocabulary.md` | | Moved | `CustomPronunciationDictionary.md` → `ASR/CustomPronunciation.md` | | Moved | `EspeakFramework.md` → `TTS/EspeakFramework.md` | | New | `Models.md` — all CoreML model pipelines | | Updated | `README.md` — new index with correct paths | | Fixed | `API.md` — removed nonexistent `compareSpeakers` | | Fixed | `Benchmarks.md` — year 2024 → 2025 | | Fixed | `TTS/SSML.md` — `TtsManager` → `TtSManager` | | Fixed | `Diarization/SpeakerManager.md` — cross-references | | Updated | `.gitignore` — benchmark artifact patterns | ## Test plan - [x] No code changes — documentation only - [ ] Verify all internal doc links resolve correctly - [ ] Review Models.md content for accuracy |