Commit Graph
16 Commits
Author SHA1 Message Date
Alex 92755e0e01 feat: add multilingual G2P model and benchmark CLI command (#367)
## Summary
- Add CharsiuG2P ByT5 CoreML multilingual G2P model
(`MultilingualG2PModel`, `MultilingualG2PLanguage`,
`MultilingualG2PError`) supporting 9 Kokoro-mapped languages
- Add `g2p-benchmark` CLI command measuring PER/WER/speed against
CharsiuG2P test set with JSON output
- Switch both English and multilingual G2P models to `cpuOnly` compute
units (benchmarked 2-3x faster than GPU/ANE for autoregressive decoding)
- Add `LevenshteinDistance` utility and `MultilingualG2PTests` (9 tests)

### Benchmark Results (M2, CPU-only, 500 words/language)

| Language | PER | WER | ms/word |
|---|---|---|---|
| Spanish | 0.1% | 0.8% | 32.6 |
| French | 0.8% | 2.0% | 26.5 |
| Italian | 2.8% | 20.0% | 20.9 |
| Hindi | 4.5% | 21.4% | 45.4 |
| Japanese | 10.5% | 23.8% | 31.7 |
| Portuguese | 8.9% | 43.2% | 24.0 |
| British English | 13.6% | 29.4% | 34.0 |
| American English | 19.0% | 38.8% | 28.2 |
| Chinese | 86.2% | 95.0% | 53.9 |

### Compute Unit Benchmarks (English BART G2P)

| Config | ms/word |
|---|---|
| cpuOnly | **13.0** |
| all (ANE+GPU+CPU) | 17.3 |
| cpuAndGPU | 23.4 |

## Test plan
- [ ] `swift build` compiles clean
- [ ] `swift test --filter MultilingualG2PTests` passes (9 tests)
- [ ] `fluidaudiocli g2p-benchmark --languages eng-us --max-words 10
--data-dir <path>` produces results
- [ ] Verify JSON output file is written correctly
<!-- devin-review-badge-begin -->

---

<a href="https://app.devin.ai/review/fluidinference/fluidaudio/pull/367"
target="_blank">
  <picture>
<source media="(prefers-color-scheme: dark)"
srcset="https://static.devin.ai/assets/gh-open-in-devin-review-dark.svg?v=1">
<img
src="https://static.devin.ai/assets/gh-open-in-devin-review-light.svg?v=1"
alt="Open with Devin">
  </picture>
</a>
<!-- devin-review-badge-end -->
2026-03-13 14:35:53 -04:00
Sachin DesaiandAlex-Wengg 8d3ce44ae1 feat: Custom vocabulary support (#251)
### Why is this change needed?
Automatic speech recognition (ASR) systems are trained on massive
datasets of general speech, which means they excel at common vocabulary
but struggle with domain-specific terminology. In business
contexts—earnings calls, medical dictation, legal proceedings—the most
critical words are often the ones the model has rarely or never seen:
company names like "Saoirse Ronan," product names like "Newrez," or
technical jargon unique to an industry. Without intervention, these
high-value terms get transcribed as phonetically similar but incorrect
common words, turning "Nequi" into "NECI" or "Bose" into "Boz."

Keyword boosting (also called context biasing or vocabulary rescoring)
addresses this gap by incorporating domain knowledge at inference time.
The system is given a list of expected vocabulary terms and uses
acoustic evidence—typically CTC log-probabilities—to determine whether
the audio actually supports replacing a transcribed word with a
vocabulary term. This isn't blind substitution; a well-designed rescorer
computes scores for both the original transcription and the candidate
vocabulary term, only making replacements when the acoustic evidence
favors the domain term. The "context-biasing weight" parameter allows
tuning how aggressively to prefer vocabulary terms.

The real-world impact is substantial. In our earnings call benchmark,
vocabulary rescoring improved F-score from baseline to 92.2%, correctly
identifying 1,094 out of 1,271 domain-specific terms. Multi-word alias
support.

This PR builds on the work from
https://github.com/FluidInference/FluidAudio/pull/240

---------

Co-authored-by: Alex-Wengg <hanweng9@gmail.com>
2026-01-28 18:26:51 -05:00
Alex 892da4f9a9 Feat: Parakeet EOU streaming ASR with 160ms/320ms chunk support (#216)
- Add Parakeet EOU 120M streaming ASR with End-of-Utterance detection
- Support 160ms and 320ms chunk sizes with automatic HuggingFace model
downloads
- benchmarks.md 
- Add GitHub Actions CI benchmark workflow for Parakeet EOU



Changes
- StreamingEouAsrManager - streaming pipeline with configurable chunk
sizes
- NeMoMelSpectrogram - native Swift mel spectrogram with vDSP
vectorization
- RnntDecoder - RNN-T greedy decoder with EOU detection
- Configurable EOU debounce (default 1280ms)

---------
2025-12-17 17:18:01 -05:00
7fd5ac5446 pyannote community-1 model for offline speaker diarization pipeline (#150)
### Why is this change needed?
<!-- Explain the motivation for this change. What problem does it solve?
-->

Keeping the streaming one around as the VBx and AHC clustering gets
pretty expensive after 30mins of audio and running it constantly gets
expensive. Its still possible to support clustering between files but
will save that for another PR.

Pyannote's Bench mark is around 11% - i increased steps to 0.2s instead
of 0.1 to double the speed but also selective fp16 results in more
operations to run on ANE but also means that we lose some precision.

```
Average DER: 14.95% | Median DER: 10.89% | Average JER: 39.27% | Median JER: 40.74% (collar=0.25s, ignoreOverlap=True)
Average RTFx: 139.63 (from 232 clips)
Metrics summary saved to: /Users/brandonweng/FluidAudioDatasets/voxconverse/metrics/test_metrics_release.json
Completed. New results: 232, Skipped existing: 0, Total attempted: 232
```

See benchmark.md for more info but compared to Pytorch model, we are
100x faster than the CPU version and ~6x faster compared to the mps
backend on mb pro 4

---------

Co-authored-by: claude[bot] <209825114+claude[bot]@users.noreply.github.com>
Co-authored-by: Brandon Weng <BrandonWeng@users.noreply.github.com>
Co-authored-by: Alex <36247722+Alex-Wengg@users.noreply.github.com>
Co-authored-by: Alex-Wengg <hanweng9@gmail.com>
2025-10-22 15:11:57 -04:00
Brandon Weng eb00809e05 Add token timings to streaming call (#135)
### Why is this change needed?
<!-- Explain the motivation for this change. What problem does it solve?
-->

Returns token timings and timestamps in the streaming implmenetatin and
uses the global offsets to align the timestamps in each chunk

https://github.com/FluidInference/FluidAudio/issues/120
2025-10-09 15:22:50 -07:00
Alex 93bd9cf49a Kokoro Text-to-Speech (#112) 2025-10-06 17:53:30 -04:00
Brandon Weng a5eeed9a0d New parakeet-tdt-v3-0.6b models, ~50% faster (#113) 2025-09-19 23:26:15 -04:00
Brandon Weng 245880345a Cleanup AudioConverter (#103) 2025-09-13 12:33:30 -04:00
Brandon Weng 85aae5f092 Normalize WER based on Huggingface WER calculation (#84)
### Why is this change needed?
<!-- Explain the motivation for this change. What problem does it solve?
-->


https://github.com/huggingface/open_asr_leaderboard/blob/main/normalizer/normalizer.py#L528

We had some things missing so WER may have seen higher than it really
was.

(Accidentally pushed the first commit to main, thats why its revert ing
the revert :P)
2025-08-29 00:44:29 +00:00
Brandon Weng 186f3580aa Revert "Normalize WER better"
This reverts commit c89144cf41.
2025-08-28 20:23:27 -04:00
Brandon Weng c89144cf41 Normalize WER better 2025-08-28 20:22:58 -04:00
Brandon WengandAlex-Wengg 6ef85af4ec nvidia/parakeet-tdt-0.6b-v3 (#77)
Introduces TDT v3 pipeline, refactors ASR internals, and reorganizes
streaming components.
- Simplifies batch ASR via a shared AudioConverter and new
ChunkProcessor.
- Marks streaming as beta/unstable and removes dedicated streaming docs
for now.

Core ASR
- New: TDT stack under Sources/FluidAudio/ASR/TDT/
- TdtConfig.swift, TdtDecoder.swift (moved/refactored),
TdtDecoderState.swift (moved/refactored), TdtHypothesis.swift.
- New: ChunkProcessor.swift for unified chunked/batch processing.
- Removed: Legacy batch processors
    - ASR/AsrBatchProcessor.swift, ASR/BatchProcessor.swift.
- Updated: AsrManager.swift, AsrModels.swift,AsrTranscription.swift,
AsrTypes.swift to use TDT config/state, zero-copy chaining, and
refactors aligning with new pipeline.

Streaming
- Moved: ASR/StreamingAsrManager.swift →
ASR/Streaming/StreamingAsrManager.swift.
- Moved: ASR/StreamingAsrSession.swift →
ASR/Streaming/StreamingAsrSession.swift.
- Status: Streaming set to beta/unstable; further tuning required.

Shared Utilities
- Moved: ASR/AudioConverter.swift → Shared/AudioConverter.swift
(centralized, public actor; 16kHz mono Float32 conversion).

Diarization
- Updated: DiarizerTypes.swift, EmbeddingExtractor.swift,
SegmentationProcessor.swift (compatibility and refinements).

CLI
- Updated: Sources/FluidAudioCLI/Commands/AsrBenchmark.swift (align with
new chunking/TDT).
- Updated:
DiarizationBenchmark.swift,DownloadCommand.swift,StreamingTranscribeCommand.swift,
TextNormalizer.swift.

Documentation
- README: Updated to mark streaming as beta/unstable; removed streaming
quick-start; added batch quick-start
and a benchmark one-liner.
- Removed: Documentation/StreamingASR.md (temporarily while tuning
streaming).

Breaking/Behavior Changes
- Batch ASR now routes through TDT-aware chunking; legacy batch
processors removed.
- Streaming APIs relocated to ASR/Streaming/ namespace; behavior still
in flux (beta).

Migration Notes
- Replace any direct use of removed batch processors with
AsrManager.transcribe(_:)
- Update imports/paths for streaming classes under ASR/Streaming/.
- Re-test streaming flows; defaults and thresholds may change as tuning
continues.


## Todos
- Need to test other lanaguages, we traded the model with english, it
could be why its not handling other languages well.

---------

Co-authored-by: Alex-Wengg <hanweng9@gmail.com>
2025-08-25 18:20:28 -04:00
Brandon WengandClaude 9d30edbc60 Remove unused FP16 methods (#69)
Cleaning up some of the AI slop that got accidentally committed. Also
simplifying the naming of the methods too

---------

Co-authored-by: Claude <noreply@anthropic.com>
2025-08-15 18:50:45 -04:00
AlexandClaude 07680d15a0 Int8 quantization + memory optimizations for faster diarization (#59)
### Summary
Performance optimizations for speaker diarization through Int8
quantization and ANE
  memory optimization. 
- int8 quantization (30 -> 85 RTFx)
- ANE memory optimization ( 85 -> 150 RTFx)
- (~0.4 %) accuracy drop

 ### Details
 - Added support for Int8 quantized WeSpeaker embedding  models
    - Reduced model size and memory footprint
    - Improved inference speed on ANE
- ANE Memory Optimizer
    - Implemented ANEMemoryOptimizer for efficient memory management
    - Optimized embedding extraction pipeline for ANE execution
    - Reduced memory allocations during inference

---------

Co-authored-by: Claude <noreply@anthropic.com>
2025-08-05 21:44:22 -04:00
Brandon Weng 6336bbec71 Swift Format (#57)
To unify everyones formatting. 

Also added a hook to Claude Code to run on completion, github jobs to
fail if its not ran

Requires Swift 6 + 
`swift format --in-place --recursive --configuration .swift-format
Sources/ Tests/ Examples/`
2025-08-02 22:40:05 -04:00
Alex a59df27384 Parakeet TDT-0.6b ASR CoreML Support (#15)
- ASR manager Introduction
  - Introduce parakeet-tdt-0.6b-v2-coreml to leverage ANE processing
  - Token Duration Transducer (TDT) supported
- ASR benchmark measures WER & RTFx and their respective means, sum and
median.
    - https://huggingface.co/spaces/hf-audio/open_asr_leaderboard 
    - https://www.openslr.org/12 dataset test-clean & test-other
  - Text normalization post-processing support  
- Added two DecoderState variables to handle microphone and system audio
separately
- CLI code reorganization 
- Main.swift file size reduction & reorganization
- Transcribe chunking introduction (Parakeet)
- RTFx testing metrics
- ASR benchmark debugging
- Create once or update existing benchmark posts instead of creating new
benchmark comments each time.
2025-07-23 12:00:12 -04:00