### Why is this change needed?
Automatic speech recognition (ASR) systems are trained on massive
datasets of general speech, which means they excel at common vocabulary
but struggle with domain-specific terminology. In business
contexts—earnings calls, medical dictation, legal proceedings—the most
critical words are often the ones the model has rarely or never seen:
company names like "Saoirse Ronan," product names like "Newrez," or
technical jargon unique to an industry. Without intervention, these
high-value terms get transcribed as phonetically similar but incorrect
common words, turning "Nequi" into "NECI" or "Bose" into "Boz."
Keyword boosting (also called context biasing or vocabulary rescoring)
addresses this gap by incorporating domain knowledge at inference time.
The system is given a list of expected vocabulary terms and uses
acoustic evidence—typically CTC log-probabilities—to determine whether
the audio actually supports replacing a transcribed word with a
vocabulary term. This isn't blind substitution; a well-designed rescorer
computes scores for both the original transcription and the candidate
vocabulary term, only making replacements when the acoustic evidence
favors the domain term. The "context-biasing weight" parameter allows
tuning how aggressively to prefer vocabulary terms.
The real-world impact is substantial. In our earnings call benchmark,
vocabulary rescoring improved F-score from baseline to 92.2%, correctly
identifying 1,094 out of 1,271 domain-specific terms. Multi-word alias
support.
This PR builds on the work from
https://github.com/FluidInference/FluidAudio/pull/240
---------
Co-authored-by: Alex-Wengg <hanweng9@gmail.com>
### Why is this change needed?
<!-- Explain the motivation for this change. What problem does it solve?
-->
Keeping the streaming one around as the VBx and AHC clustering gets
pretty expensive after 30mins of audio and running it constantly gets
expensive. Its still possible to support clustering between files but
will save that for another PR.
Pyannote's Bench mark is around 11% - i increased steps to 0.2s instead
of 0.1 to double the speed but also selective fp16 results in more
operations to run on ANE but also means that we lose some precision.
```
Average DER: 14.95% | Median DER: 10.89% | Average JER: 39.27% | Median JER: 40.74% (collar=0.25s, ignoreOverlap=True)
Average RTFx: 139.63 (from 232 clips)
Metrics summary saved to: /Users/brandonweng/FluidAudioDatasets/voxconverse/metrics/test_metrics_release.json
Completed. New results: 232, Skipped existing: 0, Total attempted: 232
```
See benchmark.md for more info but compared to Pytorch model, we are
100x faster than the CPU version and ~6x faster compared to the mps
backend on mb pro 4
---------
Co-authored-by: claude[bot] <209825114+claude[bot]@users.noreply.github.com>
Co-authored-by: Brandon Weng <BrandonWeng@users.noreply.github.com>
Co-authored-by: Alex <36247722+Alex-Wengg@users.noreply.github.com>
Co-authored-by: Alex-Wengg <hanweng9@gmail.com>
### Why is this change needed?
<!-- Explain the motivation for this change. What problem does it solve?
-->
Returns token timings and timestamps in the streaming implmenetatin and
uses the global offsets to align the timestamps in each chunk
https://github.com/FluidInference/FluidAudio/issues/120
Introduces TDT v3 pipeline, refactors ASR internals, and reorganizes
streaming components.
- Simplifies batch ASR via a shared AudioConverter and new
ChunkProcessor.
- Marks streaming as beta/unstable and removes dedicated streaming docs
for now.
Core ASR
- New: TDT stack under Sources/FluidAudio/ASR/TDT/
- TdtConfig.swift, TdtDecoder.swift (moved/refactored),
TdtDecoderState.swift (moved/refactored), TdtHypothesis.swift.
- New: ChunkProcessor.swift for unified chunked/batch processing.
- Removed: Legacy batch processors
- ASR/AsrBatchProcessor.swift, ASR/BatchProcessor.swift.
- Updated: AsrManager.swift, AsrModels.swift,AsrTranscription.swift,
AsrTypes.swift to use TDT config/state, zero-copy chaining, and
refactors aligning with new pipeline.
Streaming
- Moved: ASR/StreamingAsrManager.swift →
ASR/Streaming/StreamingAsrManager.swift.
- Moved: ASR/StreamingAsrSession.swift →
ASR/Streaming/StreamingAsrSession.swift.
- Status: Streaming set to beta/unstable; further tuning required.
Shared Utilities
- Moved: ASR/AudioConverter.swift → Shared/AudioConverter.swift
(centralized, public actor; 16kHz mono Float32 conversion).
Diarization
- Updated: DiarizerTypes.swift, EmbeddingExtractor.swift,
SegmentationProcessor.swift (compatibility and refinements).
CLI
- Updated: Sources/FluidAudioCLI/Commands/AsrBenchmark.swift (align with
new chunking/TDT).
- Updated:
DiarizationBenchmark.swift,DownloadCommand.swift,StreamingTranscribeCommand.swift,
TextNormalizer.swift.
Documentation
- README: Updated to mark streaming as beta/unstable; removed streaming
quick-start; added batch quick-start
and a benchmark one-liner.
- Removed: Documentation/StreamingASR.md (temporarily while tuning
streaming).
Breaking/Behavior Changes
- Batch ASR now routes through TDT-aware chunking; legacy batch
processors removed.
- Streaming APIs relocated to ASR/Streaming/ namespace; behavior still
in flux (beta).
Migration Notes
- Replace any direct use of removed batch processors with
AsrManager.transcribe(_:)
- Update imports/paths for streaming classes under ASR/Streaming/.
- Re-test streaming flows; defaults and thresholds may change as tuning
continues.
## Todos
- Need to test other lanaguages, we traded the model with english, it
could be why its not handling other languages well.
---------
Co-authored-by: Alex-Wengg <hanweng9@gmail.com>
Cleaning up some of the AI slop that got accidentally committed. Also
simplifying the naming of the methods too
---------
Co-authored-by: Claude <noreply@anthropic.com>
To unify everyones formatting.
Also added a hook to Claude Code to run on completion, github jobs to
fail if its not ran
Requires Swift 6 +
`swift format --in-place --recursive --configuration .swift-format
Sources/ Tests/ Examples/`
- ASR manager Introduction
- Introduce parakeet-tdt-0.6b-v2-coreml to leverage ANE processing
- Token Duration Transducer (TDT) supported
- ASR benchmark measures WER & RTFx and their respective means, sum and
median.
- https://huggingface.co/spaces/hf-audio/open_asr_leaderboard
- https://www.openslr.org/12 dataset test-clean & test-other
- Text normalization post-processing support
- Added two DecoderState variables to handle microphone and system audio
separately
- CLI code reorganization
- Main.swift file size reduction & reorganization
- Transcribe chunking introduction (Parakeet)
- RTFx testing metrics
- ASR benchmark debugging
- Create once or update existing benchmark posts instead of creating new
benchmark comments each time.