Commit Graph
12 Commits
Author SHA1 Message Date
Alex 2593f55415 Add Japanese ASR support with JSUT and Common Voice datasets (#478)
## Summary

Adds comprehensive Japanese ASR support to FluidAudio with benchmark
datasets and CLI commands.

## Changes

### Core Japanese ASR Support
- **CtcJaManager.swift** - Japanese CTC transcription manager
(actor-based)
- **CtcJaModels.swift** - Japanese model loading and management
- **ModelNames.swift** - Added Japanese model registry (`parakeetCtcJa`,
`CTCJa` enum)
- **AsrModels.swift** - Added `.ctcJa` model version (3,072 vocab, 1,024
hidden, blank_id=3072)
- **AsrManager.swift** - Added `.ctcJa` case with error directing to
`CtcJaManager`

### CLI Commands
- **JapaneseAsrBenchmark.swift** (459 lines) - New `ja-benchmark`
command
  - JSUT basic5000 dataset support
  - Mozilla Common Voice (MCV) test set support
  - Auto-download capability
  - CER (Character Error Rate) evaluation
- **DownloadCommand.swift** - Added JSUT and MCV Japanese dataset
downloads
- **TranscribeCommand.swift** - Added `.ctcJa` model version support
- **AsrBenchmark.swift** - Added `.ctcJa` switch case

### Dataset Support
- **JapaneseDatasetDownloader.swift** (387 lines) - Dataset download and
parsing
  - JSUT basic5000 (5,000 sentences, clean studio recordings)
  - Mozilla Common Voice Japanese test split
  - Efficient streaming downloads
  - Metadata extraction and validation

## Usage

### CLI Commands
```bash
# Benchmark on JSUT basic5000 (100 samples)
swift run fluidaudiocli ja-benchmark --dataset jsut --samples 100

# Benchmark on Common Voice test (500 samples, auto-download)
swift run fluidaudiocli ja-benchmark --dataset cv-test --samples 500 --auto-download

# Download datasets
swift run fluidaudiocli download --dataset jsut
swift run fluidaudiocli download --dataset cv-ja-test
```

### Swift API
```swift
// Load and use Japanese CTC transcription
let manager = try await CtcJaManager.load()
let text = try manager.transcribe(audioURL: japaneseAudioFile)
```

## Model Info
- **Repo**: `FluidInference/parakeet-ctc-0.6b-ja-coreml`
- **Architecture**: 600M parameter CTC-only
- **Vocabulary**: 3,072 Japanese SentencePiece tokens + 1 blank (id:
3072)
- **Encoder**: 1,024 hidden size
- **Expected CER**: 6.5% on JSUT basic5000, 13.3% on MCV 16.1 test

## Testing
- ✅ Builds successfully (`swift build`)
- ✅ Model loading integration tested
- ✅ CLI commands compile and link correctly
- ⏳ Runtime benchmark testing pending (requires model download)

## Related
- Mobius PR #39: Japanese CTC CoreML conversion
(https://github.com/FluidInference/mobius/pull/39)

🤖 Generated with Claude Code
<!-- devin-review-badge-begin -->

---

<a href="https://app.devin.ai/review/fluidinference/fluidaudio/pull/478"
target="_blank">
  <picture>
<source media="(prefers-color-scheme: dark)"
srcset="https://static.devin.ai/assets/gh-open-in-devin-review-dark.svg?v=1">
<img
src="https://static.devin.ai/assets/gh-open-in-devin-review-light.svg?v=1"
alt="Open with Devin">
  </picture>
</a>
<!-- devin-review-badge-end -->

---------
2026-04-04 12:57:32 -04:00
Sachin DesaiandAlex-Wengg 8d3ce44ae1 feat: Custom vocabulary support (#251)
### Why is this change needed?
Automatic speech recognition (ASR) systems are trained on massive
datasets of general speech, which means they excel at common vocabulary
but struggle with domain-specific terminology. In business
contexts—earnings calls, medical dictation, legal proceedings—the most
critical words are often the ones the model has rarely or never seen:
company names like "Saoirse Ronan," product names like "Newrez," or
technical jargon unique to an industry. Without intervention, these
high-value terms get transcribed as phonetically similar but incorrect
common words, turning "Nequi" into "NECI" or "Bose" into "Boz."

Keyword boosting (also called context biasing or vocabulary rescoring)
addresses this gap by incorporating domain knowledge at inference time.
The system is given a list of expected vocabulary terms and uses
acoustic evidence—typically CTC log-probabilities—to determine whether
the audio actually supports replacing a transcribed word with a
vocabulary term. This isn't blind substitution; a well-designed rescorer
computes scores for both the original transcription and the candidate
vocabulary term, only making replacements when the acoustic evidence
favors the domain term. The "context-biasing weight" parameter allows
tuning how aggressively to prefer vocabulary terms.

The real-world impact is substantial. In our earnings call benchmark,
vocabulary rescoring improved F-score from baseline to 92.2%, correctly
identifying 1,094 out of 1,271 domain-specific terms. Multi-word alias
support.

This PR builds on the work from
https://github.com/FluidInference/FluidAudio/pull/240

---------

Co-authored-by: Alex-Wengg <hanweng9@gmail.com>
2026-01-28 18:26:51 -05:00
Brandon Weng a5eeed9a0d New parakeet-tdt-v3-0.6b models, ~50% faster (#113) 2025-09-19 23:26:15 -04:00
Brandon Weng ad51096d0b Unified logger for CLI commands too (#97)
### Why is this change needed?
Avoid the annoying "print" when developing and testing the CLI. Also
making sure we don't miss things in the logger.error calls when running
in the CLI

`swift build` + integration tests should pass
2025-09-09 19:21:51 -04:00
Brandon WengandAlex-Wengg 6ef85af4ec nvidia/parakeet-tdt-0.6b-v3 (#77)
Introduces TDT v3 pipeline, refactors ASR internals, and reorganizes
streaming components.
- Simplifies batch ASR via a shared AudioConverter and new
ChunkProcessor.
- Marks streaming as beta/unstable and removes dedicated streaming docs
for now.

Core ASR
- New: TDT stack under Sources/FluidAudio/ASR/TDT/
- TdtConfig.swift, TdtDecoder.swift (moved/refactored),
TdtDecoderState.swift (moved/refactored), TdtHypothesis.swift.
- New: ChunkProcessor.swift for unified chunked/batch processing.
- Removed: Legacy batch processors
    - ASR/AsrBatchProcessor.swift, ASR/BatchProcessor.swift.
- Updated: AsrManager.swift, AsrModels.swift,AsrTranscription.swift,
AsrTypes.swift to use TDT config/state, zero-copy chaining, and
refactors aligning with new pipeline.

Streaming
- Moved: ASR/StreamingAsrManager.swift →
ASR/Streaming/StreamingAsrManager.swift.
- Moved: ASR/StreamingAsrSession.swift →
ASR/Streaming/StreamingAsrSession.swift.
- Status: Streaming set to beta/unstable; further tuning required.

Shared Utilities
- Moved: ASR/AudioConverter.swift → Shared/AudioConverter.swift
(centralized, public actor; 16kHz mono Float32 conversion).

Diarization
- Updated: DiarizerTypes.swift, EmbeddingExtractor.swift,
SegmentationProcessor.swift (compatibility and refinements).

CLI
- Updated: Sources/FluidAudioCLI/Commands/AsrBenchmark.swift (align with
new chunking/TDT).
- Updated:
DiarizationBenchmark.swift,DownloadCommand.swift,StreamingTranscribeCommand.swift,
TextNormalizer.swift.

Documentation
- README: Updated to mark streaming as beta/unstable; removed streaming
quick-start; added batch quick-start
and a benchmark one-liner.
- Removed: Documentation/StreamingASR.md (temporarily while tuning
streaming).

Breaking/Behavior Changes
- Batch ASR now routes through TDT-aware chunking; legacy batch
processors removed.
- Streaming APIs relocated to ASR/Streaming/ namespace; behavior still
in flux (beta).

Migration Notes
- Replace any direct use of removed batch processors with
AsrManager.transcribe(_:)
- Update imports/paths for streaming classes under ASR/Streaming/.
- Re-test streaming flows; defaults and thresholds may change as tuning
continues.


## Todos
- Need to test other lanaguages, we traded the model with english, it
could be why its not handling other languages well.

---------

Co-authored-by: Alex-Wengg <hanweng9@gmail.com>
2025-08-25 18:20:28 -04:00
Brandon WengandClaude 8cc38d7d05 Only allow proxy for macOS and add test for iOS test with xCode build (#75)
## Description

### Why is this change needed?
[breaks builds for
iOS](https://github.com/FluidInference/FluidAudio/issues/74)

### What does this change do?
Only build the proxy components for macOS, not iOS

---------

Co-authored-by: Claude <noreply@anthropic.com>
2025-08-20 16:34:59 -04:00
Brandon Weng 4ac37cbfc9 Fix melspectrogram typo (#68)
Uses new v2 model with fixes to the melspectrogram typo but also updated
the model to support shorter chunks (won't fail now) but we still need
to update the code to support the shorter chunks. The current streaming
implementation has issues here, its not properly implemented
2025-08-15 18:23:59 -04:00
Brandon Weng 875cdcf6db Fixes for ASR TDT and provide example for streaming ASR (#50)
Trying to introduce a streaming API that's similar to Apples OS 26
speech analytics one. While doing this I realized our models while the
encoder supports dynamic dimensions, the melspec model is hard coded to
10s and as a result it only really only supports 10 second chunks.

I will add a follow up PR to look more into this in order to support a
smaller window.

This PR also fixes and removes issues related to TDDT sentence decoding,
we had a lot of hard coded logic for post processing that's actually not
needed
2025-08-01 15:21:25 -04:00
Brandon Weng 0d364b99ca Add Voice dataset (subset) benchmark for vad (#43)
MUSAN is mostly used for training, it would be better to benchmark
against something like VOiCES dataset. Added a command to test and run
it on a small subset (50).

Updated the Github workflow for VAD as well.
2025-07-29 16:04:50 -04:00
Brandon Weng b36e0ac88e New version of encoder and move parts of TDT to CoreML (#41)
Seeing a ~70% improvement in RTFx when running locally on M4 Pro

Most of the improvement comes from moving one of the transpose
operations onto CoreML intead of CPU

Locally with 1000 files, RTFx went from ~40x --> 69x

The new encoder model (_v2) is also half the size of the previous one,
we could use an even smaller model but there's some sacrifices to WER
and requires iOS 18+. We may upload it in case users want that offering
since its ~300MB, 1/3 the size of the original encoder


```
2620 files per dataset • Test runtime: 4m 47s • 07/26/2025, 11:16 PM EDT
--- Benchmark Results ---
   Mode: RELEASE (optimal performance)
   Dataset: librispeech test-clean
   Files processed: 2620
   Average WER: 3.4%
   Median WER: 0.0%
   Average CER: 1.3%
   Median RTFx: 65.9x
   Overall RTFx: 69.7x (19452.5s / 279.1s)
```


```
📋 Processing 2939 files (max files limit: unlimited)

2939 files per dataset • Test runtime: 4m 54s • 07/26/2025, 11:22 PM EDT
--- Benchmark Results ---
   Mode: RELEASE (optimal performance)
   Dataset: librispeech test-other
   Files processed: 2939
   Average WER: 5.4%
   Median WER: 0.0%
   Average CER: 2.3%
   Median RTFx: 60.2x
   Overall RTFx: 65.5x (19229.6s / 293.7s)

Results saved to: asr_benchmark_results.json
ASR benchmark completed successfully
```
2025-07-27 12:15:02 -04:00
AlexandClaude 3da6d4a4e3 simplified downloadutil.swift (#36)
- Download the essential flies like config.json and the mlmodelc folders
and their files
- Simplified downloading models 
- Make sure models like vad, asr and diarizer models are downloaded once

---------

Co-authored-by: Claude <noreply@anthropic.com>
2025-07-24 22:09:18 -04:00
Alex a59df27384 Parakeet TDT-0.6b ASR CoreML Support (#15)
- ASR manager Introduction
  - Introduce parakeet-tdt-0.6b-v2-coreml to leverage ANE processing
  - Token Duration Transducer (TDT) supported
- ASR benchmark measures WER & RTFx and their respective means, sum and
median.
    - https://huggingface.co/spaces/hf-audio/open_asr_leaderboard 
    - https://www.openslr.org/12 dataset test-clean & test-other
  - Text normalization post-processing support  
- Added two DecoderState variables to handle microphone and system audio
separately
- CLI code reorganization 
- Main.swift file size reduction & reorganization
- Transcribe chunking introduction (Parakeet)
- RTFx testing metrics
- ASR benchmark debugging
- Create once or update existing benchmark posts instead of creating new
benchmark comments each time.
2025-07-23 12:00:12 -04:00