Files
Alex 7e51dc6903 refactor(parakeet): Improve consistency across ASR managers (#494)
This PR addresses three high-priority consistency improvements in the
Parakeet ASR folder from issue #457.

## Summary

- ✅ **Task 1:** Standardized lifecycle method names across all managers
(13 files)
- ✅ **Task 2:** Consolidated ~230 lines of duplicate token deduplication
logic
- ✅ **Task 3:** Extracted shared streaming code into reusable utilities

## Changes

### 1. Lifecycle Method Standardization

Unified naming conventions to eliminate confusion:

| Manager | Old Method | New Method |
|---------|-----------|------------|
| `AsrManager` | `loadModels(_:)` | `configure(models:)` |
| `SlidingWindowAsrSession` | `initialize()` | `loadModels()` |
| `SlidingWindowAsrManager` | `start()` | `startStreaming()` |
| `StreamingEouAsrManager` | `loadModelsFromHuggingFace()` |
`loadModels()` |

**Files updated:** 5 managers + 8 CLI commands

### 2. Token Deduplication Consolidation

Extracted duplicate matching algorithms into generic, type-safe
utilities:

**New Files:**
- `SequenceMatch.swift` - Data structure for sequence matches
- `SequenceMatcher.swift` - 5 reusable matching algorithms:
  - `findSuffixPrefixMatch()` - O(n) greedy boundary detection
  - `findBoundedSubstringMatch()` - Windowed search
  - `findLongestCommonSubsequence()` - O(n²) LCS via DP
  - `findContiguousMatches()` - Longest consecutive run
  - `consolidateMatches()` - Merge adjacent matches
- `TokenDeduplicationRegressionTests.swift` - 12 comprehensive tests

**Refactored:**
- `AsrManager+TokenProcessing.swift` - Reduced from ~65 to ~40 lines
(-38%)
- `ChunkProcessor.swift` - Removed ~77 lines of duplicate code

### 3. Streaming Code Extraction

Created utilities for common patterns in both `StreamingEouAsrManager`
and `StreamingNemotronAsrManager`:

**New Utilities:**
- `EncoderCacheManager` - Cache initialization and extraction
- `StreamingAsrUtils` - Audio buffering, state reset, token decoding

## Impact

| Metric | Result |
|--------|--------|
| **Duplicate code eliminated** | ~230 lines |
| **New reusable utilities** | 430 lines |
| **Test coverage** | +12 regression tests |
| **API consistency** | Unified lifecycle naming |
| **Performance** | No regression ✅ |
| **WER** | 0.4% (verified) ✅ |
| **RTFx** | 43.3x (verified) ✅ |
| **Tests** | 25/25 passing ✅ |

## Testing

```bash
# Token deduplication regression tests
swift test --filter TokenDeduplicationRegressionTests
# ✅ 12/12 tests passing

# Nemotron streaming tests
swift test --filter StreamingNemotronAsrManagerTests
# ✅ 16/16 tests passing

# ASR benchmark (no WER regression)
swift run -c release fluidaudiocli asr-benchmark --max-files 10
# ✅ WER: 0.4%, RTFx: 43.3x
```

## Breaking Changes

⚠️ This PR contains breaking API changes:
- Renamed lifecycle methods (no deprecation wrappers)
- All call sites updated in this PR

Closes #457

<!-- devin-review-badge-begin -->

---

<a href="https://app.devin.ai/review/fluidinference/fluidaudio/pull/494"
target="_blank">
  <picture>
<source media="(prefers-color-scheme: dark)"
srcset="https://static.devin.ai/assets/gh-open-in-devin-review-dark.svg?v=1">
<img
src="https://static.devin.ai/assets/gh-open-in-devin-review-light.svg?v=1"
alt="Open with Devin">
  </picture>
</a>
<!-- devin-review-badge-end -->

---------
2026-04-07 19:30:58 -04:00
..

FluidAudio CLI

The FluidAudio CLI provides various commands for audio processing, transcription, and benchmarking.

Installation

Build the CLI tool:

swift build -c release

Commands Overview

1. process - Audio Diarization

Process a single audio file to identify speakers and their segments.

# Basic usage
swift run fluidaudiocli process audio.wav

# With custom output and threshold
swift run fluidaudiocli process audio.wav --output results.json --threshold 0.7

# With debug mode
swift run fluidaudiocli process audio.wav --debug

Options:

  • --output <file>: Output JSON file (default: prints to console)
  • --threshold <value>: Clustering threshold 0.0-1.0 (default: 0.7)
  • --debug: Enable debug output

2. transcribe - Audio Transcription

Transcribe audio files using streaming ASR with real-time updates.

# Basic transcription
swift run fluidaudiocli transcribe audio.wav

# With low-latency configuration
swift run fluidaudiocli transcribe audio.wav --config low-latency

# With debug output
swift run fluidaudiocli transcribe audio.wav --debug

# Compare with direct ASR API
swift run fluidaudiocli transcribe audio.wav --compare

Options:

  • --config <type>: Configuration type: default, low-latency, high-accuracy
  • --debug: Show debug information
  • --compare: Compare streaming API with direct ASR API
  • --help, -h: Show help message

Configurations:

  • default: 2.5s chunks, 0.85 confirmation threshold
  • low-latency: 2.0s chunks, 0.75 confirmation threshold
  • high-accuracy: 3.0s chunks, 0.90 confirmation threshold

3. multi-stream - Parallel Transcription

Transcribe multiple audio files in parallel using shared ASR models.

# Process two different files
swift run fluidaudiocli multi-stream mic_audio.wav system_audio.wav

# Process same file on both streams
swift run fluidaudiocli multi-stream audio.wav

# With debug output
swift run fluidaudiocli multi-stream audio1.wav audio2.wav --debug

Options:

  • --debug: Show debug information
  • --help, -h: Show help message

4. diarization-benchmark - Speaker Diarization Benchmark

Run comprehensive benchmarks on evaluation datasets.

# Run on AMI dataset with auto-download
swift run fluidaudiocli diarization-benchmark --auto-download

# Test single file
swift run fluidaudiocli diarization-benchmark --single-file ES2004a --threshold 0.7

# Run on specific dataset
swift run fluidaudiocli diarization-benchmark --dataset ami-sdm --max-files 10

# Save results to file
swift run fluidaudiocli diarization-benchmark --output benchmark_results.json

Options:

  • --dataset <name>: Dataset to use (ami-sdm, ami-mdm, voxconverse)
  • --auto-download: Automatically download required datasets
  • --single-file <id>: Test a single file (e.g., ES2004a)
  • --threshold <value>: Clustering threshold (default: 0.7)
  • --max-files <n>: Maximum files to process
  • --output <file>: Save results to JSON file
  • --verbose: Show detailed progress

5. vad-benchmark - Voice Activity Detection Benchmark

Benchmark VAD performance on test datasets.

# Run VAD benchmark
swift run fluidaudiocli vad-benchmark --num-files 40

# With custom threshold 
swift run fluidaudiocli vad-benchmark --threshold 0.8

# Test on specific dataset
swift run fluidaudiocli vad-benchmark --dataset voices-subset

Options:

  • --num-files <n>: Number of files to test (or --all-files)
  • --threshold <value>: VAD threshold (default: 0.3)
  • --dataset <name>: Dataset to use (e.g., mini50, voices-subset, musan-full)
  • --debug: Verbose logging and per-file RTFx

6. asr-benchmark - ASR Benchmark

Benchmark ASR performance on LibriSpeech or other datasets.

# Run on LibriSpeech test-clean
swift run fluidaudiocli asr-benchmark --subset test-clean --max-files 100

# Run on test-other subset
swift run fluidaudiocli asr-benchmark --subset test-other --max-files 50

# With verbose output
swift run fluidaudiocli asr-benchmark --verbose

Options:

  • --subset <name>: LibriSpeech subset (test-clean, test-other)
  • --max-files <n>: Maximum files to process
  • --verbose: Show detailed progress

7. parakeet-eou - Streaming ASR

Real-time streaming transcription with end-of-utterance detection.

swift run fluidaudiocli parakeet-eou --input audio.wav --use-cache
swift run fluidaudiocli parakeet-eou --benchmark --chunk-size 160 --max-files 100 --use-cache

Options: --input <path>, --benchmark, --max-files <n>, --chunk-size <160|320|1600>, --eou-debounce <ms>, --use-cache, --models <path>, --output <path>, --verbose

8. download - Download Datasets

swift run fluidaudiocli download --dataset ami-sdm
swift run fluidaudiocli download --list

Options: --dataset <name>, --list

Output Formats

Diarization Output (JSON)

{
  "audioFile": "audio.wav",
  "durationSeconds": 300.5,
  "speakerCount": 3,
  "segments": [
    {
      "speakerId": "Speaker 1",
      "startTimeSeconds": 0.0,
      "endTimeSeconds": 45.2,
      "qualityScore": 0.85
    }
  ],
  "processingTimeSeconds": 15.2,
  "realTimeFactor": 0.05
}

Transcription Output

  • Real-time updates showing volatile and confirmed text
  • Final transcription with performance metrics
  • RTFx (Real-Time Factor) showing processing speed

Performance Notes

  • All commands process audio as fast as possible (no artificial delays)
  • Multi-stream command demonstrates parallel processing with shared models
  • Benchmarks provide detailed performance metrics including DER, WER, and RTFx

Examples

Complete Workflow Example

# 1. Download dataset
swift run fluidaudiocli download --dataset ami-sdm

# 2. Run diarization benchmark
swift run fluidaudiocli diarization-benchmark --dataset ami-sdm --output results.json

# 3. Process individual file
swift run fluidaudiocli process audio.wav --threshold 0.7

# 4. Transcribe audio
swift run fluidaudiocli transcribe audio.wav --config low-latency

# 5. Multi-stream transcription
swift run fluidaudiocli multi-stream mic.wav system.wav

Quick Test

# Test with included sample files
swift run fluidaudiocli transcribe medical.wav
swift run fluidaudiocli process IS1001a.Mix-Headset.wav --threshold 0.7

Troubleshooting

  1. Model Download Issues: The CLI will automatically download required models on first use. Ensure you have internet connectivity.

  2. Memory Usage: For long audio files, ensure sufficient memory is available.

  3. Performance: Use release build (swift build -c release) for best performance.

  4. Audio Format: The CLI automatically handles various audio formats and sample rates.