This PR addresses three high-priority consistency improvements in the Parakeet ASR folder from issue #457. ## Summary - ✅ **Task 1:** Standardized lifecycle method names across all managers (13 files) - ✅ **Task 2:** Consolidated ~230 lines of duplicate token deduplication logic - ✅ **Task 3:** Extracted shared streaming code into reusable utilities ## Changes ### 1. Lifecycle Method Standardization Unified naming conventions to eliminate confusion: | Manager | Old Method | New Method | |---------|-----------|------------| | `AsrManager` | `loadModels(_:)` | `configure(models:)` | | `SlidingWindowAsrSession` | `initialize()` | `loadModels()` | | `SlidingWindowAsrManager` | `start()` | `startStreaming()` | | `StreamingEouAsrManager` | `loadModelsFromHuggingFace()` | `loadModels()` | **Files updated:** 5 managers + 8 CLI commands ### 2. Token Deduplication Consolidation Extracted duplicate matching algorithms into generic, type-safe utilities: **New Files:** - `SequenceMatch.swift` - Data structure for sequence matches - `SequenceMatcher.swift` - 5 reusable matching algorithms: - `findSuffixPrefixMatch()` - O(n) greedy boundary detection - `findBoundedSubstringMatch()` - Windowed search - `findLongestCommonSubsequence()` - O(n²) LCS via DP - `findContiguousMatches()` - Longest consecutive run - `consolidateMatches()` - Merge adjacent matches - `TokenDeduplicationRegressionTests.swift` - 12 comprehensive tests **Refactored:** - `AsrManager+TokenProcessing.swift` - Reduced from ~65 to ~40 lines (-38%) - `ChunkProcessor.swift` - Removed ~77 lines of duplicate code ### 3. Streaming Code Extraction Created utilities for common patterns in both `StreamingEouAsrManager` and `StreamingNemotronAsrManager`: **New Utilities:** - `EncoderCacheManager` - Cache initialization and extraction - `StreamingAsrUtils` - Audio buffering, state reset, token decoding ## Impact | Metric | Result | |--------|--------| | **Duplicate code eliminated** | ~230 lines | | **New reusable utilities** | 430 lines | | **Test coverage** | +12 regression tests | | **API consistency** | Unified lifecycle naming | | **Performance** | No regression ✅ | | **WER** | 0.4% (verified) ✅ | | **RTFx** | 43.3x (verified) ✅ | | **Tests** | 25/25 passing ✅ | ## Testing ```bash # Token deduplication regression tests swift test --filter TokenDeduplicationRegressionTests # ✅ 12/12 tests passing # Nemotron streaming tests swift test --filter StreamingNemotronAsrManagerTests # ✅ 16/16 tests passing # ASR benchmark (no WER regression) swift run -c release fluidaudiocli asr-benchmark --max-files 10 # ✅ WER: 0.4%, RTFx: 43.3x ``` ## Breaking Changes ⚠️ This PR contains breaking API changes: - Renamed lifecycle methods (no deprecation wrappers) - All call sites updated in this PR Closes #457 <!-- devin-review-badge-begin --> --- <a href="https://app.devin.ai/review/fluidinference/fluidaudio/pull/494" target="_blank"> <picture> <source media="(prefers-color-scheme: dark)" srcset="https://static.devin.ai/assets/gh-open-in-devin-review-dark.svg?v=1"> <img src="https://static.devin.ai/assets/gh-open-in-devin-review-light.svg?v=1" alt="Open with Devin"> </picture> </a> <!-- devin-review-badge-end --> ---------
FluidAudio CLI
The FluidAudio CLI provides various commands for audio processing, transcription, and benchmarking.
Installation
Build the CLI tool:
swift build -c release
Commands Overview
1. process - Audio Diarization
Process a single audio file to identify speakers and their segments.
# Basic usage
swift run fluidaudiocli process audio.wav
# With custom output and threshold
swift run fluidaudiocli process audio.wav --output results.json --threshold 0.7
# With debug mode
swift run fluidaudiocli process audio.wav --debug
Options:
--output <file>: Output JSON file (default: prints to console)--threshold <value>: Clustering threshold 0.0-1.0 (default: 0.7)--debug: Enable debug output
2. transcribe - Audio Transcription
Transcribe audio files using streaming ASR with real-time updates.
# Basic transcription
swift run fluidaudiocli transcribe audio.wav
# With low-latency configuration
swift run fluidaudiocli transcribe audio.wav --config low-latency
# With debug output
swift run fluidaudiocli transcribe audio.wav --debug
# Compare with direct ASR API
swift run fluidaudiocli transcribe audio.wav --compare
Options:
--config <type>: Configuration type:default,low-latency,high-accuracy--debug: Show debug information--compare: Compare streaming API with direct ASR API--help, -h: Show help message
Configurations:
default: 2.5s chunks, 0.85 confirmation thresholdlow-latency: 2.0s chunks, 0.75 confirmation thresholdhigh-accuracy: 3.0s chunks, 0.90 confirmation threshold
3. multi-stream - Parallel Transcription
Transcribe multiple audio files in parallel using shared ASR models.
# Process two different files
swift run fluidaudiocli multi-stream mic_audio.wav system_audio.wav
# Process same file on both streams
swift run fluidaudiocli multi-stream audio.wav
# With debug output
swift run fluidaudiocli multi-stream audio1.wav audio2.wav --debug
Options:
--debug: Show debug information--help, -h: Show help message
4. diarization-benchmark - Speaker Diarization Benchmark
Run comprehensive benchmarks on evaluation datasets.
# Run on AMI dataset with auto-download
swift run fluidaudiocli diarization-benchmark --auto-download
# Test single file
swift run fluidaudiocli diarization-benchmark --single-file ES2004a --threshold 0.7
# Run on specific dataset
swift run fluidaudiocli diarization-benchmark --dataset ami-sdm --max-files 10
# Save results to file
swift run fluidaudiocli diarization-benchmark --output benchmark_results.json
Options:
--dataset <name>: Dataset to use (ami-sdm, ami-mdm, voxconverse)--auto-download: Automatically download required datasets--single-file <id>: Test a single file (e.g., ES2004a)--threshold <value>: Clustering threshold (default: 0.7)--max-files <n>: Maximum files to process--output <file>: Save results to JSON file--verbose: Show detailed progress
5. vad-benchmark - Voice Activity Detection Benchmark
Benchmark VAD performance on test datasets.
# Run VAD benchmark
swift run fluidaudiocli vad-benchmark --num-files 40
# With custom threshold
swift run fluidaudiocli vad-benchmark --threshold 0.8
# Test on specific dataset
swift run fluidaudiocli vad-benchmark --dataset voices-subset
Options:
--num-files <n>: Number of files to test (or--all-files)--threshold <value>: VAD threshold (default: 0.3)--dataset <name>: Dataset to use (e.g.,mini50,voices-subset,musan-full)--debug: Verbose logging and per-file RTFx
6. asr-benchmark - ASR Benchmark
Benchmark ASR performance on LibriSpeech or other datasets.
# Run on LibriSpeech test-clean
swift run fluidaudiocli asr-benchmark --subset test-clean --max-files 100
# Run on test-other subset
swift run fluidaudiocli asr-benchmark --subset test-other --max-files 50
# With verbose output
swift run fluidaudiocli asr-benchmark --verbose
Options:
--subset <name>: LibriSpeech subset (test-clean, test-other)--max-files <n>: Maximum files to process--verbose: Show detailed progress
7. parakeet-eou - Streaming ASR
Real-time streaming transcription with end-of-utterance detection.
swift run fluidaudiocli parakeet-eou --input audio.wav --use-cache
swift run fluidaudiocli parakeet-eou --benchmark --chunk-size 160 --max-files 100 --use-cache
Options: --input <path>, --benchmark, --max-files <n>, --chunk-size <160|320|1600>, --eou-debounce <ms>, --use-cache, --models <path>, --output <path>, --verbose
8. download - Download Datasets
swift run fluidaudiocli download --dataset ami-sdm
swift run fluidaudiocli download --list
Options: --dataset <name>, --list
Output Formats
Diarization Output (JSON)
{
"audioFile": "audio.wav",
"durationSeconds": 300.5,
"speakerCount": 3,
"segments": [
{
"speakerId": "Speaker 1",
"startTimeSeconds": 0.0,
"endTimeSeconds": 45.2,
"qualityScore": 0.85
}
],
"processingTimeSeconds": 15.2,
"realTimeFactor": 0.05
}
Transcription Output
- Real-time updates showing volatile and confirmed text
- Final transcription with performance metrics
- RTFx (Real-Time Factor) showing processing speed
Performance Notes
- All commands process audio as fast as possible (no artificial delays)
- Multi-stream command demonstrates parallel processing with shared models
- Benchmarks provide detailed performance metrics including DER, WER, and RTFx
Examples
Complete Workflow Example
# 1. Download dataset
swift run fluidaudiocli download --dataset ami-sdm
# 2. Run diarization benchmark
swift run fluidaudiocli diarization-benchmark --dataset ami-sdm --output results.json
# 3. Process individual file
swift run fluidaudiocli process audio.wav --threshold 0.7
# 4. Transcribe audio
swift run fluidaudiocli transcribe audio.wav --config low-latency
# 5. Multi-stream transcription
swift run fluidaudiocli multi-stream mic.wav system.wav
Quick Test
# Test with included sample files
swift run fluidaudiocli transcribe medical.wav
swift run fluidaudiocli process IS1001a.Mix-Headset.wav --threshold 0.7
Troubleshooting
-
Model Download Issues: The CLI will automatically download required models on first use. Ensure you have internet connectivity.
-
Memory Usage: For long audio files, ensure sufficient memory is available.
-
Performance: Use release build (
swift build -c release) for best performance. -
Audio Format: The CLI automatically handles various audio formats and sample rates.