## Summary
Systematic cleanup of the ASR module addressing tech debt items from
#457. Net reduction of ~430 lines while fixing real bugs and improving
maintainability.
### Bug fixes
- **`enableFP16` silently ignored** —
`optimizedConfiguration(enableFP16:)` delegated to a shared factory that
hardcoded `allowLowPrecisionAccumulationOnGPU = true`, ignoring the
caller's parameter
- **`MLArrayCache.returnArray` only reset float32 data** — cached arrays
of other types (float16, int32) retained stale data from previous use
- **CTC model auto-detection broken** —
`Repo.parakeetCtc110m.folderName` returned `"parakeet-ctc-110m"` instead
of `"parakeet-ctc-110m-coreml"` because the `folderName` switch fell
through to a `default` case that stripped the `-coreml` suffix. Same for
`parakeetCtc06b`.
- **Duplicate tokens at chunk merge boundary** — `mergeByMidpoint` used
`<=`/`>=` so tokens exactly at the cutoff appeared in both left and
right chunks
### Dead code removal
- Deleted `ANEOptimizer` indirection layer (166 lines) — was a
pass-through wrapping `MLModel` with no optimization
- Deleted `PerformanceMonitor` actor and `AggregatedMetrics` — never
instantiated, component times hardcoded to 0
- Deleted `getFloat16Array` from MLArrayCache — never called
- Deleted `sliceEncoderOutput` from AsrTranscription — never called (30
lines)
- Deleted `loadWithANEOptimization` from AsrModels — never called
- Removed unused `tokenTimings` parameter chain through
`processTranscriptionResult`
- Removed unused `import OSLog` / `import CoreML` across 5 files
- Removed `nonisolated(unsafe)` from SlidingWindowAsrManager (types
already Sendable)
### Duplication elimination
- Extracted `clearCachedCtcData()` helper (replaced 3× triple-nil
assignments)
- Extracted `decoderState(for:)` / `setDecoderState(_:for:)` (replaced
4× switch blocks)
- Extracted `frameAlignedAudio()` (replaced 2× duplicated
frame-alignment blocks)
- Added `ASRConstants.secondsPerEncoderFrame` (replaced 5× magic `0.08`)
- Replaced hardcoded `16_000` with `config.sampleRate` /
`ASRConstants.sampleRate`
- Extracted `MLModelConfigurationUtils.defaultConfiguration()` (replaced
5× copy-pasted config methods)
- Extracted `MLModelConfigurationUtils.defaultModelsDirectory()`
(replaced 3× copy-pasted directory methods)
- Consolidated duplicate `vocabularyFile` / `vocabularyFileArray`
constants
### File organization
- Moved `PerformanceMetrics.swift`, `ProgressEmitter.swift`,
`MLArrayCache.swift` from `ASR/Parakeet/` to `Shared/` (used by multiple
modules)
- Renamed `StreamingAudioSourceFactory` → `AudioSourceFactory`,
`StreamingAudioSampleSource` → `AudioSampleSource` (types used by both
ASR and Diarizer)
- Renamed files to match type names: `SortformerDiarizerPipeline.swift`
→ `SortformerDiarizer.swift`, `LSEENDDiarizerAPI.swift` →
`LSEENDDiarizer.swift`, `NemotronPipeline.swift` →
`NemotronStreamingAsrManager+Pipeline.swift`
- Replaced force unwraps in `RnntDecoder.swift` with `guard let` +
descriptive errors
- Removed stale TODO about decoder state in AsrManager
### Benchmark script
- Added `Scripts/run_parakeet_benchmarks.sh` — runs all 6 benchmarks
(v3, v2, TDT-CTC-110M, CTC earnings, EOU 320ms, Nemotron 1120ms) with
WER comparison against `benchmarks100.md` baselines and regression
detection
- Referenced from `Documentation/ASR/benchmarks100.md`
## Verified — no regressions
```
Model Baseline Current Delta
Parakeet TDT v3 (0.6B) 2.6% 2.64% +0.04%
Parakeet TDT v2 (0.6B) 3.8% 3.79% -0.01%
CTC-TDT 110M 3.6% 3.56% -0.04%
CTC Earnings 16.54% 16.51% -0.03%
EOU 320ms (120M) 7.11% 7.11% +0.00%
Nemotron 1120ms (0.6B) 1.99% 1.99% +0.00%
```
## Test plan
- [x] `swift build` passes
- [x] `swift test` passes (all existing tests, updated for removed dead
code)
- [x] All 6 ASR benchmarks match baselines (100 files each)
- [ ] `swift format lint` passes
- Update Package.swift to swift-tools-version: 6.0
- Add @preconcurrency import CoreML/AVFoundation throughout codebase
- Make structs Sendable (AppLogger, DownloadConfig, etc.)
- Use nonisolated(unsafe) for static mutable state
- Fix AudioStream by removing @unchecked Sendable, making AsyncCallback
@Sendable
- Fix Task closures with proper explicit captures
- Convert concurrent tests to sequential where types aren't Sendable
- Add @MainActor to test methods using waitForExpectations
🤖 Generated with [Claude Code](https://claude.com/claude-code)
### Why is this change needed?
<!-- Explain the motivation for this change. What problem does it solve?
-->
resolve#231
- support audio streaming for speaker diarization
- Fixed `SegmentProcessor.getSegments` internal access level
- Added audio conversion method for `CMSampleBuffer`
- Added a `startTime` argument to
`DiarizerManager.performCompleteDiarization` to offset segment
timestamps.
- Added `AudioStream` struct to simplify timestamp calculations for a
streamed/online audio input source with a lot of overlap
---------
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
### Why is this change needed?
<!-- Explain the motivation for this change. What problem does it solve?
-->
Keeping the streaming one around as the VBx and AHC clustering gets
pretty expensive after 30mins of audio and running it constantly gets
expensive. Its still possible to support clustering between files but
will save that for another PR.
Pyannote's Bench mark is around 11% - i increased steps to 0.2s instead
of 0.1 to double the speed but also selective fp16 results in more
operations to run on ANE but also means that we lose some precision.
```
Average DER: 14.95% | Median DER: 10.89% | Average JER: 39.27% | Median JER: 40.74% (collar=0.25s, ignoreOverlap=True)
Average RTFx: 139.63 (from 232 clips)
Metrics summary saved to: /Users/brandonweng/FluidAudioDatasets/voxconverse/metrics/test_metrics_release.json
Completed. New results: 232, Skipped existing: 0, Total attempted: 232
```
See benchmark.md for more info but compared to Pytorch model, we are
100x faster than the CPU version and ~6x faster compared to the mps
backend on mb pro 4
---------
Co-authored-by: claude[bot] <209825114+claude[bot]@users.noreply.github.com>
Co-authored-by: Brandon Weng <BrandonWeng@users.noreply.github.com>
Co-authored-by: Alex <36247722+Alex-Wengg@users.noreply.github.com>
Co-authored-by: Alex-Wengg <hanweng9@gmail.com>
### Why is this change needed?
The post processing logic we had previously was convoluted and hacked
together.
- Remove the normalization
- Remove the post processing of the output from the model
- Refactor the model name and centralize in ModelNames
- Batch processing of the chunks, RTFx, 100 RTx --> 1200 RTFx
- Accelerate for array operations helped a ton here too, increased by
100 RTFx -> 220 RTFx
### Why is this change needed?
Avoid the annoying "print" when developing and testing the CLI. Also
making sure we don't miss things in the logger.error calls when running
in the CLI
`swift build` + integration tests should pass
## Summary
This PR significantly improves speaker diarization performance and
simplifies the API by introducing
comprehensive speaker management and streaming capabilities.
### Key Improvements
- **17.7% DER Achievement**: Optimized clustering threshold and
parameters deliver state-of-the-art
performance
- **Streaming Support**: Real-time diarization with first-occurrence
speaker mapping for production use
#### SpeakerManager
- `assignSpeaker()` - assigns embeddings to existing speakers or creates
new ones
- `initializeKnownSpeakers()` - Preload known speaker profiles for
recognition
- `reset()` - Clear session data for new recordings
#### Renamed for Clarity
- `minSpeechDuration` (was minDurationOn) - Minimum speech segment
duration
- `minSilenceGap` (was minDurationOff) - Minimum gap between speakers
- `clusteringThreshold` - Optimal at 0.7 for 17.7% DER
#### Performance Parameters
- `speakerThreshold` - For matching existing speakers (default: 0.65)
- `embeddingThreshold` - For updating embeddings (default: 0.45)
- Real-time capable: 140+ RTFx on GitHub Actions, 150+ RTFx on Apple
Silicon
#### Others
- Removed Hungarian algorithm dependency
- Consolidated multiple test files into single comprehensive suite
- Parameter renames for clarity
Trying to introduce a streaming API that's similar to Apples OS 26
speech analytics one. While doing this I realized our models while the
encoder supports dynamic dimensions, the melspec model is hard coded to
10s and as a result it only really only supports 10 second chunks.
I will add a follow up PR to look more into this in order to support a
smaller window.
This PR also fixes and removes issues related to TDDT sentence decoding,
we had a lot of hard coded logic for post processing that's actually not
needed
- Download the essential flies like config.json and the mlmodelc folders
and their files
- Simplified downloading models
- Make sure models like vad, asr and diarizer models are downloaded once
---------
Co-authored-by: Claude <noreply@anthropic.com>
- ASR manager Introduction
- Introduce parakeet-tdt-0.6b-v2-coreml to leverage ANE processing
- Token Duration Transducer (TDT) supported
- ASR benchmark measures WER & RTFx and their respective means, sum and
median.
- https://huggingface.co/spaces/hf-audio/open_asr_leaderboard
- https://www.openslr.org/12 dataset test-clean & test-other
- Text normalization post-processing support
- Added two DecoderState variables to handle microphone and system audio
separately
- CLI code reorganization
- Main.swift file size reduction & reorganization
- Transcribe chunking introduction (Parakeet)
- RTFx testing metrics
- ASR benchmark debugging
- Create once or update existing benchmark posts instead of creating new
benchmark comments each time.