Fixes#415
## Summary
Converts `AsrManager` from a class to an actor to fix Swift 6 strict
concurrency checking errors reported in issue #415. This eliminates data
race warnings when compiling with Xcode 16.4 RC's stricter concurrency
enforcement.
## Problem
With Swift 6 strict concurrency checking enabled, the compiler correctly
flags the following pattern as unsafe:
```swift
if let asrManager = asrManager {
try await asrManager.resetDecoderState(for: audioSource)
}
```
The `nonisolated(unsafe)` workaround was hiding real data race risks.
## Solution
Convert `AsrManager` to an actor, which:
- Makes it automatically `Sendable`
- Provides compiler-enforced data race safety
- Eliminates the need for unsafe workarounds
- Ensures all external access is properly isolated with `await`
## Changes
### Core Conversion
- **AsrManager.swift**: Changed `public final class AsrManager` →
`public actor AsrManager`
- Refactored `initializeDecoderState(decoderState: inout
TdtDecoderState)` to `initializeDecoderState(for: AudioSource)` to
handle actor isolation
- Modified `transcribeWithState` to take `source: AudioSource` instead
of `inout` decoder state
### Removed Unsafe Workarounds
- **StreamingAsrManager.swift**: Removed `nonisolated(unsafe)` from
`asrManager` property
### Updated Call Sites
- Added `await` to all actor method calls in:
- `StreamingAsrManager.swift` (3 locations)
- `ChunkProcessor.swift` (3 locations)
- `TranscribeCommand.swift` (1 location)
- `TTSCommand.swift` (2 locations)
### Marked Pure Functions as Nonisolated
- `extractFeatureValue`, `extractFeatureValues` - ML feature extraction
utilities
- `padAudioIfNeeded` - Audio padding helper
- `calculateStartFrameOffset` - Deprecated test compatibility helper
### Test Updates
- **AsrTranscriptionTests.swift**: Made test functions async and created
`setupMockVocabulary()` helper
## Testing
✅ All CI tests pass (13 tests, 0 failures)
```
Test Suite 'CITests' passed
Executed 13 tests, with 0 failures in 1.030 seconds
```
## Impact
- **Breaking Change**: Yes - external calls to `AsrManager` methods now
require `await`
- **Performance**: No impact - actor isolation has minimal overhead
- **Safety**: Significantly improved - compiler-enforced data race
safety
- **Compatibility**: Requires Swift 6 for full benefits
## Migration Guide
For users of FluidAudio:
```swift
// Before
let manager = AsrManager()
try await manager.initialize(models: models)
let result = try await manager.transcribe(audioBuffer)
manager.cleanup()
// After
let manager = AsrManager()
try await manager.initialize(models: models)
let result = try await manager.transcribe(audioBuffer)
await manager.cleanup() // Add await
```
<!-- devin-review-badge-begin -->
---
<a href="https://app.devin.ai/review/fluidinference/fluidaudio/pull/419"
target="_blank">
<picture>
<source media="(prefers-color-scheme: dark)"
srcset="https://static.devin.ai/assets/gh-open-in-devin-review-dark.svg?v=1">
<img
src="https://static.devin.ai/assets/gh-open-in-devin-review-light.svg?v=1"
alt="Open with Devin">
</picture>
</a>
<!-- devin-review-badge-end -->
## Summary
- Add voice cloning capability using Mimi encoder model
- Clone voices from audio files or raw samples
- Use cloned voices directly for synthesis without file I/O
- Save/load cloned voice data for persistence
## New API
```swift
let manager = PocketTtsManager()
try await manager.initialize()
// Clone a voice
let voiceData = try await manager.cloneVoice(from: audioURL)
// Use immediately
let audio = try await manager.synthesize(text: "Hello!", voiceData: voiceData)
// Or save for later
try manager.saveClonedVoice(voiceData, to: outputURL)
```
## Changes
- `ModelNames.swift`: Add `mimiEncoder` model name
- `PocketTtsModelStore.swift`: Add lazy loading of Mimi encoder
- `PocketTtsSynthesizer.swift`: Add synthesize overload accepting
`PocketTtsVoiceData`
- `PocketTtsManager.swift`: Add public voice cloning API
- **New**: `PocketTtsVoiceCloner.swift`: Core voice cloning
implementation
## Test plan
- [ ] Build succeeds
- [ ] Existing TTS tests pass
- [ ] Voice cloning produces valid conditioning data
- [ ] Cloned voice synthesis produces audio
## Summary
- Add PocketTTS as a new TTS backend — flow-matching language model with
autoregressive streaming synthesis
- Pure Swift implementation using 4 CoreML models (cond_step,
flowlm_step, flow_decoder, mimi_decoder)
- iOS 17 compatible — no `scaled_dot_product_attention` ops (avoids BNNS
crash)
- Add audio post-processor with de-esser for reducing sibilant harshness
## Test plan
- [x] Short sentence: WER 0, 3.44s audio
- [x] Long sentence: WER 0, 6.64s audio
- [x] Fresh HuggingFace download works end-to-end
- [x] iOS build succeeds (`xcodebuild -destination
'generic/platform=iOS'`)
- [x] macOS build succeeds (`swift build -c release`)
## Summary
- Add audio post-processing to reduce harsh sibilant sounds (s, sh, z)
in Kokoro TTS output
- De-esser uses a biquad high-shelf filter at 6kHz with -3dB reduction,
plus 80Hz high-pass for rumble removal
- Feature is on by default with `--no-deess` CLI flag to disable
## Test plan
- [x] Build succeeds
- [ ] Generate TTS audio and verify reduced sibilance compared to
previous versions
- [ ] Test `--no-deess` flag to ensure it bypasses the de-esser
### Why is this change needed?
Custom lexicons exist to let users override pronunciation for
domain-specific terminology that general-purpose G2P and built-in
dictionaries either do not support of mishandle —this is especially
common in finance (tickers, company names, jargon like “EBITDA”),
healthcare (drug names, procedures, acronyms), legal (Latinisms, case
citations), and tech (product names, abbreviations). The goal is for
these overrides to be dependable and to take effect exactly when the
user expects resulting in fluid audio generation.
- Make TTS optional to avoid GPL by default; enable with
`FLUIDAUDIO_ENABLE_TTS=1.`
- Split TTS into FluidAudioTTS target; CLI imports it only when enabled.
- Expose MLModel.compatPrediction as public for cross‑module use.
---------
Co-authored-by: Brandon Weng <18161326+BrandonWeng@users.noreply.github.com>
### Why is this change needed?
<!-- Explain the motivation for this change. What problem does it solve?
-->
Pulling some of the changes from this PR here to break it down into
smaller PRs.
https://github.com/FluidInference/FluidAudio/pull/150