Commit Graph
12 Commits
Author SHA1 Message Date
Alex 7e51dc6903 refactor(parakeet): Improve consistency across ASR managers (#494)
This PR addresses three high-priority consistency improvements in the
Parakeet ASR folder from issue #457.

## Summary

- ✅ **Task 1:** Standardized lifecycle method names across all managers
(13 files)
- ✅ **Task 2:** Consolidated ~230 lines of duplicate token deduplication
logic
- ✅ **Task 3:** Extracted shared streaming code into reusable utilities

## Changes

### 1. Lifecycle Method Standardization

Unified naming conventions to eliminate confusion:

| Manager | Old Method | New Method |
|---------|-----------|------------|
| `AsrManager` | `loadModels(_:)` | `configure(models:)` |
| `SlidingWindowAsrSession` | `initialize()` | `loadModels()` |
| `SlidingWindowAsrManager` | `start()` | `startStreaming()` |
| `StreamingEouAsrManager` | `loadModelsFromHuggingFace()` |
`loadModels()` |

**Files updated:** 5 managers + 8 CLI commands

### 2. Token Deduplication Consolidation

Extracted duplicate matching algorithms into generic, type-safe
utilities:

**New Files:**
- `SequenceMatch.swift` - Data structure for sequence matches
- `SequenceMatcher.swift` - 5 reusable matching algorithms:
  - `findSuffixPrefixMatch()` - O(n) greedy boundary detection
  - `findBoundedSubstringMatch()` - Windowed search
  - `findLongestCommonSubsequence()` - O(n²) LCS via DP
  - `findContiguousMatches()` - Longest consecutive run
  - `consolidateMatches()` - Merge adjacent matches
- `TokenDeduplicationRegressionTests.swift` - 12 comprehensive tests

**Refactored:**
- `AsrManager+TokenProcessing.swift` - Reduced from ~65 to ~40 lines
(-38%)
- `ChunkProcessor.swift` - Removed ~77 lines of duplicate code

### 3. Streaming Code Extraction

Created utilities for common patterns in both `StreamingEouAsrManager`
and `StreamingNemotronAsrManager`:

**New Utilities:**
- `EncoderCacheManager` - Cache initialization and extraction
- `StreamingAsrUtils` - Audio buffering, state reset, token decoding

## Impact

| Metric | Result |
|--------|--------|
| **Duplicate code eliminated** | ~230 lines |
| **New reusable utilities** | 430 lines |
| **Test coverage** | +12 regression tests |
| **API consistency** | Unified lifecycle naming |
| **Performance** | No regression ✅ |
| **WER** | 0.4% (verified) ✅ |
| **RTFx** | 43.3x (verified) ✅ |
| **Tests** | 25/25 passing ✅ |

## Testing

```bash
# Token deduplication regression tests
swift test --filter TokenDeduplicationRegressionTests
# ✅ 12/12 tests passing

# Nemotron streaming tests
swift test --filter StreamingNemotronAsrManagerTests
# ✅ 16/16 tests passing

# ASR benchmark (no WER regression)
swift run -c release fluidaudiocli asr-benchmark --max-files 10
# ✅ WER: 0.4%, RTFx: 43.3x
```

## Breaking Changes

⚠️ This PR contains breaking API changes:
- Renamed lifecycle methods (no deprecation wrappers)
- All call sites updated in this PR

Closes #457

<!-- devin-review-badge-begin -->

---

<a href="https://app.devin.ai/review/fluidinference/fluidaudio/pull/494"
target="_blank">
  <picture>
<source media="(prefers-color-scheme: dark)"
srcset="https://static.devin.ai/assets/gh-open-in-devin-review-dark.svg?v=1">
<img
src="https://static.devin.ai/assets/gh-open-in-devin-review-light.svg?v=1"
alt="Open with Devin">
  </picture>
</a>
<!-- devin-review-badge-end -->

---------
2026-04-07 19:30:58 -04:00
Alex ea50062181 ASR architecture cleanup: naming, dead code, file organization 29/03/2026 (#457) (#468)
## Summary

Addresses #457 — ASR architecture inconsistencies, tech debt, and
misplaced code.

### Naming consistency
- Standardized `Manager` suffix: `StreamingAsrEngine` →
`StreamingAsrManager` (protocol)
- Streaming-first prefix: `EouStreamingAsrManager` →
`StreamingEouAsrManager`, `NemotronStreamingAsrManager` →
`StreamingNemotronAsrManager`
- `AsrManager.initialize(models:)` → `loadModels(_:)` (matches streaming
managers)
- `AsrManager.resetState()` → `reset()`

### Dead code removal
- Removed CTC logit caching from `AsrManager` (~60 lines) —
`SlidingWindowAsrManager` never read the cache, it runs its own CTC
inference via `CtcKeywordSpotter`
- Removed `StreamingAsrManagerFactory` — moved `createManager()` onto
`StreamingModelVariant` enum

### Lifecycle consistency
- Added `cleanup()` to `StreamingAsrManager` protocol and all
implementations
- Every ASR manager now has both `reset()` and `cleanup()`

### File organization
- Split `AsrManager+Transcription.swift` (441 lines) into:
  - `+Transcription.swift` (129 lines) — high-level API
  - `+Pipeline.swift` (152 lines) — CoreML inference
  - `+TokenProcessing.swift` (170 lines) — confidence, timings, dedup
- Moved `MLMultiArray.reset(to:)` to
`Shared/MLMultiArray+Extensions.swift`
- Made `transcribeChunk()` internal

## Verification

6 benchmarks × 100 files, zero WER regressions:

| Model | Baseline | Current | Delta |
|-------|----------|---------|-------|
| Parakeet TDT v3 | 2.6% | 2.64% | +0.04% |
| Parakeet TDT v2 | 3.8% | 3.79% | -0.01% |
| CTC-TDT 110M | 3.6% | 3.56% | -0.04% |
| CTC Earnings | 16.54% | 16.51% | -0.03% |
| EOU 320ms | 7.11% | 7.11% | +0.00% |
| Nemotron 1120ms | 1.99% | 1.99% | +0.00% |

## Test plan
- [x] `swift build` passes
- [x] All 6 subset benchmarks pass with zero WER regressions
- [ ] `swift test` CI passes

🤖 Generated with [Claude Code](https://claude.com/claude-code)
<!-- devin-review-badge-begin -->

---

<a href="https://app.devin.ai/review/fluidinference/fluidaudio/pull/468"
target="_blank">
  <picture>
<source media="(prefers-color-scheme: dark)"
srcset="https://static.devin.ai/assets/gh-open-in-devin-review-dark.svg?v=1">
<img
src="https://static.devin.ai/assets/gh-open-in-devin-review-light.svg?v=1"
alt="Open with Devin">
  </picture>
</a>
<!-- devin-review-badge-end -->
2026-03-29 20:29:50 -04:00
Alex aa800cb963 Convert AsrManager to actor for Swift 6 concurrency safety (#419)
Fixes #415

## Summary

Converts `AsrManager` from a class to an actor to fix Swift 6 strict
concurrency checking errors reported in issue #415. This eliminates data
race warnings when compiling with Xcode 16.4 RC's stricter concurrency
enforcement.

## Problem

With Swift 6 strict concurrency checking enabled, the compiler correctly
flags the following pattern as unsafe:

```swift
if let asrManager = asrManager {
    try await asrManager.resetDecoderState(for: audioSource)
}
```

The `nonisolated(unsafe)` workaround was hiding real data race risks.

## Solution

Convert `AsrManager` to an actor, which:
- Makes it automatically `Sendable` 
- Provides compiler-enforced data race safety
- Eliminates the need for unsafe workarounds
- Ensures all external access is properly isolated with `await`

## Changes

### Core Conversion
- **AsrManager.swift**: Changed `public final class AsrManager` →
`public actor AsrManager`
- Refactored `initializeDecoderState(decoderState: inout
TdtDecoderState)` to `initializeDecoderState(for: AudioSource)` to
handle actor isolation
- Modified `transcribeWithState` to take `source: AudioSource` instead
of `inout` decoder state

### Removed Unsafe Workarounds
- **StreamingAsrManager.swift**: Removed `nonisolated(unsafe)` from
`asrManager` property

### Updated Call Sites
- Added `await` to all actor method calls in:
  - `StreamingAsrManager.swift` (3 locations)
  - `ChunkProcessor.swift` (3 locations)
  - `TranscribeCommand.swift` (1 location)
  - `TTSCommand.swift` (2 locations)

### Marked Pure Functions as Nonisolated
- `extractFeatureValue`, `extractFeatureValues` - ML feature extraction
utilities
- `padAudioIfNeeded` - Audio padding helper
- `calculateStartFrameOffset` - Deprecated test compatibility helper

### Test Updates
- **AsrTranscriptionTests.swift**: Made test functions async and created
`setupMockVocabulary()` helper

## Testing

✅ All CI tests pass (13 tests, 0 failures)

```
Test Suite 'CITests' passed
Executed 13 tests, with 0 failures in 1.030 seconds
```

## Impact

- **Breaking Change**: Yes - external calls to `AsrManager` methods now
require `await`
- **Performance**: No impact - actor isolation has minimal overhead
- **Safety**: Significantly improved - compiler-enforced data race
safety
- **Compatibility**: Requires Swift 6 for full benefits

## Migration Guide

For users of FluidAudio:

```swift
// Before
let manager = AsrManager()
try await manager.initialize(models: models)
let result = try await manager.transcribe(audioBuffer)
manager.cleanup()

// After
let manager = AsrManager()
try await manager.initialize(models: models)
let result = try await manager.transcribe(audioBuffer)
await manager.cleanup()  // Add await
```
<!-- devin-review-badge-begin -->

---

<a href="https://app.devin.ai/review/fluidinference/fluidaudio/pull/419"
target="_blank">
  <picture>
<source media="(prefers-color-scheme: dark)"
srcset="https://static.devin.ai/assets/gh-open-in-devin-review-dark.svg?v=1">
<img
src="https://static.devin.ai/assets/gh-open-in-devin-review-light.svg?v=1"
alt="Open with Devin">
  </picture>
</a>
<!-- devin-review-badge-end -->
2026-03-24 17:26:08 -04:00
AlexandSachin Desai e318089cf5 Replace eSpeak with CoreML G2P model, rename to FluidAudioTTS (#350)
## Summary

- **Replace eSpeak NG C library** with a CoreML BART encoder-decoder
model for grapheme-to-phoneme conversion. G2P now runs entirely through
CoreML — no native C dependency.
- **Rename `FluidAudioEspeak`** target/product/module to
**`FluidAudioTTS`** and `EspeakG2P` class to `G2PModel`.
- **Remove `ESpeakNG.xcframework`** (~30MB binary) from the repo
entirely.
- **Add morphological stemming** as a fallback before G2P for inflected
words (-s/-ed/-ing), reducing CoreML inference calls.

### Breaking changes
- `FluidAudioEspeak` module is now `FluidAudioTTS` — update any `import
FluidAudioEspeak` to `import FluidAudioTTS`.

### Details
| Change | Files |
|--------|-------|
| CoreML G2P replaces espeak C API | `G2PModel.swift` (was
`EspeakG2P.swift`) |
| Morphological stemming (-s/-ed/-ing) | `KokoroChunker.swift` (+170
lines) |
| Module rename | `Package.swift`, all imports |
| Remove dead `#if canImport(ESpeakNG)` guards | 4 test files |
| Delete framework link tests | `FrameworkLinkTests.swift` |
| Remove framework binary | `Frameworks/ESpeakNG.xcframework/` |
| 24 new stemming unit tests | `KokoroChunkerStemTests.swift` |

## Test plan
- [x] `swift build` passes
- [x] `swift build --build-tests` passes
- [x] `swift test --filter KokoroChunkerStemTests` — 24/24 pass
- [x] `swift format lint` clean (no new warnings)
- [x] TTS .wav synthesis verified with common words (lexicon path)
- [x] TTS .wav synthesis verified with nonsense words (G2P model path)
- [x] iOS generation tested separately before branch push
<!-- devin-review-badge-begin -->

---

<a href="https://app.devin.ai/review/fluidinference/fluidaudio/pull/350"
target="_blank">
  <picture>
<source media="(prefers-color-scheme: dark)"
srcset="https://static.devin.ai/assets/gh-open-in-devin-review-dark.svg?v=1">
<img
src="https://static.devin.ai/assets/gh-open-in-devin-review-light.svg?v=1"
alt="Open with Devin">
  </picture>
</a>
<!-- devin-review-badge-end -->

---------

Co-authored-by: Sachin Desai <sdesai@salesforce.com>
2026-03-06 17:42:15 -05:00
Alex 80fec6feef refactor: rename FluidAudioTTS to FluidAudioEspeak (#302)
## Summary

- Rename `FluidAudioTTS` product/target to `FluidAudioEspeak`
- Clarifies that this product contains the ESpeakNG GPL dependency
- PocketTTS is now in core `FluidAudio` (MIT licensed, #301)

## Migration

```swift
// Before
import FluidAudioTTS

// After  
import FluidAudioEspeak
```

```swift
// Package.swift - Before
.product(name: "FluidAudioTTS", package: "FluidAudio")

// Package.swift - After
.product(name: "FluidAudioEspeak", package: "FluidAudio")
```
2026-02-11 22:18:23 -05:00
Alex a43f66f168 feat: add voice cloning support for PocketTTS (#289)
## Summary
- Add voice cloning capability using Mimi encoder model
- Clone voices from audio files or raw samples
- Use cloned voices directly for synthesis without file I/O
- Save/load cloned voice data for persistence

## New API

```swift
let manager = PocketTtsManager()
try await manager.initialize()

// Clone a voice
let voiceData = try await manager.cloneVoice(from: audioURL)

// Use immediately
let audio = try await manager.synthesize(text: "Hello!", voiceData: voiceData)

// Or save for later
try manager.saveClonedVoice(voiceData, to: outputURL)
```

## Changes
- `ModelNames.swift`: Add `mimiEncoder` model name
- `PocketTtsModelStore.swift`: Add lazy loading of Mimi encoder
- `PocketTtsSynthesizer.swift`: Add synthesize overload accepting
`PocketTtsVoiceData`
- `PocketTtsManager.swift`: Add public voice cloning API
- **New**: `PocketTtsVoiceCloner.swift`: Core voice cloning
implementation

## Test plan
- [ ] Build succeeds
- [ ] Existing TTS tests pass
- [ ] Voice cloning produces valid conditioning data
- [ ] Cloned voice synthesis produces audio
2026-02-11 19:35:55 -05:00
Alex 9fcdf2f32c feat: add PocketTTS backend for lightweight text-to-speech (#273)
## Summary
- Add PocketTTS as a new TTS backend — flow-matching language model with
autoregressive streaming synthesis
- Pure Swift implementation using 4 CoreML models (cond_step,
flowlm_step, flow_decoder, mimi_decoder)
- iOS 17 compatible — no `scaled_dot_product_attention` ops (avoids BNNS
crash)
- Add audio post-processor with de-esser for reducing sibilant harshness

## Test plan
- [x] Short sentence: WER 0, 3.44s audio
- [x] Long sentence: WER 0, 6.64s audio
- [x] Fresh HuggingFace download works end-to-end
- [x] iOS build succeeds (`xcodebuild -destination
'generic/platform=iOS'`)
- [x] macOS build succeeds (`swift build -c release`)
2026-02-02 23:49:47 -05:00
Alex 0afbabca21 feat: add TTS de-esser to reduce sibilant harshness (#267)
## Summary
- Add audio post-processing to reduce harsh sibilant sounds (s, sh, z)
in Kokoro TTS output
- De-esser uses a biquad high-shelf filter at 6kHz with -3dB reduction,
plus 80Hz high-pass for rumble removal
- Feature is on by default with `--no-deess` CLI flag to disable

## Test plan
- [x] Build succeeds
- [ ] Generate TTS audio and verify reduced sibilance compared to
previous versions
- [ ] Test `--no-deess` flag to ensure it bypasses the de-esser
2026-01-28 02:30:34 -05:00
Sachin Desai 3ff14b9eaf add support for custom vocabulary (#213)
### Why is this change needed?
Custom lexicons exist to let users override pronunciation for
domain-specific terminology that general-purpose G2P and built-in
dictionaries either do not support of mishandle —this is especially
common in finance (tickers, company names, jargon like “EBITDA”),
healthcare (drug names, procedures, acronyms), legal (Latinisms, case
citations), and tech (product names, abbreviations). The goal is for
these overrides to be dependable and to take effect exactly when the
user expects resulting in fluid audio generation.
2025-12-15 14:09:02 -05:00
AlexandBrandon Weng f5f30a4940 optionalize TTS via FluidAudioTTS target (#186)
- Make TTS optional to avoid GPL by default; enable with
`FLUIDAUDIO_ENABLE_TTS=1.`
- Split TTS into FluidAudioTTS target; CLI imports it only when enabled.
- Expose MLModel.compatPrediction as public for cross‑module use.

---------

Co-authored-by: Brandon Weng <18161326+BrandonWeng@users.noreply.github.com>
2025-11-24 00:55:57 -05:00
Brandon Weng eec3d961f7 Clean up unneeded version checks (#152)
### Why is this change needed?
<!-- Explain the motivation for this change. What problem does it solve?
-->

Pulling some of the changes from this PR here to break it down into
smaller PRs.

https://github.com/FluidInference/FluidAudio/pull/150
2025-10-20 22:58:43 -04:00
Alex 93bd9cf49a Kokoro Text-to-Speech (#112) 2025-10-06 17:53:30 -04:00