Commit Graph
6 Commits
Author SHA1 Message Date
7fd5ac5446 pyannote community-1 model for offline speaker diarization pipeline (#150)
### Why is this change needed?
<!-- Explain the motivation for this change. What problem does it solve?
-->

Keeping the streaming one around as the VBx and AHC clustering gets
pretty expensive after 30mins of audio and running it constantly gets
expensive. Its still possible to support clustering between files but
will save that for another PR.

Pyannote's Bench mark is around 11% - i increased steps to 0.2s instead
of 0.1 to double the speed but also selective fp16 results in more
operations to run on ANE but also means that we lose some precision.

```
Average DER: 14.95% | Median DER: 10.89% | Average JER: 39.27% | Median JER: 40.74% (collar=0.25s, ignoreOverlap=True)
Average RTFx: 139.63 (from 232 clips)
Metrics summary saved to: /Users/brandonweng/FluidAudioDatasets/voxconverse/metrics/test_metrics_release.json
Completed. New results: 232, Skipped existing: 0, Total attempted: 232
```

See benchmark.md for more info but compared to Pytorch model, we are
100x faster than the CPU version and ~6x faster compared to the mps
backend on mb pro 4

---------

Co-authored-by: claude[bot] <209825114+claude[bot]@users.noreply.github.com>
Co-authored-by: Brandon Weng <BrandonWeng@users.noreply.github.com>
Co-authored-by: Alex <36247722+Alex-Wengg@users.noreply.github.com>
Co-authored-by: Alex-Wengg <hanweng9@gmail.com>
2025-10-22 15:11:57 -04:00
Alex 89a0a875c5 streaming speaker diarization management support (#63)
## Summary

This PR significantly improves speaker diarization performance and
simplifies the API by introducing
  comprehensive speaker management and streaming capabilities.

  ### Key Improvements

- **17.7% DER Achievement**: Optimized clustering threshold and
parameters deliver state-of-the-art
  performance
- **Streaming Support**: Real-time diarization with first-occurrence
speaker mapping for production use

  #### SpeakerManager
- `assignSpeaker()` - assigns embeddings to existing speakers or creates
new ones
- `initializeKnownSpeakers()` - Preload known speaker profiles for
recognition
  - `reset()` - Clear session data for new recordings

  #### Renamed for Clarity
- `minSpeechDuration` (was minDurationOn) - Minimum speech segment
duration
  - `minSilenceGap` (was minDurationOff) - Minimum gap between speakers
  - `clusteringThreshold` - Optimal at 0.7 for 17.7% DER

  #### Performance Parameters
  - `speakerThreshold` - For matching existing speakers (default: 0.65)
  - `embeddingThreshold` - For updating embeddings (default: 0.45)
- Real-time capable: 140+ RTFx on GitHub Actions, 150+ RTFx on Apple
Silicon

#### Others
  - Removed Hungarian algorithm dependency
  - Consolidated multiple test files into single comprehensive suite
  - Parameter renames for clarity
2025-08-14 02:29:33 -04:00
Brandon Weng 6336bbec71 Swift Format (#57)
To unify everyones formatting. 

Also added a hook to Claude Code to run on completion, github jobs to
fail if its not ran

Requires Swift 6 + 
`swift format --in-place --recursive --configuration .swift-format
Sources/ Tests/ Examples/`
2025-08-02 22:40:05 -04:00
Brandon Weng 875cdcf6db Fixes for ASR TDT and provide example for streaming ASR (#50)
Trying to introduce a streaming API that's similar to Apples OS 26
speech analytics one. While doing this I realized our models while the
encoder supports dynamic dimensions, the melspec model is hard coded to
10s and as a result it only really only supports 10 second chunks.

I will add a follow up PR to look more into this in order to support a
smaller window.

This PR also fixes and removes issues related to TDDT sentence decoding,
we had a lot of hard coded logic for post processing that's actually not
needed
2025-08-01 15:21:25 -04:00
panv-kw 91020294b8 Add a models parameter to DiarizerManager.initialize (#28)
Only the last commit is new.

Builds on #25 by adding a `models` parameter to
`DiarizerManager.initialize`.

The `DiarizerModels` type gives users APIs to download models, including
to custom directories, and load models that are bundled with the app or
downloaded from elsewhere. Models can also be loaded with user-specified
options, such as the compute units to use.

This unlocks some new use-cases but comes with some minor API changes.
IMO, it is actually quite nice to let users know that initialising the
diarization manager might involve a download, so I feel the extra
clarity is worth the extra characters.

The `DiarizerModels` APIs make the previous
`DiarizerManager.downloadModels` and
`DiarizerConfig.modelCacheDirectory` APIs obsolete, so they have also
been removed.
2025-07-24 18:43:08 +00:00
Alex a59df27384 Parakeet TDT-0.6b ASR CoreML Support (#15)
- ASR manager Introduction
  - Introduce parakeet-tdt-0.6b-v2-coreml to leverage ANE processing
  - Token Duration Transducer (TDT) supported
- ASR benchmark measures WER & RTFx and their respective means, sum and
median.
    - https://huggingface.co/spaces/hf-audio/open_asr_leaderboard 
    - https://www.openslr.org/12 dataset test-clean & test-other
  - Text normalization post-processing support  
- Added two DecoderState variables to handle microphone and system audio
separately
- CLI code reorganization 
- Main.swift file size reduction & reorganization
- Transcribe chunking introduction (Parakeet)
- RTFx testing metrics
- ASR benchmark debugging
- Create once or update existing benchmark posts instead of creating new
benchmark comments each time.
2025-07-23 12:00:12 -04:00