mirror of
https://github.com/FluidInference/FluidAudio.git
synced 2026-06-11 20:24:36 +00:00
## Summary - Add CharsiuG2P ByT5 CoreML multilingual G2P model (`MultilingualG2PModel`, `MultilingualG2PLanguage`, `MultilingualG2PError`) supporting 9 Kokoro-mapped languages - Add `g2p-benchmark` CLI command measuring PER/WER/speed against CharsiuG2P test set with JSON output - Switch both English and multilingual G2P models to `cpuOnly` compute units (benchmarked 2-3x faster than GPU/ANE for autoregressive decoding) - Add `LevenshteinDistance` utility and `MultilingualG2PTests` (9 tests) ### Benchmark Results (M2, CPU-only, 500 words/language) | Language | PER | WER | ms/word | |---|---|---|---| | Spanish | 0.1% | 0.8% | 32.6 | | French | 0.8% | 2.0% | 26.5 | | Italian | 2.8% | 20.0% | 20.9 | | Hindi | 4.5% | 21.4% | 45.4 | | Japanese | 10.5% | 23.8% | 31.7 | | Portuguese | 8.9% | 43.2% | 24.0 | | British English | 13.6% | 29.4% | 34.0 | | American English | 19.0% | 38.8% | 28.2 | | Chinese | 86.2% | 95.0% | 53.9 | ### Compute Unit Benchmarks (English BART G2P) | Config | ms/word | |---|---| | cpuOnly | **13.0** | | all (ANE+GPU+CPU) | 17.3 | | cpuAndGPU | 23.4 | ## Test plan - [ ] `swift build` compiles clean - [ ] `swift test --filter MultilingualG2PTests` passes (9 tests) - [ ] `fluidaudiocli g2p-benchmark --languages eng-us --max-words 10 --data-dir <path>` produces results - [ ] Verify JSON output file is written correctly <!-- devin-review-badge-begin --> --- <a href="https://app.devin.ai/review/fluidinference/fluidaudio/pull/367" target="_blank"> <picture> <source media="(prefers-color-scheme: dark)" srcset="https://static.devin.ai/assets/gh-open-in-devin-review-dark.svg?v=1"> <img src="https://static.devin.ai/assets/gh-open-in-devin-review-light.svg?v=1" alt="Open with Devin"> </picture> </a> <!-- devin-review-badge-end -->