mirror of
https://github.com/FluidInference/FluidAudio.git
synced 2026-06-11 20:24:36 +00:00
## Summary - Add support for Nemotron streaming ASR with 160ms and 80ms chunk sizes - Expose chunk size variants that were already available on HuggingFace but not in the public API ## Changes - **NemotronChunkSize**: Add `.ms160` and `.ms80` enum cases - **ModelNames**: Add `nemotronStreaming160` and `nemotronStreaming80` to `Repo` enum with correct subdirectory mappings - **CLI Commands**: Update `NemotronTranscribe` and `NemotronBenchmark` to accept 160 and 80ms options - **Tests**: Update `NemotronChunkSizeTests` to verify all 4 chunk size variants ## Available Chunk Sizes | Chunk Size | Latency | Use Case | |------------|---------|----------| | 1120ms | 1.12s | Best accuracy & speed (original) | | 560ms | 0.56s | Lower latency | | 160ms | 0.16s | Very low latency | | 80ms | 0.08s | Ultra low latency | ## Usage Examples \`\`\`bash # Transcribe with 160ms chunks fluidaudio nemotron-transcribe --input audio.wav --chunk 160 # Benchmark with 80ms chunks fluidaudio nemotron-benchmark --chunk 80 --max-files 50 \`\`\` ## Test Plan - ✅ All `NemotronChunkSizeTests` pass - ✅ Build completes successfully - ✅ swift-format compliance verified <!-- devin-review-badge-begin --> --- <a href="https://app.devin.ai/review/fluidinference/fluidaudio/pull/490" target="_blank"> <picture> <source media="(prefers-color-scheme: dark)" srcset="https://static.devin.ai/assets/gh-open-in-devin-review-dark.svg?v=1"> <img src="https://static.devin.ai/assets/gh-open-in-devin-review-light.svg?v=1" alt="Open with Devin"> </picture> </a> <!-- devin-review-badge-end -->