nanosamur.ai now supports multiple models across all three transcription stages. Each selected model produces its own transcript, called a track. You can use one model for live text, another for the final recording, or several models at each stage.
Our recent Qwen3-ASR vs Faster-Whisper comparison showed one of the reasons why this matters. Tracks let you use the best model or host multiple models and route your audio to the one that's best for given job! Also, you can use tracks for model evaluation and benchmarking - to compare those differences directly.
Choose models for each stage
- Real-time: text appears while you speak.
- Refined (semi-realtime): workers process short audio windows during the session, with more context and a delay behind live text.
- Final (batch): workers process the complete recording after you stop.
The supported model families cover these stages:
| Model family | Real-time | Refined | Final |
|---|---|---|---|
| Whisper | Faster-Whisper | WhisperX | WhisperX |
| Qwen | Qwen3-ASR | Qwen3-ASR | Qwen3-ASR |
| Nemotron | Nemotron 3.5 ASR | Not available | Not available |
| Parakeet | Not available | Parakeet TDT | Parakeet TDT |
Faster-Whisper and WhisperX run by default. Qwen, Nemotron, and Parakeet are optional. For example, you could compare Faster-Whisper and Nemotron live, then compare WhisperX and Parakeet on the completed recording. Each result keeps its model label.
The client sends audio once. SamuraiBFF routes it to the selected live services and publishes it once to Kafka. Refinement workers read audio windows, while the recorder saves the session to S3-compatible storage. Each selected final worker reads that recording. Refined and final tracks are saved in PostgreSQL. The transcription lifecycle guide explains the flow.
Before starting audio, open Session settings and select the available models for each stage. Disable stages you do not need. Keep recording storage enabled when selecting multiple final tracks.
Follow the model setup guide to add models and combine tracks. Refined and final model selection needs newer service versions than the base stack currently pins. Each running model also needs its own RAM and GPU memory; clearing a session selection does not unload it.