Choose Your Model - Sanas Developer Hub
Who is listening?
A person is listening
You need a Human ↔ Human model. Voice quality and naturalness matter.
A machine is listening
You need a Human ↔ Machine model. Relative Word Error Rate (RWERR) reduction and ASR accuracy matter.
Human ↔ Human
| Model | Model Name | Scenario | Description |
|---|---|---|---|
| Speech Enhancement · Voice Isolation (General) | VI_G_SE |
Noisy environment, human listener | SE · Voice Isolation (General) |
| Speech Enhancement Standard | SE2.1 |
Telephony audio restoration, low CPU | Restores and enhances voice quality, outputs 8kHz |
| Speech Enhancement with Full-Fidelity | SE2.2 |
Full-fidelity audio restoration | Bandwidth extension to ultra-fidelity 24kHz |
| Accent Translation | AT5.2 |
Accent Normalization | Add Accent Translation to your application to modifies global accents in real-time, allowing your teams to be instantly understood. |
| Language Translation | LT | Voice-preserving Language translation | Add real-time speech-to-speech language translation to your application that preserves your speakers’ voices, tone, and intent. |
Human ↔ Machine
| Model | Model Name | Scenario | Description |
|---|---|---|---|
| Agentic SE · Voice Isolation (General) | AGENTIC_VI_G_SE |
Single speaker → ASR (16kHz) | Removes environmental noise and distant human voices. Single-speaker isolation for voice agents, IVR, phone bots |
| Agentic SE · Voice Isolation (Telephony) | AGENTIC_VI_GT_SE |
Single speaker → ASR (8kHz telephony) | Telephony-optimized variant of Voice Isolation for 8kHz narrowband audio |
| Agentic SE · Standard | AGENTIC_ST_SE |
Multiple speakers → ASR | Agentic SE · Standard Removes environmental noise, preserves distant human speech. Multi-speaker environments where background conversations carry context |
More models and model updates are coming soon.
Still Not Sure?
Explore audio samples, full specifications, use cases, and code examples:
Language Translation
LT Real-time language translation that preserves your speakers’ voices, tone, and intent.