Choose Your Model - Sanas Developer Hub

Who is listening?

A person is listening

You need a Human ↔ Human model. Voice quality and naturalness matter.

A machine is listening

You need a Human ↔ Machine model. Relative Word Error Rate (RWERR) reduction and ASR accuracy matter.

Human ↔ Human

Model Model Name Scenario Description
Speech Enhancement · Voice Isolation (General) VI_G_SE Noisy environment, human listener SE · Voice Isolation (General)
Speech Enhancement Standard SE2.1 Telephony audio restoration, low CPU Restores and enhances voice quality, outputs 8kHz
Speech Enhancement with Full-Fidelity SE2.2 Full-fidelity audio restoration Bandwidth extension to ultra-fidelity 24kHz
Accent Translation AT5.2 Accent Normalization Add Accent Translation to your application to modifies global accents in real-time, allowing your teams to be instantly understood.
Language Translation LT Voice-preserving Language translation Add real-time speech-to-speech language translation to your application that preserves your speakers’ voices, tone, and intent.

Human ↔ Machine

Model Model Name Scenario Description
Agentic SE · Voice Isolation (General) AGENTIC_VI_G_SE Single speaker → ASR (16kHz) Removes environmental noise and distant human voices. Single-speaker isolation for voice agents, IVR, phone bots
Agentic SE · Voice Isolation (Telephony) AGENTIC_VI_GT_SE Single speaker → ASR (8kHz telephony) Telephony-optimized variant of Voice Isolation for 8kHz narrowband audio
Agentic SE · Standard AGENTIC_ST_SE Multiple speakers → ASR Agentic SE · Standard Removes environmental noise, preserves distant human speech. Multi-speaker environments where background conversations carry context

More models and model updates are coming soon.

Still Not Sure?

Explore audio samples, full specifications, use cases, and code examples:

Speech Enhancement · Voice Isolation (General)
VI_G_SE Isolates intended speech by removing background noise and voices. Optimized for human listeners.

Speech Enhancement · Standard
SE2.1 Restores and enhances voice quality for telephony audio. Low CPU footprint.

Speech Enhancement · Full Fidelity
SE2.2 Full-fidelity speech enhancement with bandwidth extension to ultra-fidelity 24kHz.

Agentic Speech Enhancement · Voice Isolation (General)
AGENTIC_VI_G_SE Removes background noise and distant voices for complete voice isolation of the primary speaker’s audio stream.

Agentic Speech Enhancement · Voice Isolation (Telephony)
AGENTIC_VI_GT_SE Telephony-optimized variant of Voice Isolation for 8kHz narrowband audio.

Agentic Speech Enhancement · Standard
AGENTIC_ST_SE Removes background noise while preserving all human speech for multi-speaker environments.

Accent Translation
AT5.2 Accent Translation modifies global accents in real-time, allowing your teams to be instantly understood while preserving what makes every voice unique.

Language Translation
LT Real-time language translation that preserves your speakers’ voices, tone, and intent.