
A real-time voice model and API that handles interruptions, adapts mid-sentence, and carries conversations the way people actually talk. It leads major deployed real-time APIs on audio benchmarks, supports 100+ languages, and is compatible with the OpenAI Realtime API.


