Boson AI logo
Building voice AI that feels live: a hands-on guide to Higgs Realtime
Aug. 20, 2026
Building voice AI that feels live: a hands-on guide to Higgs Realtime

Real-time voice AI is easy to demo and much harder to build well. The Higgs Realtime API Tutorial is a build-it-yourself guide that constructs a browser voice assistant one layer at a time — live audio, interruptions, out-of-order events, tool calling — with a Git checkpoint and acceptance test at every stage.

Higgs Realtime: cost-efficient, real-time speech-to-speech model
Aug. 4, 2026
Higgs Realtime: cost-efficient, real-time speech-to-speech model

A real-time voice model and API that handles interruptions, adapts mid-sentence, and carries conversations the way people actually talk. It leads major deployed real-time APIs on audio benchmarks, supports 100+ languages, and is compatible with the OpenAI Realtime API.

Higgs Avatar API is now available
Jun. 25, 2026
Higgs Avatar API is now available

Developers can now generate talking-head avatar videos from a still image and either an audio clip or Higgs TTS text input.

Higgs TTS 3: beyond reading, toward real speech for voice AI
Jun. 4, 2026
Higgs TTS 3: beyond reading, toward real speech for voice AI

Higgs TTS 3 is built for voice chat: it speaks, not just reads. It turns model responses into expressive conversational speech across 100 languages, with zero-shot voice cloning and inline control over emotion, style, prosody, pauses, and sound effects.

Meet Higgs Avatar: Real-time avatars for voice agents
May. 13, 2026
Meet Higgs Avatar: Real-time avatars for voice agents

A real-time foundation model that brings human-like digital presence to customer conversations, virtual assistants, training, and interactive experiences

Boson AI Launches Higgs STT 3 Speech-to-Text Model
Mar. 18, 2026
Boson AI Launches Higgs STT 3 Speech-to-Text Model

Today, we are publicly releasing Higgs STT 3, a state-of-the-art Speech-to-Text (STT / ASR) foundation model. It supports 94 languages with sophisticated language detection, advanced sentiment and semantic understanding, and outperforms whisper-v3-large by a large margin on key languages.