Alex Smola, the machine learning researcher who co-founded Boson AI in 2023, is preparing to release the company’s first speech-to-speech model called Higgs RealTime. His thesis is straightforward: voice AI models have finally crossed the quality threshold needed to fundamentally change how humans interact with machines.
What Boson AI is building
Boson AI, headquartered in Santa Clara, California, was co-founded by Smola and Mu Li with a focus on creating foundation models for natural, responsive voice interactions.
The company’s Higgs TTS 3, released on June 4, 2026, supports expressive conversational speech in over 100 languages. It features zero-shot voice cloning, meaning it can replicate a voice without needing extensive training samples, and inline emotion control that lets developers dial up or down the emotional register of generated speech in real time.
Alongside the text-to-speech model, Boson AI launched its Higgs Avatar API in June 2026. That tool generates real-time talking-head video from a single still image paired with audio or text inputs.
The earlier Higgs TTS 2 models were open-sourced on Hugging Face back in May 2025, having been trained on over 10 million hours of audio data. That open-source move helped build developer goodwill and community engagement, a strategy the company reinforced by co-hosting a Higgs Audio Hackathon with Eigen AI in Mountain View from March 20-22, 2026.
The new Higgs RealTime model represents the next logical step: moving from text-to-speech into full speech-to-speech processing, eliminating the intermediate text conversion step that introduces latency and strips away vocal nuance.
The competitive landscape and what investors should watch
Boson AI isn’t operating in a vacuum. Major players like OpenAI, Google, and ElevenLabs have all made significant moves in voice AI. OpenAI’s GPT-4o demonstrated real-time voice capabilities that captured public imagination, while ElevenLabs has built a substantial business around voice synthesis and cloning.
What differentiates Smola’s approach is the combination of his deep academic pedigree in machine learning, the decision to open-source earlier model versions to build ecosystem adoption, and the specific focus on production-grade, low-latency applications. Training on over 10 million hours of data for just the TTS 2 model alone signals the kind of compute investment typically associated with well-funded operations.
Boson AI has no known direct connections to blockchain or digital assets at this stage. But the infrastructure it is building, low-latency voice processing, emotion-aware speech generation, real-time avatar creation, represents exactly the kind of tooling that Web3 applications will eventually need to deliver user experiences that don’t feel like a step backward from Web2.
Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our Editorial Policy.

2 hours ago
9









English (US) ·