Smallest.ai secures $13 million for voice agents that listen and respond in real time

Smallest.ai, founded in late 2024, has raised $13 million in a Series A led by Seligman Ventures, taking its total funding above $21 million. The startup is developing a compact voice model that listens, reasons and speaks simultaneously to make enterprise AI calls feel human.
Why latency breaks the illusion
Large language models typically receive a complete prompt before generating a response. That sequence works in text, but even a short pause can make a telephone conversation feel artificial, particularly when customers expect quick acknowledgment or need to interrupt.
Smallest.ai instead treats speech as a continuous exchange. Its specialized model begins processing information while the caller is still talking and can respond or interject without waiting for a full audio segment.
“While I’m speaking to you, you’re already thinking, and you might interrupt me if I talk for too long.”
A two-model architecture for support calls
The smaller model acts as a real-time intelligence layer for conversations within a defined knowledge domain, with virtually zero response lag. If a query falls outside that domain, the system hands it to a large foundational model and briefly places the caller on hold while it researches the answer.
CEO Sudarshan Kamath expects voice agents to converge on this architecture: a compact model for immediate dialogue and an “offline” LLM for complex tasks. Smallest.ai concentrates on accents, dozens of languages and noisy environments rather than building a general-purpose foundation model.
RingCentral and Truecaller are among its customers. The company competes with ElevenLabs, Cartesia and regional specialists such as Sarvam, but focuses exclusively on real-time enterprise voice agents rather than dubbing or podcast production. The financing follows Gradium’s $100 million voice AI funding round and signals continued investor interest in specialized speech infrastructure.
For businesses, the practical lesson is to assess voice automation as a system, not merely a convincing synthetic voice. Latency, interruption handling, multilingual accuracy, noise tolerance and a reliable fallback path will determine whether an agent can resolve real customer issues without damaging trust.

