Retroactive Grade (July 2026): Mostly correct. By New Year’s the count was six or seven voice-AI unicorns, not ten: ElevenLabs, Sesame, Parloa, Speak, maybe Suno, plus incumbents Uniphore and Cresta. You only reach ten if you count agent companies that happen to talk, like Sierra and Abridge. But ElevenLabs went from $3.3B in January 2025 to $11B in February 2026, which is more than the ten unicorns I predicted combined.

“Voice is the oldest medium,” my friend and former boss Paul Davison would say all the time at Clubhouse. Humans have been using voice to share knowledge and build communities for thousands of years. We hear voices before we’re born, and they’re often the last thing we experience. In 2025, I believe we’ll see voice return as our primary interface with technology, powered by AI.

Why voice matters now

The paradox of voice has always been its accessibility versus its efficiency. Speaking comes naturally to us, but voice carries less information density than text or images. This trade-off has historically limited voice interfaces, forcing us to choose between convenience and capability.

But AI is changing this equation. Large language models can now transform our natural speech into information-dense formats while preserving the nuance and context that make voice communication so powerful. This breakthrough enables a new generation of tools that combine the ergonomic benefits of voice with the precision of digital interfaces.

The technical breakthrough

What’s different now is the emergence of true multimodal AI models. Previously, voice interfaces required a cumbersome chain of transformations: speech-to-text, text processing, and text-to-speech. Each step added latency and reduced quality. Modern AI models can process voice natively as input and output, enabling real-time, natural conversations.

This architectural shift, combined with widespread WebRTC adoption and improved voice quality, has finally pushed us past the uncanny valley. Even subtle elements like “ums” and “ahs” that once betrayed artificial speech are now handled naturally.

The next wave of voice applications

We’re already seeing promising applications emerge:

  • Lindy: Automating survey and feedback collection through voice
  • Otter: Real-time transcription and meeting intelligence
  • SuperWhisper: Privacy-focused, local voice processing
  • Drillbit: Automated voice agent for services businesses
  • ChatGPT: Launched Voice mode and, just today, 1-800-CHATGPT

But these are just the beginning. I predict we’ll see at least 10 voice-AI unicorns emerge in 2025, primarily at the application layer. I expect most of them to be voice-enabled agents that transform existing industries, not standalone apps.

Think about every profession with “agent” in the title: insurance agents, travel agents, real estate agents. Each is a complex workflow of gathering information, processing it, and taking action. Voice AI is perfectly positioned to augment or reimagine these roles.

The path to a billion users

Voice AI’s most profound impact won’t be technological. It will be in bringing AI to the next billion users. My aunt, in her 70s, recently had a natural conversation in Serbian with an AI assistant. That moment of genuine connection and delight showed me how voice can make AI accessible to people who might never use a chat interface.

The key is that we no longer need artificial constraints like phone trees or menu systems. Voice AI can understand context, maintain conversation flow, and take actions, all through the most natural interface humans have ever known.

Building the voice-first future

For developers and entrepreneurs looking to build in this space, focus on workflows where:

  1. Repetitive voice interactions are already happening
  2. Text-based interfaces feel cumbersome or unnatural
  3. Information needs to be gathered, transformed, and acted upon

The most successful applications will reimagine entire processes around voice-first interactions rather than use voice as a feature.

Looking ahead

As we enter 2025, voice AI is a return to our most fundamental form of communication, augmented by AI’s ability to understand, process, and act. The companies that succeed will create experiences that feel as natural as having a conversation with a friend.

The future of human-computer interaction is computers finally speaking our language.