8.7.2026

Voice is the Gateway Drug to Next-Gen AI

Jordan Crook

“By 2026, the input box as the main interface for AI applications will disappear.” That’s what A16Z’s Marc Andrusko had to say in December of last year. Around then, Wispr’s valuation was a third of what it is today. 

The input box is not gone, but that interface is no longer the core primitive that builders use to design their products. Something fundamental changed between the failed voice interfaces of 2018 and the investment consensus around voice in 2025 and… spoiler alert: it has nothing to do with transcription accuracy. 

Computers could always hear us. What changed is that the mess of spoken word, rather than written text, is suddenly useful. The AI on the other side of that transcription can function as a relatively competent intermediary between sloppy, stream-of-consciousness input and some more refined output, either as instructions for more AI or content to be consumed by other humans. 

As such, voice as an interface is likely to unlock a new wave of AI adoption and, as a consequence, a ripple of unforeseen use cases and second order effects. It’s the gateway drug. 

The very sloppiness that prevented adoption, that eroded its efficiency, is now set to anoint voice as the most context-rich interface for AI native products. Perhaps more surprising, that adoption may not come from the demographics you’d expect. 

Cognitive Slop = Context 

For the first time in the history of computing, sloppy spoken input may be more valuable than precise written language (ie, structured data). Thinking is inherently messy and, rather than trimming that away in the writing process, voice encourages us to leave it in as context. 

Think about it. The keyboard forces you to pre-compress your thinking into clean instructions or outlines. Even in a brainstorming session, where you might be jotting down your thoughts as they come, it’s still more structured and sparse than what you might be able to say aloud. 

If the spectrum we’re discussing has thought on one end, and written language on the other, voice is much closer to the former than the latter. It more likely includes the process by which a crystallized idea is formed. Curiosity A leads to speculative answer B, which gets anecdotal support from evidence C, which surfaces some applicable insight D. 

When you type, you skip straight to D, because that’s the valuable thing. But those earlier steps contain the context that makes D legible. It’s not just what you think, or even why you think that… it’s how you think. 

If you’re wondering whether the context around voice is valuable, look no further than a little company called Granola. In three years, the company has risen to a $1.5B valuation and, without saying too much, let’s just say that the growth of the business is reflective of that pace. It did this flanked by venture-backed startups on one side (Otter, Fireflies, Fathom) and incumbent platforms on the other (Microsoft Teams, Google Meet, Zoom). 

Then there’s Plaud, which was doing $250 million in annualized revenue in 2025 (I hear that number has nearly doubled, but that’s just venture rumor – I can’t independently verify). 

It’s not surprising to find that these products are doing so well. It’s simply easier and more efficient to use voice in most cases (spoken WPM is around 165 compared to typed WPM which is around 65).  

Wispr’s adoption strategy centers on eliminating the need to correct output, because the adoption barrier is built out of years of small transcription failures. Users who stick with Wispr Flow for six months end up typing 72% of their characters by voice instead of keyboard. That’s not preference, it’s wholesale behavioral replacement. The gateway drug worked.

But two things are happening simultaneously – the context afforded to these systems via voice is making AI more valuable to its users at the exact time that users are offering up more and more context because voice is convenient. 

The latter feeds the former, creating a loop of engagement that improves naturally over time. 

The Ring of Trust

Some of us talk to ourselves (very guilty!), but the essence of spoken language is multiplayer. We speak to be heard, the presumption of an audience ever present, and more often than not we expect some response. It’s a loop.

When working alone, usually writing or clicking rather than speaking, every thought has to pass a threshold test before we commit to it: Are my hands free? Is it worth pulling out my phone? Voice nullifies that threshold, and creates an Other on the other side of that content, that can, in the right circumstances, build trust and increase engagement. 

The “no bad ideas” mantra—the premise that even the most trivial or half-baked thought accrues more value by being externalized—is the engine of collaboration, because it treats raw, unfiltered input as fuel for a system that turns the messy overlap of perspectives into insights greater than their sum. 

Humans compute far more than language when they hear spoken words. They’re listening for emotional cues in tone and pace; that voice can activate visual and sensory parts of the brain with descriptive language. The brain waves of a storyteller and a listener have been shown to neurally pair. There’s a subconscious emotional response to voice that text simply doesn’t trigger—the sound of a mother’s voice on the phone releases the same oxytocin as a hug. 

ElevenLabs has consolidated quite a bit of the value in the ‘sounds human’ category of AI. The voice models they’ve developed create the illusion of human behavior and developers are coming up with more and more creative ways to engender even more trust. 

For example, Sandbar (a Betaworks portfolio company releasing an AI ring in the fall) has used Eleven Labs voice models to clone the voice of a user, lowering it by a half step. It’s meant to create a closed-loop environment that feels like the voice in your own head rather than a conversation with an Other. 

But the valuation of ElevenLabs, a lofty $11 billion, is actually only reflective of a small slice of the trust pie for AI and agents. 

The performance of humanness is not enough for real trust. There must also be the performance of competence. That is achieved through actual competence (accurate transcription which will lay the foundation for an accurate output) and statefulness. 

Statefulness is one of the distinctions enjoyed by agents, rather than AI applications. It means that the AI has some sense of memory, of context sharing across sessions, that compounds its utility over time. Statefulness is not the only required criteria for agents – most believe they should also be able to act on your behalf to complete a task, often working through the problem on their own. 

Regardless, voice paired with the competence of agents is a recipe for increased trust for users, despite losing the traceability available via text. 

The trust claim has limits. ICTworks documented voice interface failures in African marketplaces, where code-mixing and environmental noise make transcription unreliable enough that voice may raise digital barriers rather than lower them. I was recently told by a Korean LP that Granola is a non-starter until it does a better job with Korean. 

Language-fluency support is a related counterpoint that will fade over time but matters now. The more structural problem is accountability: a screen gives you receipts a voice exchange doesn’t, which gets painful when a model hallucinates. Voice makes it harder to hide who you are; text makes it harder to hide what was said. Those are two different accountability channels, and voice trades one to unlock the other.

On balance the benefits win, and tools will emerge to close the traceability gap. But the trust threshold isn’t crossed through accuracy alone—it’s crossed through repeated small successes that show the system can execute on delegation, not just follow instructions. 

Trust Who? 

This argument will hold for the usual suspects: early adopters and millennials who are less averse to this technology. But voice may be the ever critical lynchpin that unlocks two much more complicated demographics: Boomers and Gen Z. 

Boomers have historically struggled with technology adoption not because they don’t like it, but because they were raised in a voice-native world. In the 70s, when boomers were in their formative years, the average household had one television screen and that’s it. 

Trustworthy voice agents may unlock an interest from that generation that, despite their age, would be a meaningful audience for just about any company. They represent one-fifth of the population here in the U.S. 

Gen Z, conversely, has no issue with technology adoption but is in their ‘save the world’ phase. They see AI as soulless, a threat to their livelihood, and, perhaps more importantly, a massive threat to our planetary health. Voice may not overcome all that, but it does come to them in their native language. Gen Z are the Snap generation, the TikTok generation, and are in many ways, post-text. Whereas Millennials grew up mostly reading the internet, Gen Z had full high-resolution images and video fresh out of the womb. 

Perhaps voice, and particularly voice-centric AI agents, can meet these two challenging demographics where they’re at. 

The Agent Economy 

So what does multiplication look like once someone crosses that threshold? The VC consensus is that agents will resemble employees in 2026, because production requires accountability and workflows have to stabilize before AI leverage compounds. 

Voice isn’t where you’d start that stabilization—tuning agents to be personalized, focused, and reliable is a heavy lift best done with text and traceability first. But the reason Granola is growing as fast as it is comes down to a simple fact: the conversations employees have with one another, via voice, are how businesses actually move. It’s not the PRD or the strategy doc or the P&L that pushes a company into the next quarter—it’s the ideas, delegation, decision-making, and organization of minds that happen in meetings. Once voice is unlocked inside agentic AI at work, agents can participate in that momentum in a way text-based AI can’t.

The second-order effects are what matter. Once voice is unlocked—and with it trust, engagement, and imagination on the part of users—the new economy starts to take shape. The biggest companies and products of the AI era may turn out to be responses to problems that don’t fully exist yet, and getting there requires AI to be adopted by the masses who can unleash their full imagination onto it. Voice is the gateway drug. Not because it transcribes better, but because it captures the thought you’d never bother to type. And for the first time, something on the other side is smart enough to know what to do with it.

REad more

Back to writings