.jpg)
For 10 years, every car launch promised voice would change the cabin. Voice that replaces the buttons. Voice that holds a conversation. Voice that finally works.
It almost never did. The honest answer was always the same: it worked if you said “navigate to Schlossstrasse 7, Stuttgart” exactly like that, and it fell over the moment you said anything a human would say.
That part is over. AI voice assistants are now shipping in production vehicles, and the conversational layer is genuinely good enough. Chinese new-energy brands are putting multi-turn voice at the center of their interior story. European OEMs are launching with LLM-based assistants as the headline cabin feature. The platform players have re-pointed their automotive voice efforts at large language models.
So, voice in cars works. Good. We can stop treating that as the interesting part of the story.
Because solving it exposed the actual problem.
What changed, and what didn’t
Voice quality used to be the bottleneck. Type “voice in cars 2018” into any analyst archive and you’ll find 500 articles about why the recognition stack couldn’t handle accents, couldn’t hold context, couldn’t deal with two people in a cabin, couldn’t survive road noise.
That stack of complaints is largely solved. The current generation handles fuzzy phrasing, multi-turn dialogue, accent diversity and ambient noise at a quality the 2018 stack couldn’t approach.
Solving voice quality moves the constraint somewhere new. Specifically, to the thing the voice is talking to.
The Tier 1 commodity tier
Most 2027 voice launches will use that newly capable voice layer to do the same things voice has always done in cars. Navigation. Climate. Calls. Vehicle settings. Read me the messages.
I’m convinced that’s fine, and also a commodity. Tier 1 suppliers and platform players already cover most of it, and the rest is on a 12-to-18-month catch-up curve. By the next vehicle generation, no OEM differentiates on “voice can change the temperature.” Every OEM can.
Voice as a button replacement earns its keep. It doesn’t earn the brand anything.
The half that earns the brand something is voice for content. And almost nobody is set up to deliver that half.
What voice for content actually has to do
Three jobs, none of which a wake word with a search box solves.
Discovery. The driver says, “find me something light, 20 minutes, the kids are in the back.” A keyword system hears “comedy.” A real discovery layer turns fuzzy intent into one specific result, in a moving cabin, without scrolling into a grid. The translation work happens at the content layer, not at the wake word.
Interaction. One query rarely lands the right thing. “Something else.” “Shorter.” “No, the new season.” The voice session must stay warm across a 20-minute drive, remember what was just rejected, and not make the driver start over every prompt. A single-shot voice search is a faster version of a broken experience.
Personalization. Same driver, same car, very different drives. School run on Tuesday morning is a different context to Friday night with the partner in the passenger seat. What voice returns must be conditioned on who’s listening, what time it is, what was already watched, what the household has paid for. Personalization is what makes the cabin feel like it knows you.
These three jobs share a common dependency. They all live in the content layer. The voice layer can’t fake them.
Why most launches will still feel broken
Here’s the architecture problem nobody likes to talk about.
In most cabins today, the content “experience” is a screen of partner tiles. Each tile has its own catalogue, its own search, its own recommendation logic, its own personalization signals it won’t share. The screen calls itself an experience. It’s a launcher with a paint job.
Put a fluent voice assistant in front of that launcher, and what comes back is: “I found 4 results across 11 apps, here they are.” Which is the screen, read aloud.
The driver doesn’t want 4 results across 11 apps. The driver wants 1. The right 1, for this trip, for this person, for the 20 minutes available before the school gate.
That answer doesn’t come from the voice layer. It comes from a content layer built to be queryable, conversational, and personalized across the whole catalogue, not per app.
If that layer isn’t there, the voice assistant has nothing useful to say.
The brand argument
There’s a strategic reading of this that matters for anyone running OEM product or in-cabin experience.
Nav and climate are commodities within 18 months. Voice for the cabin’s mechanical functions is a feature war that levels out across the industry on the same release cadence as every other interior tech war.
Content experience is different. The OEM whose car lands the right show in the right moment for the right driver builds a kind of attachment that doesn’t transfer to the next vehicle generation. It travels with the customer.
Vehicle ownership cycles are getting shorter. The driver who walks out at the end of a 3-year lease takes their music habits, their kids’ show histories, their morning commute patterns with them. The brand that picked that thread up cleanly across a few model years has a customer reason to come back that the spec sheet never shows. The brand that handed the content conversation to a third-party assistant doesn’t.
One more piece of brand-savvy OEMs are starting to get right. Keep the wake word yours. The driver should still be talking to the brand, whatever the assistant is being asked for. A brand-owned wake word for navigation and a generic third-party assistant for content is a customer journey that hands the relationship away halfway through the trip.
The wake word stays at the front door. What sits behind it is a stack the driver never sees. That distinction is where the competitive position lives.
What “good” looks like by 2027
A short version, for anyone working on this now.
A voice layer that understands intent rather than matches keywords. A content layer underneath that’s unified across sources, queryable across the catalogue, and ready to hold a conversation about its own content. Personalization that knows who’s in which seat, what time it is, what got watched yesterday, what the household has paid for, what’s available on this trip’s connectivity. All of them are white-labeled. The driver hears the OEM. The OEM owns the relationship. The platform underneath does the work.
That’s a real bet. It’s also a content-layer problem dressed up as a voice problem, which is why a lot of 2026 budgets will buy the wrong half of it.
If you’re working through this
We’ve spent years building this kind of conversational content layer on the TV side. We’re bringing it into the cabin now, white-labeled, behind the OEM’s wake word. Discovery, multi-turn interaction, personalization, on top of one unified content layer.
If voice is on your 2027 roadmap and the content-layer question is starting to surface in your reviews, get in touch. Happy to compare notes.
Felix Walter is Chief Growth Officer at 3SS, leading the 3Ready Automotive business unit.


