Back to Blog
AI Search13 min read18 September 2026

Voice Search Optimization: What Actually Works for AI Search

Speakable schema only works for US news publishers on Google Home. Here's what actually determines whether an AI voice assistant gets your business right.

By AI Schema Gen Team

Most advice on "voice search optimization" was written for a voice assistant that, as of this year, barely exists anymore. Google is retiring Google Assistant on Android phones and tablets, and shutting down the standalone iOS app, moving everyone to Gemini, an actual language model, not the command-matching system Assistant used to be. Amazon's overhaul, Alexa+, runs on a model agnostic routing system that draws on multiple models, including its own Nova line and Anthropic's Claude, rather than one fixed system. The tips built for the old generation, matching "conversational long-tail keywords," a checklist for smart speakers, one schema type applied broadly, are being applied to assistants that no longer work anything like that.

So the real question isn't "how do I rank for voice search," since there's no separate voice ranking system left to rank in. It's narrower and more useful: when someone asks an AI assistant a spoken question about your business, what actually determines whether it answers correctly? That turns out to hinge on the same handful of things this whole site is about, applied to a specific constraint that only shows up when an answer has to be spoken instead of displayed rather than read from a screen.

The One Schema Type Built for Voice, and Its Real Scope

Schema.org has exactly one property built specifically for audio: speakable. If you've read that structured data helps with voice search, this is almost always what's meant, and almost every explanation of it overstates what it actually covers.

Google's own documentation is specific about the limits: speakable "identifies sections within an article or webpage that are best suited for audio playback using text-to-speech," and it applies only to Article and WebPage types. The restrictions go further than the type limit. Google states plainly that "the speakable property works for users in the U.S. that have Google Home devices set to English, and publishers that publish content in English," and that the Google Assistant uses it "to answer topical news queries on smart speaker devices." That last part is not a throwaway detail: this feature is built for news publishers answering topical news questions, the kind of thing someone asks a smart speaker to catch them up on, not for a local business, an ecommerce store, or a SaaS company answering questions about themselves.

If you're not a news publisher, speakable schema isn't a lever available to you, regardless of how many "voice SEO checklist" articles list it as step one. Marking it up on a page it doesn't apply to does nothing, since Google's systems only consume it in the news-query context it was built for.

Voice Isn't a Separate Ranking Problem, It's a Different Answer Constraint

For a "near me" style query, retrieval happens somewhere else entirely, a maps or directory layer that your own website's schema can't reach, covered in full in how AI actually handles "near me" searches. This post is about a different, narrower situation: someone asks about your business by name, or asks a factual question your site could plausibly answer, and the assistant has to speak a response.

That single constraint, having to produce spoken words instead of a screen, changes three things that don't show up when the same question gets answered as text.

There's no list to read from. A text-based AI Overview or a chat answer can show three options and let you pick. An assistant giving a spoken answer generally commits to one. If your business shares a name with something else, or if there's any real ambiguity about which entity is being asked about, a spoken answer has nowhere to hedge the way a written one can. This is the same disambiguation problem covered in what is an entity in SEO and AI search, just with the stakes raised: a text answer that's slightly off can still show useful context around it. A spoken answer that's wrong is just wrong, out loud, with nothing else on the page to correct it.

The fact has to be looked up, not paraphrased live. A voice assistant answering "what time do you close" out loud isn't reading your homepage's sentence about "swinging by after work," it's producing a specific, spoken value, and it has far less room to hedge with vague phrasing than a written answer does. The gap between a fact a system can read off a declared field and one it has to infer from a sentence, covered in full in structured data vs. unstructured data, matters more here, not less: an inferred fact spoken aloud with total confidence is a specific, immediate way for an assistant to get something wrong in front of the person asking.

The answer has to be short and self-contained. Nobody wants a paragraph read aloud. Assistants compress toward the shortest true statement they can produce, which means the same discipline that makes an FAQ answer quotable in text, covered in how to write FAQs that get cited by AI search, matters just as much for a spoken answer, arguably more, since there's no follow-up link to click if the short version leaves something out.

Why This Is Converging With AI Search Generally, Not Staying Separate

The reason "voice search optimization" is starting to look like a subset of AI search readiness rather than its own discipline is that the assistants themselves have converged, all three major ones, in the same twelve-month stretch. Google is moving Assistant to Gemini. Amazon rebuilt Alexa around a multi-model routing layer. Apple's turn came at WWDC 2026, where it unveiled a rebuilt Siri that, according to reporting on the overhaul, moved from a "command-and-response system" to an "LLM-powered architecture," routing complex queries to cloud infrastructure running Google's Gemini models under a compute-only agreement rather than the older approach of chaining together narrow, task-specific models.

Gemini answering a spoken question on your phone and Gemini answering a typed question in a browser tab are, functionally, the same system reading the same web. Alexa+ making a dinner reservation or answering a factual question runs through the same model-agnostic routing layer either way. There's no separate "voice index" these systems consult that a screen-based query skips. The device changed. The underlying mechanism, an AI model reading whatever it can find about your business and either finding a declared fact or reconstructing one, did not.

That's also why the older conversational-keyword advice has aged badly. It assumed a much simpler system: a smart speaker matching a spoken phrase against an index of pages, where phrasing your content the way people talk out loud gave you an edge in the match. A language model reading your page doesn't need you to have anticipated its exact phrasing. It needs the underlying fact to actually be there, in a form it can trust, which is precisely the entity-grounding argument made in entity grounding: guessing vs. knowing.

One Question, Two Outcomes: A Worked Example

Take an ordinary spoken question: "does Riverside Dental take walk-ins?" Here's how that plays out depending on whether the fact exists as a declared value on the practice's site or only inside a paragraph of prose.

If the fact only exists in a sentence ("we're happy to see walk-in patients when our schedule allows"), an assistant answering out loud has to compress that hedge into something speakable. It might say yes. It might say "usually," or leave the qualifier out entirely, or, if the page is ambiguous about which of two same-named practices it's describing, apply the wrong practice's policy altogether. None of those failures show up anywhere the practice would notice. Nothing flags it, nothing bounces back, just a spoken answer that may or may not match reality, said out loud with the same confident tone whether it happens to be right or wrong.

If the same fact is declared explicitly, through a structured field the practice's schema states directly rather than implies in prose, the assistant has something to read instead of something to interpret. The spoken answer comes out the same way every time, on every assistant, because it was never a guess to begin with. Nothing about the page's visible content needs to change for a visitor. What changes is whether the one-sentence answer a machine gives out loud is standing on a fact or a hedge.

The Properties That Carry the Most Weight for a Spoken Answer

Not every field on an entity profile gets asked about out loud with equal frequency. A handful of properties account for most of the factual questions someone would plausibly speak to an assistant, which makes them the highest-priority ones to have genuinely declared rather than left in prose.

Identity fields (name, address, telephone) matter first, because a spoken answer to almost any other question depends on the assistant having already resolved which specific business is being asked about. Get this wrong or leave it inconsistent across your own site, and every other fact inherits that ambiguity.

openingHoursSpecification answers the single most common spoken business question there is: are you open right now, or when do you close. It's also the property where the gap between a declared fact and an inferred one is most visible, since "closes at 9 most weeknights" and a structured closes: 21:00 value produce very different confidence levels in a spoken answer.

priceRange covers the second most common category: how much does this cost, roughly. A vague answer here ("depends on the project") is often accurate for a human reader browsing a page, but gives an assistant nothing concrete to say out loud.

FAQPage entries for the handful of yes/no or short-factual questions specific to your business (do you take walk-ins, do you ship internationally, is parking available) give an assistant a pre-written, already-concise answer to draw from instead of having to compress a longer explanation down to something speakable on the fly.

sameAs links do quieter work: they're part of how an assistant confirms it has resolved your specific entity rather than a similarly-named one before it commits to speaking an answer about you at all.

Common Mistakes

Marking up speakable schema for a business that isn't a news publisher. As covered above, this does nothing outside the topical-news context Google built it for. It's not harmful, but it's also not the lever most articles present it as.

Chasing "conversational long-tail keywords" as if today's assistants still do simple phrase matching. A language model reading your page isn't rewarding you for guessing how someone might phrase a spoken question. It's working from whatever facts and structure actually exist on the page, phrased however you naturally write it.

Treating voice as a reason to write separate content. The fixes that help a spoken answer, unambiguous entity identity, structured facts instead of buried prose, tight FAQ-style answers, are the same fixes that help every other AI search surface. A second, "voice-optimized" version of your content isn't the goal; a site that's already well-grounded for AI search generally answers spoken questions correctly as a side effect.

Confusing this with the "near me" retrieval problem. If an assistant isn't finding your business at all for a category-plus-location query, that's a directory and maps listing problem, not a voice or schema problem, and it needs the different fix covered in how AI actually handles "near me" searches.

How to Check This on Your Own Site

Pick three or four factual questions someone might genuinely ask out loud about your business: your hours, whether you offer a specific service, your price range, what makes you different from a competitor with a similar name. Ask an AI assistant each one by name, out loud if you have one handy, or through a chat interface set to give a short, spoken-style answer.

For each answer, ask the same question this whole series keeps coming back to: did that trace back to something your site actually declares, or did the assistant reconstruct it from a sentence and get lucky? If you can find the exact fact as a structured, declared value on your site, a spoken answer built on it is standing on solid ground. If you can't find it declared anywhere, the correct-sounding answer you just heard was a guess, and a different assistant, or the same one after an update, isn't guaranteed to guess the same way twice.

It's worth repeating this same handful of questions across more than one assistant, Gemini, Alexa+, and Siri if you have access to all three, rather than testing once and assuming the result generalizes. A fact that's genuinely declared on your site should produce the same answer regardless of which assistant is asked, since all three are reading the same underlying data. A fact that's inconsistent across assistants, or that changes when you ask again a few weeks later, is a reliable sign that whatever answer they gave the first time was reconstructed from context rather than read from a stable, declared source.

Frequently Asked Questions

Is your site ready for AI?

Get a free readiness score in under a minute. No signup, no card.

Run the free check