The 2024-2025 Voice-AI Convergence: A Qualitative Timeline
Voice search before 2023 was mostly keyword matching with text-to-speech output. Between 2023 and 2025, that quietly stopped being true.
Four landmark moments mark the convergence — all verifiable, all within an 18-month window:
- September 2023. OpenAI released ChatGPT Voice mode for paid users. It became free for all users in 2024.
- June 2024 (WWDC). Apple announced Apple Intelligence and Siri's integration with ChatGPT. The rollout spanned 2024-2025.
- 2024-2025. Google began transitioning Google Assistant experiences to Gemini-powered models on Android and Pixel.
- February 2025. Amazon announced Alexa Plus, the generative-AI version of Alexa.
For marketers, the practical consequence is simple. The voice surface is now an LLM surface. Content that wins citations from ChatGPT, Perplexity, and Google AI Overviews is the same content that gets read aloud through voice assistants. Teams that want a partner to earn those citations engage our ai seo services, and our comparison of ai content optimization tools covers the authoring stack that produces voice-ready structure.
ChatGPT Voice + Siri Integration (June 2024 WWDC)
In short
ChatGPT Voice mode launched in September 2023 and became free for all users in 2024. Apple announced Siri's integration with ChatGPT at WWDC in June 2024 as part of Apple Intelligence, with phased rollout across 2024-2025.
ChatGPT Voice is a real-time spoken-conversation mode that lets users talk to the model directly. It launched for paid users in September 2023 and was extended free to all users during 2024.
The Apple side of the story is equally specific. At WWDC in June 2024, Apple announced Apple Intelligence — its on-device and cloud AI stack — together with a partnership that lets Siri hand off certain queries to ChatGPT, with user permission.
Two practical implications for content owners:
- The same retrieval applies. When ChatGPT answers a voice query, it can pull from ChatGPT Search results — which use the Bing index plus OpenAI's OAI-SearchBot. Voice optimization is ChatGPT optimization.
- Hand-offs change attribution. A user might begin with Siri and end up answered by ChatGPT. Your analytics will not show "Siri referral." It shows up as part of the broader AI search shift.
None of this requires new optimization tactics. It is the same playbook: entity clarity, structured data, FAQ blocks, citation-rich writing, and freshness.
Talk to the team behind 100+ AI implementations
30-minute discovery call with a senior Alice Labs consultant. No slide deck, no sales pitch — just a scoping conversation.
Book a Discovery CallAmazon Alexa Plus and Google Assistant -> Gemini
In short
Amazon announced Alexa Plus, a generative-AI version of Alexa, in February 2025. Google has been transitioning Google Assistant experiences to Gemini-powered models across 2024-2025. Both moves replace older speech-pipeline assistants with LLM-driven ones.
Amazon announced Alexa Plus in February 2025. It is described as a generative-AI version of Alexa that can handle more conversational, multi-turn requests than the original product.
Google has been moving Google Assistant toward Gemini across 2024-2025. On Android and Pixel devices, Gemini increasingly replaces the legacy Assistant for voice interactions and broader help tasks.
The pattern is consistent across all four major assistants:
- Old core, new core. Each assistant kept its brand and surface (the wake word, the device fleet) and swapped the underlying model for an LLM.
- Web answers, not skills. The early Alexa-skills / Actions-on-Google ecosystem mattered less. LLM-driven assistants answer from web content and tool calls.
- Citations vary by surface. Some voice surfaces read out a single answer with no on-screen citations; others show source links on a paired screen.
For marketers, this changes the question. It is no longer "do I build an Alexa skill?" It is "is my web content the source the LLM picks up when a user asks aloud?"
Optimization: Schema.org Speakable Property
In short
Schema.org Speakable is a property designed to mark which parts of a page are suitable for voice readout. It provides explicit signals to assistants and AI systems but adoption is still limited — making it cheap upside for early movers.
Schema.org Speakable is a property — typically used on Article or WebPage — that tells a search system which sections of the page are most appropriate for being read aloud.
It works through CSS selectors or XPath that point to the voice-friendly sections of the page. The intent is explicit: this is the bit you would read aloud.
Three things to know before deploying it:
- It is a signal, not a guarantee. No major assistant has publicly committed to honoring Speakable as the sole determinant of what gets read aloud.
- Adoption is limited. Most sites do not implement Speakable. That is precisely what makes it interesting — early movers stand out.
- Pair it with quick-answer blocks. Speakable works best when the marked-up section is already a self-contained answer of two to three sentences.
For an Alice Labs client implementation, the rule of thumb is simple. If your page already has a quick-answer block, point Speakable at it. The marginal cost is minutes; the downside is zero.
Want to know if voice assistants would read your page?
We audit your top pages for quick-answer blocks, Schema.org Speakable, and FAQ structure — the same signals voice surfaces and LLMs both prefer.
Request a voice + AI readiness auditQuick Answer Text as the Voice Readout Candidate
In short
Voice assistants often read aloud the same text that wins featured snippets and quick-answer boxes in traditional search. A 40-50 word self-contained answer at the top of a page is the strongest dual-purpose voice and AI optimization element.
Across both traditional and AI search, one content element does disproportionate work: a self-contained, two-to-three sentence answer placed at the top of the page.
That same block is what voice assistants reach for. When a Siri, Alexa, or Google Assistant query maps to your page, the readout tends to come from the most extractable text — usually a quick answer or FAQ entry.
What makes a quick-answer block voice-ready:
- 40 to 55 words. Short enough to read aloud in one breath.
- Self-contained. Doesn't require the surrounding paragraph to make sense.
- Direct answer first. Lead with the answer, then add nuance.
- No "click here" or "see below". Cross-references to other parts of the page break voice readout.
The Aggarwal et al. 2024 GEO paper (arXiv:2311.09735) found that content with inline citations, specific statistics, and authoritative language saw up to 40% improved visibility in generative engines. The same writing style transfers cleanly into voice readouts.
Strategic Implications: Should Brands Optimize Specifically for Voice?
In short
In 2026, dedicated 'voice search optimization' as a separate budget line is rarely justified. Voice has converged with LLM answers — the same content fundamentals (entity clarity, FAQ blocks, structured data, freshness) drive both. Treat voice as a downstream surface of LLMO, not a parallel program.
For a decade, agencies pitched "voice search optimization" as a standalone deliverable. In 2026, that pitch is mostly outdated.
The practical answer for most brands:
- Don't build a separate voice content program.The same FAQ pages, quick-answer blocks, and Schema.org markup that win AI citations also win voice readouts.
- Do add Speakable to high-traffic pillar pages.Low cost, no downside, signals voice-readiness explicitly.
- Do write quick-answer blocks in spoken-friendly form.40-55 words, direct answer first, no on-page cross-references.
- Don't waste budget on Alexa skills or custom Actions.The LLM-driven assistants pull from the open web, not skill ecosystems.
For Alice Labs clients running LLMO programs, voice readouts are a "free" downstream effect of the citation work. We don't bill voice optimization separately. We make the existing pages voice-friendly and let the LLM-mediated voice surfaces benefit.
The exception: if voice is core to your product (in-car, smart home, accessibility), invest in voice UX research, not voice SEO. Those are different disciplines.
About the Authors & Reviewers

CEO & Co-Founder, Alice Labs
CEO & Co-Founder at Alice Labs. Author of 7 research reports on AI adoption, governance and labor markets cited across EU, OECD and US benchmarks.
- 8+ years in AI strategy & implementation
- Top-5 AI Speaker, Sweden (Mindley 2025)
- 100+ enterprise AI engagements

Co-Founder, Alice Labs
Co-Founder at Alice Labs. Builds AI automation, agent workflows and integration systems that hold up in real business operations.
- AI automation & agent systems lead
- Workflow design across 100+ deployments
- Specialist in RAG, integrations & APIs
Frequently Asked Questions
When did ChatGPT Voice launch?
ChatGPT Voice mode launched in September 2023 for paid ChatGPT Plus users. OpenAI extended Voice mode to free users during 2024, making spoken conversation with the model broadly available.
What is Apple Intelligence and how does it relate to Siri?
Apple Intelligence is Apple's AI stack, announced at WWDC in June 2024. It includes on-device and cloud models plus a partnership with OpenAI that lets Siri hand off certain queries to ChatGPT with user permission. The rollout spanned 2024-2025.
What is Amazon Alexa Plus?
Amazon announced Alexa Plus in February 2025 as a generative-AI version of Alexa, designed for more conversational, multi-turn interactions than the original assistant. It layers an LLM on top of the existing Alexa device fleet.
Is Google Assistant being replaced by Gemini?
Across 2024-2025, Google has been transitioning Google Assistant experiences toward Gemini-powered models, especially on Android and Pixel devices. Gemini increasingly handles voice interactions that were previously routed through the legacy Assistant.
What is Schema.org Speakable?
Schema.org Speakable is a property that marks which sections of a page are suitable for being read aloud by voice assistants and AI systems. It uses CSS selectors or XPath to point to specific page sections. Adoption is still limited, which makes it relatively easy to differentiate with.
Do I need a separate voice search SEO program in 2026?
For most brands, no. Voice has converged with LLM answers — the same content fundamentals (entity clarity, FAQ blocks, Schema.org structured data, quick-answer blocks, freshness) drive both. Treat voice as a downstream surface of your LLMO program rather than a parallel budget line.
Should I still build Alexa skills or Google Assistant Actions?
Generally no, unless you have a specific in-app reason. With Alexa Plus and Gemini-driven Assistant moving to LLM-mediated answers, the older skill and Actions ecosystems have less leverage. Web content and structured data return more visibility than skill development for most use cases.
How long are voice queries compared to typed ones?
Voice queries are typically longer and more conversational than typed queries. Users tend to ask full natural-language questions and follow up more readily. This makes FAQ pages and quick-answer blocks especially valuable — they match the question shape voice users actually use.
AI Search Users 2026: Adoption Data, Milestones & Behavior
Next in AI Search & LLMOAI Overview Trigger Rate: Which Queries Show AI Answers?
Further reading
- Schema.org — Speakable property documentation· schema.org
- OpenAI — official site (ChatGPT Voice mode announcements)· openai.com
- Apple — Apple Intelligence overview· apple.com
- Aggarwal et al. — GEO: Generative Engine Optimization (arXiv:2311.09735)· arxiv.org
- SparkToro — 2024 zero-click search study· sparktoro.com
Related reading
AI Search Optimization: Complete Guide for 2026
Full playbook covering ChatGPT, Perplexity, Claude, and Google AI Overviews — voice surfaces sit downstream of this work.
14 min glossaryWhat Is LLMO? Large Language Model Optimization Explained
Glossary definition of LLMO — the overarching discipline behind voice and AI search visibility.
7 min guideSchema.org for AI: Which Types Actually Matter
Practical guide to high-leverage Schema.org types for AI search, including Speakable for voice readouts.
9 minSources
- OpenAI — ChatGPT Voice mode launch (September 2023; free for all users 2024)(accessed 2026-05-06)
- Apple — Apple Intelligence and Siri + ChatGPT integration (announced WWDC, June 2024; rolled out 2024-2025)(accessed 2026-05-06)
- Amazon — Alexa Plus announcement (February 2025)(accessed 2026-05-06)
- Google — Google Assistant transitioning to Gemini-powered experiences (2024-2025)(accessed 2026-05-06)
- Schema.org — Speakable property documentation(accessed 2026-05-06)
- Aggarwal et al. — GEO: Generative Engine Optimization (arXiv:2311.09735, 2024)(accessed 2026-05-06)
- Jeremy Howard / Answer.AI — llms.txt proposal (September 2024)(accessed 2026-05-06)
- SparkToro / Datos — 2024 zero-click search analysis (~60% of Google searches end without a click)(accessed 2026-05-06)
- Alice Labs — LLMO Citation Benchmark (100+ Nordic implementations)(accessed 2026-05-06)
Next scheduled review: