Xtrusio AEO/GEO Audit

50 answers name ElevenLabs.

Seven name Fish Audio. None put it first.

Across 20 buyer-intent questions tested on ChatGPT, Claude and Gemini, Fish Audio is named in 7 of 60 responses (11.7%) and records zero first placements. That visibility is low relative to the company’s reported scale — more than 8 million users and $21 million in ARR — while ElevenLabs appears in 50 responses, Deepgram and Cartesia in 32 each, and Inworld in 28. The gap measured here is AI recommendation visibility, not product capability.

The findings below come from Xtrusio, an AI visibility audit system built specifically for B2B buyer-intent testing. Every result was produced by running 20 real prospect queries across three generative AI platforms and recording which vendors were named.

Scored against the voice AI and text-to-speech infrastructure category, where buyers compare vendors on latency, deployment model, language coverage and per-character cost.

September 2026
20 Queries • 3 Platforms
Fish Audio
20%
Claude
4 of 20 queries
0× #1 RANKINGS
10%
Gemini
2 of 20 queries
0× #1 RANKINGS
5%
ChatGPT
1 of 20 queries
LOWEST VISIBILITY
Proof Exists. Visibility Does Not.

Fish Audio already has measurable proof. It is not consistently surfacing in these answers.

Fish Audio publishes benchmark results, pricing, technical documentation and open-weight model information. Competitors including ElevenLabs, Deepgram and Cartesia nevertheless surfaced far more often across the 60 responses tested — 50, 32 and 32 respectively, against Fish Audio’s 7. This audit establishes that difference in AI visibility. It does not establish why it occurs. A separate source-authority and retrieval analysis would be needed to determine whether third-party coverage, enterprise documentation, model association or another factor contributes to the gap.

Section 2

Platform Scorecard

Where Fish Audio was named, and where it was not

Fish Audio Mention Rate by Platform
Claude
20%
Gemini
10%
ChatGPT
5%
Vendor Comparison — Mentions Across All 60 Responses
ElevenLabs
83%
Deepgram
53%
Cartesia
53%
Inworld
47%
Azure Speech
30%
Google
27%
Fish Audio
12%

Scored across all 60 responses. Deepgram and Cartesia both finish on 32, ahead of Inworld — the infrastructure and compliance questions carry them. Mention frequency is a visibility measure, not a quality ranking.

Claude Is Fish Audio’s Strongest Platform In This Audit — 20%
Claude named Fish Audio on 4 of 20 queries — per-byte pricing, cloning reference length, long-form narration and open weights. All four sit in the same narrow band, and none was a first placement.
ChatGPT: Lowest Visibility — 1 of 20
The single ChatGPT appearance is on the open-weights question, at rank 6, where the model names Fish Audio and correctly flags its Research License and the need for separate commercial licensing before recommending a more permissively licensed alternative. It is a qualified mention, counted toward visibility but not an endorsement.
Section 3

AI Visibility Leaderboard

Who the AI answers name when buyers ask about voice infrastructure

Platform-by-Platform Breakdown
Claude
4/20
Fish Audio named
Gemini
2/20
Fish Audio named
ChatGPT
1/20
Fish Audio named
Fish Audio
1
4
2
7
ElevenLabs
17
19
14
50
Deepgram
11
12
9
32
Cartesia
9
11
12
32
Inworld
13
11
4
28
Azure Speech
5
9
4
18
Google
6
8
2
16
ChatGPT
Claude
Gemini

All bars scaled to the highest-mentioned vendor (50 of 60).

Share of 60 Responses
ElevenLabs: 50 of 60 responses (83%) Deepgram: 32 of 60 responses (53%) Fish Audio: 7 of 60 responses (11.7%)
12%
Fish Audio
ElevenLabs50
Deepgram32
Fish Audio7
Mention Intensity Heatmap
ChatGPT
Claude
Gemini
Total
#1s
ElevenLabs
17
19
14
50
18
Deepgram
11
12
9
32
11
Cartesia
9
11
12
32
3
Inworld
13
11
4
28
8
Azure Speech
5
9
4
18
4
Google
6
8
2
16
1
Fish Audio
1
4
2
7
0

First placements counted as the first vendor named in an answer. The six competitors above account for 45 of the 60 first placements; the remaining 15 went to vendors outside the primary comparison set.

Zero First Positions, Every Platform, Every Session
Across 60 responses Fish Audio was never the first vendor named. ElevenLabs took 18 first placements, Deepgram 11 and Inworld 8. All counts are reproducible from the heatmap above. A mention that arrives fourth or sixth in a list carries very different weight to a buyer than the name that opens the answer.
Deepgram Is The Comparison That Matters
Deepgram appears in 32 of 60 responses and takes 11 first positions, second only to ElevenLabs, winning the infrastructure and compliance questions outright. Cartesia matches it on 32 mentions. Both are consistently associated with published latency, compliance and deployment evidence in the tested answers. Establishing whether that association is causal would require a separate source-authority and retrieval audit.
Section 4

AI Positioning Audit

20 buyer-intent queries — click any row to see the exact question

Each query was written from the perspective of a real decision-maker researching voice AI during discovery, before any vendor shortlist exists. These are the buyers whose AI answers determine whether Fish Audio is considered at all.

Target Buyer Sector VP-level AI leaders at regulated healthcare and benefits companies, and production leaders at audiobook publishers
Buyer Personas
3 verified LinkedIn decision-makers behind these 20 queries
Name, title and company are verified from public profiles. Pain points are buyer scenarios constructed from role and company context, not statements made by these individuals.
SL
VP, Head of AI — AI Strategy & Execution
HealthEquity • Health Benefits • Alpharetta, GA
10queries
Pain Points
Member audio cannot leave the environment, so any vendor without an in-VPC or air-gapped option is disqualified before the demo. The current support voice reads flat on emotionally loaded claims calls. Procurement wants one fixed annual number, an uptime commitment and documented data residency — not a credit balance that moves with usage.
“self-hosted voice AI HIPAA BAA”“TTS latency under production load”
Queries 1, 2, 4–7, 12, 15, 17, 19Profile
ED
Head of Production
Podium Audio • Audiobook Publishing • Toronto
5queries
Pain Points
Narration cost per finished hour has become the largest line in the production budget across hundreds of titles a year. Invented fantasy and science fiction names get mispronounced with no way to lock a correction across a whole book. The Asian fantasy catalogue needs Japanese and Korean that sounds native, not an English model reading a translation.
“AI narration cost per character”“custom pronunciation dictionary audiobook”
Queries 3, 10, 11, 13, 18Profile
TF
Senior Producer
Simon & Schuster Audio • Publishing • New York
5queries
Pain Points
An Audie-winning producer who directs human voice talent in the studio, so AI narration has to take direction line by line rather than reading everything at one level. Full-cast scenes need genuine multi-speaker output instead of stitched exports. Narrator consent and rights have to be airtight before any voice is cloned.
“AI voice consent verification”“self-hostable open weight TTS”
Queries 8, 9, 14, 16, 20Profile
#Query TopicProduct LineClaudeChatGPTGemini
1Sub-300ms latency under loadReal-Time Agent✗✗✗
Exact question asked across all AI platforms:

“We’re building a customer-support voice agent and callers keep noticing the pause before the bot replies. Which text-to-speech APIs actually hold sub-300ms time-to-first-audio under real production load, not just in a demo?”

2Measured vs published latencyReal-Time Agent✗✗✗
Exact question asked across all AI platforms:

“Which text-to-speech provider has the lowest measured time-to-first-audio for streaming voice agents, and how do the independent benchmarks compare to the vendors’ own claimed numbers?”

3Flat per-character pricingNarration, STT & Pricing✓✗✓
Exact question asked across all AI platforms:

“Our voice agent product is scaling past 50 million characters a month and the TTS line item is eating our gross margin. Which speech vendors offer flat per-character pricing at that volume instead of credit systems that get more expensive as you grow?”

4Streaming framework integrationsReal-Time Agent✗✗✗
Exact question asked across all AI platforms:

“We run our voice pipeline on an open-source orchestration framework and need a speech vendor with a documented WebSocket streaming integration. Which TTS providers have first-party support for real-time agent frameworks?”

5Per-sentence tone controlEmotion Control✗✗✗
Exact question asked across all AI platforms:

“Our support agent sounds correct but emotionally flat, so customers hang up. Are there speech models where I can control tone and delivery per sentence from the API rather than picking one fixed voice?”

6Bundled speech-to-textNarration, STT & Pricing✗✗✗
Exact question asked across all AI platforms:

“We need speech-to-text and text-to-speech from the same vendor to cut a hop out of our voice pipeline. Which providers do both well, and what does that usually cost per audio hour?”

7SIP telephony at scaleReal-Time Agent✗✗✗
Exact question asked across all AI platforms:

“We’re deploying voice agents over SIP telephony at contact-centre scale with thousands of concurrent calls. Which speech vendors are proven for that kind of telephony workload?”

8Inline performance tagsEmotion Control✗✗✗
Exact question asked across all AI platforms:

“I’m building an AI companion app and the voice has to laugh, sigh and whisper on cue instead of reading everything in one flat tone. Which speech platforms let you direct the performance with tags inside the script?”

9Large commercial voice libraryCloning & Library✗✗✗
Exact question asked across all AI platforms:

“We need dozens of distinct character voices for an interactive story app without hiring dozens of actors. Which voice platforms have large ready-made voice libraries you can license and use commercially?”

10Minimum clone reference clipCloning & Library✓✗✓
Exact question asked across all AI platforms:

“How short can a reference clip be before voice cloning quality falls apart, and which vendors can produce a usable clone from well under a minute of audio?”

11Native Japanese and KoreanMultilingual & Dubbing✗✗✗
Exact question asked across all AI platforms:

“Our app is biggest in Japan and Korea and the AI voices we’ve tried sound like an English model reading Japanese. Which speech vendors have genuinely native-quality Japanese and Korean voices?”

12Full character layer for gamesReal-Time Agent✗✗✗
Exact question asked across all AI platforms:

“For a game with NPCs that need voice, memory and dialogue behaviour, are there platforms that handle the whole character layer rather than just converting text into audio?”

13Long-form narration consistencyNarration, STT & Pricing✓✗✗
Exact question asked across all AI platforms:

“I’m producing a 12-hour audiobook and need the narrator’s voice and pacing to stay consistent from chapter one to the end. Which AI narration tools hold up over long-form content?”

14Broadcast-grade English fidelityEmotion Control✗✗✗
Exact question asked across all AI platforms:

“For a premium brand video where the voiceover has to be indistinguishable from a professional English voice actor, which AI voice platform currently produces the highest fidelity output?”

15In-VPC and air-gapped deploymentSelf-Host & Compliance✗✗✗
Exact question asked across all AI platforms:

“Our compliance team won’t let patient audio leave our own environment. Which speech vendors let you run the actual model inside your own VPC or an air-gapped data centre?”

16Open published weightsSelf-Host & Compliance✓⚠✗
Exact question asked across all AI platforms:

“Are there any production-grade text-to-speech models with openly published weights that a company can self-host and fine-tune, rather than being locked into a hosted API?”

17HIPAA, BAA, SOC 2 and ZDRSelf-Host & Compliance✗✗✗
Exact question asked across all AI platforms:

“What should a healthcare organisation check before approving an AI voice vendor — HIPAA, BAA, SOC 2, zero data retention — and which speech vendors actually meet those today?”

18Voice-preserving translationMultilingual & Dubbing✗✗✗
Exact question asked across all AI platforms:

“We need to localise internal training and customer content into a dozen languages while keeping the same speaker’s voice across all of them. Which AI dubbing or voice translation tools do this reliably?”

19Enterprise SLA and data residencySelf-Host & Compliance✗✗✗
Exact question asked across all AI platforms:

“We’re replacing a legacy speech synthesis vendor and need predictable annual cost, an uptime SLA and data residency options in the US, EU and Asia. Which AI voice vendors are set up for that kind of enterprise contract?”

20Consent and rights managementCloning & Library✗✗✗
Exact question asked across all AI platforms:

“Our legal team is worried about voice rights and consent when cloning an employee’s or spokesperson’s voice. Which AI voice vendors have the strongest consent verification and rights-management controls?”

MENTIONS4/20 (20%)1/20 (5%)2/20 (10%)
FIRST PLACEMENTS000

▹ Query 16 carries a warning marker on ChatGPT. Fish Audio was named there, then flagged for a research licence restricting commercial use, with Kokoro-82M recommended instead. It is counted as a mention and separately flagged as qualified; it is not treated as an endorsement. Ordinal rank was captured for that response only, at position 6. Ranks for the Claude and Gemini mentions were not recorded during scoring; no Fish Audio mention on any platform was a first placement.

Section 5

Seven Mentions, Zero First Placements

Where Fish Audio surfaced, and on what terms

The seven mentions fall on four questions only: per-byte pricing, cloning reference length, long-form narration and open weights. Sixteen of the twenty questions produced no Fish Audio mention on any platform. On the four where it did appear, it was never the first vendor named.

“Are there any production-grade text-to-speech models with openly published weights that a company can self-host and fine-tune, rather than being locked into a hosted API?”

— Query 16. Fish Audio is technically relevant here because it publishes downloadable model weights and supports self-hosted deployment. ChatGPT names it at rank 6 and Claude names it, both flagging that the current Fish Audio Research License does not grant unrestricted commercial use and that separate licensing is required. Gemini does not name it at all. The caveat the models raise is consistent with the published licence terms.

“We need dozens of distinct character voices for an interactive story app without hiring dozens of actors. Which voice platforms have large ready-made voice libraries you can license and use commercially?”

— Query 9. All three platforms name ElevenLabs and describe its library as the largest ready-made commercial catalogue, in the range of 10,000 to 11,000 voices. Fish Audio advertises more than 2,000,000 community-uploaded voices and is named by none of them. The two figures are not directly comparable without separating community uploads from commercially cleared inventory.

“Our support agent sounds correct but emotionally flat, so customers hang up. Are there speech models where I can control tone and delivery per sentence from the API rather than picking one fixed voice?”

— Query 5, and Query 8 repeats the pattern. Both describe inline, script-level performance direction accurately and name other vendors. Fish Audio ships a tag-based direction system and appears on neither.
16 of 20 Queries: No Mention On Any Platform
Latency under load, measured benchmarks, streaming framework integrations, bundled speech-to-text, SIP telephony, character layers for games, inline performance tags, voice library scale, Japanese and Korean quality, broadcast fidelity, in-VPC deployment, HIPAA and BAA, voice-preserving translation, enterprise SLAs, dubbing and consent controls. Silent on all sixteen, across all three platforms.
Query 16: Opportunity And Caveat Together
Fish Audio is the only vendor in the Query 16 answers whose entry pairs an accurate capability description with a licensing qualification. Kokoro, Parler-TTS, CosyVoice and MeloTTS received unqualified recommendations. Three of four sessions reported the same commercial-use restriction, which is consistent with the published Research License rather than a misreading of it. Commercial self-hosting is available from Fish Audio under separate licensing; that path is not reflected in the answers.
Capability Recognised. Vendor Not Named.

On the questions describing inline emotion tags, open-weight self-hosting and large commercial voice libraries, the tested answers describe the capability accurately and name other vendors. Across three platforms the competitor set shifted substantially between sessions, with only ElevenLabs, Cartesia and Deepgram appearing consistently throughout. Fish Audio was absent from every variation. This is a measured association gap. Determining its cause would require a source-authority and retrieval audit, which is outside the scope of these 60 responses.

Section 6

AI Topic Authority Map

Query heatmap — product line × platform

TopicAI LeaderFish Audio Status
Real-time latency under loadInworld / CartesiaINVISIBLE (0/3)
SIP telephony and contact-centre scaleDeepgramINVISIBLE (0/3)
Inline emotion and performance tagsElevenLabs / InworldINVISIBLE (0/3)
Voice library scaleElevenLabsINVISIBLE (0/3)
Japanese and Korean qualityElevenLabs / AzureINVISIBLE (0/3)
Compliance, SLA and data residencyDeepgram / AzureINVISIBLE (0/3)
Voice cloning reference minimumsInworld / CartesiaClaude + Gemini (2/3)
Per-character pricing at volumeDeepgramClaude + Gemini (2/3)
Open weights and self-hostingKokoro-82MClaude + ChatGPT (2/3)
Long-form audiobook narrationElevenLabsClaude only (1/3)
Product Line
Claude
ChatGPT
Gemini
Real-Time Voice Agent TTS
5 queries
0%
0%
0%
Expressive TTS & Emotion Control
3 queries
0%
0%
0%
Voice Cloning & Voice Library
3 queries
33%
0%
33%
Multilingual & Dubbing
2 queries
0%
0%
0%
Self-Host & Enterprise Compliance
4 queries
25%
25%
0%
Narration, STT & Pricing
3 queries
67%
0%
33%

▹ Narration, STT & Pricing is the only Fish Audio product line to clear 50% on any platform. Real-Time Voice Agent TTS and Expressive TTS & Emotion Control both return zero across all three platforms, and together they account for 8 of the 20 queries tested.

Real-Time Voice Agent TTS • 5 queries
Claude0%
ChatGPT0%
Gemini0%
Expressive TTS & Emotion Control • 3 queries
Claude0%
ChatGPT0%
Gemini0%
Voice Cloning & Voice Library • 3 queries
Claude33%
ChatGPT0%
Gemini33%
Multilingual & Dubbing • 2 queries
Claude0%
ChatGPT0%
Gemini0%
Self-Host & Enterprise Compliance • 4 queries
Claude25%
ChatGPT25%
Gemini0%
Narration, STT & Pricing • 3 queries
Claude67%
ChatGPT0%
Gemini33%
Zero Product Lines Above 70% On Any Platform
Not one of the six product lines reaches the threshold that would normally indicate topic ownership. Four of the six sit at or near zero across all three platforms.
Real-Time Voice Agent TTS: 0% Across 5 Queries
This is the line the enterprise motion depends on and the line the $52M raise is funding. Inworld, Cartesia and Deepgram split every one of those five answers between them.
Section 7

Methodology

How this Xtrusio AEO/GEO Audit was conducted, and what it does and does not measure

ParameterDefinition used in this audit
Test windowSeptember 2026
Scored baseline20 queries × 3 platforms = 60 scored responses
Platforms and modeChatGPT, Claude and Gemini. All three scored in conversational, search-augmented mode
SessionsFour sessions run, three scored. A Gemini custom-configuration session and an earlier ChatGPT session were superseded by conversational re-audits on the same locked question set
Mention ruleCounted when Fish Audio is explicitly named as a relevant provider or option in the answer body
Rank ruleOrdinal position of the vendor in the answer. A first placement means the vendor is the first provider named
Qualified mentionsA mention that carries a caveat or a recommendation against the vendor still counts toward visibility, is flagged separately, and is not treated as an endorsement. One qualified mention occurred, at Query 16 on ChatGPT
Source URLsNot logged. This audit measures vendor naming and rank. It does not identify which source caused a model to surface a vendor
VolatilityPoint-in-time snapshot. Generative outputs are non-deterministic and can vary on rerun. Two platforms returned different scores across sessions on the same questions
Not measuredProduct quality, latency, accuracy or fitness for purpose. No claim is made about why any vendor surfaced more often than another
Query Design
Twenty discovery-phase questions, none of which names Fish Audio or any competitor. Questions mirror how a VP of AI or a head of audiobook production researches vendors before a shortlist exists, across real-time voice agents, expressive speech, cloning, multilingual delivery, self-hosting and compliance.
Competitor Scope
The primary comparison set was defined before testing: ElevenLabs, Inworld, Deepgram, Cartesia and Microsoft Azure Speech. The leaderboard also includes Google, which surfaced repeatedly in the scored responses. It is not an exhaustive market census. Deepgram appears because the question set includes real-time voice-agent, speech infrastructure and enterprise deployment questions; not every appearance represents a direct text-to-speech substitution.
Client Research
Product and positioning reviewed at fish.audio across text-to-speech, cloning, speech-to-text, self-hosting and enterprise tiers, including published benchmark work, pricing and licensing terms. Fish Audio publishes first-party blind-testing and benchmark results; independently hosted benchmarks can differ materially by measurement setup, and both are treated as context rather than as a ranking.
Persona Selection Exceptions
All three personas are real, currently-serving decision-makers verified on LinkedIn. Name, title and company are verified from public profiles; the pain points shown are buyer scenarios constructed from role and company context. Two exceptions were applied under the standard rubric. Emily Derr shows no posting activity in the last 30 days; the exception was granted on the basis of 14 years of continuous tenure. Tiffany Frarey holds a title below the usual Director threshold; the exception was granted on the basis of documented end-to-end production ownership including budget. The persona set scored 90 of 100 on the eight-criterion rubric.
Section 8

Recommendations

Each action is a hypothesis to test against a rerun of the same locked 20 queries

Phase 1 — 0–30 Days
Fix Category Association Around The Proof That Already Exists
  • Publish a single canonical licensing page stating which models carry which terms, what commercial self-hosting costs and what is permitted without a paid licence. Two platforms surfaced the Research License restriction accurately; the commercial path exists but is not visible alongside it
  • Turn the existing blind-test and benchmark work into buyer-facing comparison assets with stated test methodology, model and version dates, and independent replication where possible
  • Separate community-uploaded voices from commercially cleared inventory in all public materials, and publish a dated count for each. Every platform named ElevenLabs on the commercial-library question
  • Correct the pricing unit in public materials to $15 per 1M UTF-8 bytes, roughly 1M English characters depending on the text
Phase 2 — 30–90 Days
Strengthen And Standardise The Enterprise Proof That Is Already Marketed
  • Consolidate the existing VPC, on-premises and air-gapped deployment options into clear buyer-facing pages written in the language of Queries 15, 17 and 19, substantiating each claim with contractual or certification detail
  • Document current security and compliance posture, available deployment controls, and any certifications or BAAs that can be independently verified. Avoid asserting certification status that cannot be shown
  • Publish the voice-rights and consent workflow as a buyer-facing page. Query 20 returned zero mentions across all three platforms
  • Document the tag-based direction system as a named, citable capability with a public reference list. Queries 5 and 8 describe it precisely and name other vendors
Phase 3 — 90+ Days
Build Independently Verifiable Authority, Then Retest
  • Pursue inclusion in third-party voice-AI benchmarks, technical comparisons and category roundups. This is a hypothesis, not a guarantee: no third-party coverage, schema or benchmark placement can be shown to cause a model to recommend a vendor
  • Publish evaluated Japanese, Korean, Mandarin and Cantonese comparisons. Query 11 named ElevenLabs and Azure on every platform despite this being a stated strength
  • Rerun the same locked 20 queries under the same conditions and measure whether Fish Audio moves from mentions into higher placements. Quarterly Xtrusio re‑audits track movement off the 11.7% baseline
Continuous AI Visibility Tracking
Brands can improve their AI discovery using generative engine optimization tools like Xtrusio.

Strong product traction. Weak AI recommendation visibility.

The next step is making existing proof easier to verify and to associate with the buyer questions that already match the product.

This research report was generated using the Xtrusio Company Intelligence Module.