50 answers name ElevenLabs.
Seven name Fish Audio. None put it first.
Across 20 buyer-intent questions tested on ChatGPT, Claude and Gemini, Fish Audio is named in 7 of 60 responses (11.7%) and records zero first placements. That visibility is low relative to the company’s reported scale — more than 8 million users and $21 million in ARR — while ElevenLabs appears in 50 responses, Deepgram and Cartesia in 32 each, and Inworld in 28. The gap measured here is AI recommendation visibility, not product capability.
The findings below come from Xtrusio, an AI visibility audit system built specifically for B2B buyer-intent testing. Every result was produced by running 20 real prospect queries across three generative AI platforms and recording which vendors were named.
Scored against the voice AI and text-to-speech infrastructure category, where buyers compare vendors on latency, deployment model, language coverage and per-character cost.
Fish Audio already has measurable proof. It is not consistently surfacing in these answers.
Fish Audio publishes benchmark results, pricing, technical documentation and open-weight model information. Competitors including ElevenLabs, Deepgram and Cartesia nevertheless surfaced far more often across the 60 responses tested — 50, 32 and 32 respectively, against Fish Audio’s 7. This audit establishes that difference in AI visibility. It does not establish why it occurs. A separate source-authority and retrieval analysis would be needed to determine whether third-party coverage, enterprise documentation, model association or another factor contributes to the gap.
Platform Scorecard
Where Fish Audio was named, and where it was not
Scored across all 60 responses. Deepgram and Cartesia both finish on 32, ahead of Inworld — the infrastructure and compliance questions carry them. Mention frequency is a visibility measure, not a quality ranking.
AI Visibility Leaderboard
Who the AI answers name when buyers ask about voice infrastructure
All bars scaled to the highest-mentioned vendor (50 of 60).
First placements counted as the first vendor named in an answer. The six competitors above account for 45 of the 60 first placements; the remaining 15 went to vendors outside the primary comparison set.
AI Positioning Audit
20 buyer-intent queries — click any row to see the exact question
Each query was written from the perspective of a real decision-maker researching voice AI during discovery, before any vendor shortlist exists. These are the buyers whose AI answers determine whether Fish Audio is considered at all.
| # | Query Topic | Product Line | Claude | ChatGPT | Gemini |
|---|---|---|---|---|---|
| 1 | Sub-300ms latency under load | Real-Time Agent | ✗ | ✗ | ✗ |
Exact question asked across all AI platforms: “We’re building a customer-support voice agent and callers keep noticing the pause before the bot replies. Which text-to-speech APIs actually hold sub-300ms time-to-first-audio under real production load, not just in a demo?” | |||||
| 2 | Measured vs published latency | Real-Time Agent | ✗ | ✗ | ✗ |
Exact question asked across all AI platforms: “Which text-to-speech provider has the lowest measured time-to-first-audio for streaming voice agents, and how do the independent benchmarks compare to the vendors’ own claimed numbers?” | |||||
| 3 | Flat per-character pricing | Narration, STT & Pricing | ✓ | ✗ | ✓ |
Exact question asked across all AI platforms: “Our voice agent product is scaling past 50 million characters a month and the TTS line item is eating our gross margin. Which speech vendors offer flat per-character pricing at that volume instead of credit systems that get more expensive as you grow?” | |||||
| 4 | Streaming framework integrations | Real-Time Agent | ✗ | ✗ | ✗ |
Exact question asked across all AI platforms: “We run our voice pipeline on an open-source orchestration framework and need a speech vendor with a documented WebSocket streaming integration. Which TTS providers have first-party support for real-time agent frameworks?” | |||||
| 5 | Per-sentence tone control | Emotion Control | ✗ | ✗ | ✗ |
Exact question asked across all AI platforms: “Our support agent sounds correct but emotionally flat, so customers hang up. Are there speech models where I can control tone and delivery per sentence from the API rather than picking one fixed voice?” | |||||
| 6 | Bundled speech-to-text | Narration, STT & Pricing | ✗ | ✗ | ✗ |
Exact question asked across all AI platforms: “We need speech-to-text and text-to-speech from the same vendor to cut a hop out of our voice pipeline. Which providers do both well, and what does that usually cost per audio hour?” | |||||
| 7 | SIP telephony at scale | Real-Time Agent | ✗ | ✗ | ✗ |
Exact question asked across all AI platforms: “We’re deploying voice agents over SIP telephony at contact-centre scale with thousands of concurrent calls. Which speech vendors are proven for that kind of telephony workload?” | |||||
| 8 | Inline performance tags | Emotion Control | ✗ | ✗ | ✗ |
Exact question asked across all AI platforms: “I’m building an AI companion app and the voice has to laugh, sigh and whisper on cue instead of reading everything in one flat tone. Which speech platforms let you direct the performance with tags inside the script?” | |||||
| 9 | Large commercial voice library | Cloning & Library | ✗ | ✗ | ✗ |
Exact question asked across all AI platforms: “We need dozens of distinct character voices for an interactive story app without hiring dozens of actors. Which voice platforms have large ready-made voice libraries you can license and use commercially?” | |||||
| 10 | Minimum clone reference clip | Cloning & Library | ✓ | ✗ | ✓ |
Exact question asked across all AI platforms: “How short can a reference clip be before voice cloning quality falls apart, and which vendors can produce a usable clone from well under a minute of audio?” | |||||
| 11 | Native Japanese and Korean | Multilingual & Dubbing | ✗ | ✗ | ✗ |
Exact question asked across all AI platforms: “Our app is biggest in Japan and Korea and the AI voices we’ve tried sound like an English model reading Japanese. Which speech vendors have genuinely native-quality Japanese and Korean voices?” | |||||
| 12 | Full character layer for games | Real-Time Agent | ✗ | ✗ | ✗ |
Exact question asked across all AI platforms: “For a game with NPCs that need voice, memory and dialogue behaviour, are there platforms that handle the whole character layer rather than just converting text into audio?” | |||||
| 13 | Long-form narration consistency | Narration, STT & Pricing | ✓ | ✗ | ✗ |
Exact question asked across all AI platforms: “I’m producing a 12-hour audiobook and need the narrator’s voice and pacing to stay consistent from chapter one to the end. Which AI narration tools hold up over long-form content?” | |||||
| 14 | Broadcast-grade English fidelity | Emotion Control | ✗ | ✗ | ✗ |
Exact question asked across all AI platforms: “For a premium brand video where the voiceover has to be indistinguishable from a professional English voice actor, which AI voice platform currently produces the highest fidelity output?” | |||||
| 15 | In-VPC and air-gapped deployment | Self-Host & Compliance | ✗ | ✗ | ✗ |
Exact question asked across all AI platforms: “Our compliance team won’t let patient audio leave our own environment. Which speech vendors let you run the actual model inside your own VPC or an air-gapped data centre?” | |||||
| 16 | Open published weights | Self-Host & Compliance | ✓ | ⚠ | ✗ |
Exact question asked across all AI platforms: “Are there any production-grade text-to-speech models with openly published weights that a company can self-host and fine-tune, rather than being locked into a hosted API?” | |||||
| 17 | HIPAA, BAA, SOC 2 and ZDR | Self-Host & Compliance | ✗ | ✗ | ✗ |
Exact question asked across all AI platforms: “What should a healthcare organisation check before approving an AI voice vendor — HIPAA, BAA, SOC 2, zero data retention — and which speech vendors actually meet those today?” | |||||
| 18 | Voice-preserving translation | Multilingual & Dubbing | ✗ | ✗ | ✗ |
Exact question asked across all AI platforms: “We need to localise internal training and customer content into a dozen languages while keeping the same speaker’s voice across all of them. Which AI dubbing or voice translation tools do this reliably?” | |||||
| 19 | Enterprise SLA and data residency | Self-Host & Compliance | ✗ | ✗ | ✗ |
Exact question asked across all AI platforms: “We’re replacing a legacy speech synthesis vendor and need predictable annual cost, an uptime SLA and data residency options in the US, EU and Asia. Which AI voice vendors are set up for that kind of enterprise contract?” | |||||
| 20 | Consent and rights management | Cloning & Library | ✗ | ✗ | ✗ |
Exact question asked across all AI platforms: “Our legal team is worried about voice rights and consent when cloning an employee’s or spokesperson’s voice. Which AI voice vendors have the strongest consent verification and rights-management controls?” | |||||
| MENTIONS | 4/20 (20%) | 1/20 (5%) | 2/20 (10%) | ||
| FIRST PLACEMENTS | 0 | 0 | 0 | ||
▹ Query 16 carries a warning marker on ChatGPT. Fish Audio was named there, then flagged for a research licence restricting commercial use, with Kokoro-82M recommended instead. It is counted as a mention and separately flagged as qualified; it is not treated as an endorsement. Ordinal rank was captured for that response only, at position 6. Ranks for the Claude and Gemini mentions were not recorded during scoring; no Fish Audio mention on any platform was a first placement.
Seven Mentions, Zero First Placements
Where Fish Audio surfaced, and on what terms
The seven mentions fall on four questions only: per-byte pricing, cloning reference length, long-form narration and open weights. Sixteen of the twenty questions produced no Fish Audio mention on any platform. On the four where it did appear, it was never the first vendor named.
“Are there any production-grade text-to-speech models with openly published weights that a company can self-host and fine-tune, rather than being locked into a hosted API?”
“We need dozens of distinct character voices for an interactive story app without hiring dozens of actors. Which voice platforms have large ready-made voice libraries you can license and use commercially?”
“Our support agent sounds correct but emotionally flat, so customers hang up. Are there speech models where I can control tone and delivery per sentence from the API rather than picking one fixed voice?”
On the questions describing inline emotion tags, open-weight self-hosting and large commercial voice libraries, the tested answers describe the capability accurately and name other vendors. Across three platforms the competitor set shifted substantially between sessions, with only ElevenLabs, Cartesia and Deepgram appearing consistently throughout. Fish Audio was absent from every variation. This is a measured association gap. Determining its cause would require a source-authority and retrieval audit, which is outside the scope of these 60 responses.
AI Topic Authority Map
Query heatmap — product line × platform
| Topic | AI Leader | Fish Audio Status |
|---|---|---|
| Real-time latency under load | Inworld / Cartesia | INVISIBLE (0/3) |
| SIP telephony and contact-centre scale | Deepgram | INVISIBLE (0/3) |
| Inline emotion and performance tags | ElevenLabs / Inworld | INVISIBLE (0/3) |
| Voice library scale | ElevenLabs | INVISIBLE (0/3) |
| Japanese and Korean quality | ElevenLabs / Azure | INVISIBLE (0/3) |
| Compliance, SLA and data residency | Deepgram / Azure | INVISIBLE (0/3) |
| Voice cloning reference minimums | Inworld / Cartesia | Claude + Gemini (2/3) |
| Per-character pricing at volume | Deepgram | Claude + Gemini (2/3) |
| Open weights and self-hosting | Kokoro-82M | Claude + ChatGPT (2/3) |
| Long-form audiobook narration | ElevenLabs | Claude only (1/3) |
5 queries
3 queries
3 queries
2 queries
4 queries
3 queries
▹ Narration, STT & Pricing is the only Fish Audio product line to clear 50% on any platform. Real-Time Voice Agent TTS and Expressive TTS & Emotion Control both return zero across all three platforms, and together they account for 8 of the 20 queries tested.
Methodology
How this Xtrusio AEO/GEO Audit was conducted, and what it does and does not measure
This research is based on Xtrusio’s proprietary AI visibility analysis framework.
| Parameter | Definition used in this audit |
|---|---|
| Test window | September 2026 |
| Scored baseline | 20 queries × 3 platforms = 60 scored responses |
| Platforms and mode | ChatGPT, Claude and Gemini. All three scored in conversational, search-augmented mode |
| Sessions | Four sessions run, three scored. A Gemini custom-configuration session and an earlier ChatGPT session were superseded by conversational re-audits on the same locked question set |
| Mention rule | Counted when Fish Audio is explicitly named as a relevant provider or option in the answer body |
| Rank rule | Ordinal position of the vendor in the answer. A first placement means the vendor is the first provider named |
| Qualified mentions | A mention that carries a caveat or a recommendation against the vendor still counts toward visibility, is flagged separately, and is not treated as an endorsement. One qualified mention occurred, at Query 16 on ChatGPT |
| Source URLs | Not logged. This audit measures vendor naming and rank. It does not identify which source caused a model to surface a vendor |
| Volatility | Point-in-time snapshot. Generative outputs are non-deterministic and can vary on rerun. Two platforms returned different scores across sessions on the same questions |
| Not measured | Product quality, latency, accuracy or fitness for purpose. No claim is made about why any vendor surfaced more often than another |
Recommendations
Each action is a hypothesis to test against a rerun of the same locked 20 queries
- Publish a single canonical licensing page stating which models carry which terms, what commercial self-hosting costs and what is permitted without a paid licence. Two platforms surfaced the Research License restriction accurately; the commercial path exists but is not visible alongside it
- Turn the existing blind-test and benchmark work into buyer-facing comparison assets with stated test methodology, model and version dates, and independent replication where possible
- Separate community-uploaded voices from commercially cleared inventory in all public materials, and publish a dated count for each. Every platform named ElevenLabs on the commercial-library question
- Correct the pricing unit in public materials to $15 per 1M UTF-8 bytes, roughly 1M English characters depending on the text
- Consolidate the existing VPC, on-premises and air-gapped deployment options into clear buyer-facing pages written in the language of Queries 15, 17 and 19, substantiating each claim with contractual or certification detail
- Document current security and compliance posture, available deployment controls, and any certifications or BAAs that can be independently verified. Avoid asserting certification status that cannot be shown
- Publish the voice-rights and consent workflow as a buyer-facing page. Query 20 returned zero mentions across all three platforms
- Document the tag-based direction system as a named, citable capability with a public reference list. Queries 5 and 8 describe it precisely and name other vendors
- Pursue inclusion in third-party voice-AI benchmarks, technical comparisons and category roundups. This is a hypothesis, not a guarantee: no third-party coverage, schema or benchmark placement can be shown to cause a model to recommend a vendor
- Publish evaluated Japanese, Korean, Mandarin and Cantonese comparisons. Query 11 named ElevenLabs and Azure on every platform despite this being a stated strength
- Rerun the same locked 20 queries under the same conditions and measure whether Fish Audio moves from mentions into higher placements. Quarterly Xtrusio re‑audits track movement off the 11.7% baseline
Strong product traction. Weak AI recommendation visibility.
The next step is making existing proof easier to verify and to associate with the buyer questions that already match the product.
This research report was generated using the Xtrusio Company Intelligence Module.


