
Fish Audio is a text-to-speech and voice cloning platform built around word-level emotion control. It clones a voice from a ten-second sample, streams at sub-300ms latency, and publishes its S2 model weights openly — though under a research licence that excludes commercial use. Operated by Hanabi AI Inc.
Used this tool? Rate it
Used this tool? Rate it
Fish Audio is a voice AI platform built around one core idea: that getting the words right and getting the delivery right are two different problems, and that most text-to-speech systems have only ever solved the first. The site presents three capability lines on its homepage — text to speech, voice cloning, and speech to text — with a model family called S2 at the centre of all three.
The product's defining feature is word-level expressive control. Rather than generating a flat read of your script, the system accepts inline natural-language tags that direct how a line is delivered. The homepage demo exposes an emotion set including [angry], [sad], [embarrassed], [emphasis], [whispering], [soft], [breathy] and [excited], alongside a special set covering [laughing], [chuckling], [clear throat], [sobbing], [sighing], [panting] and both short and long pauses. The technical framing on the project's own repository is more precise than the marketing: it describes support for sub-word level fine-grained control of prosody and emotion using natural language tags, while natively supporting multi-speaker and multi-turn conversation generation.
The company is not anonymous, though its naming is worth untangling. The site footer reads "© 2026 Hanabi AI Inc. All rights reserved." The open-source licence file, meanwhile, requires an attribution notice naming a different entity: "Copyright © 39 AI, INC." Both names are documented here rather than reconciled, since the site does not explain the relationship between them.
Its funding and scale are publicly stated in the company's own announcement. Fish Audio raised $52 million in seed funding on its first anniversary, in a round led by Coreline Ventures and Capital Today with participation from 359 Capital, Play Time, HF0, 645 Ventures and others. The same post reports scaling from three people to twenty-two, going from zero to $21M in annual recurring revenue, 8M+ creators and developers on the platform, and 2M+ voice models in the community library. Those business metrics are self-reported and have not been independently audited, so they should be read as company statements rather than verified figures. The post also names the leadership: CEO Rissa Cao, who cites years in voice AI at Amazon Alexa and Meta, and chief scientist Shijia Liao, formerly a research engineer at NVIDIA who began training the models on a single 4090 in his bedroom.
Two things about this product deserve more scrutiny than a feature list provides, and both are covered in depth below rather than glossed. The first is what the platform actually requires regarding consent when you clone someone's voice. The second is what "open source" means here, because the licence is not one of the permissive ones the phrase usually implies. Neither answer is a reason to avoid the product, but both change how you should use it.
The cloning pitch is specific and, unusually for this category, comes with published numbers rather than adjectives. The product page states that ten seconds of reference audio is enough, that streaming latency is sub-300ms end-to-end, and that cloning is zero-shot cross-lingual across thirteen languages — meaning you clone once and the same voice speaks other supported languages without retraining or a second recording session.
That last property is the one that changes workflows most. Traditional dubbing requires either a bilingual performer or a separate voice per language, which means a piece of content sounds like a different person in every market. Zero-shot cross-lingual cloning means one recorded identity can carry across a catalogue.
The TTS side is where the emotion tag system does its work. Rather than choosing from a fixed list of "styles", you annotate the script itself, and the tags apply at the point where they appear. This is why the sub-word claim matters: emphasis placed on one word inside a sentence is a different capability from setting a sentence-level mood, and it is the difference between narration that sounds read and narration that sounds performed.
The platform is considerably broader than a TTS endpoint. The products menu lists eight distinct tools: Text-to-Speech, Speech-to-Text, Voice Cloning, Voice Changer, Story Studio, Audio Separation, Audio Translation, and Sound Effects. The speech-to-text side is pitched as including multispeaker handling, emotion tags and natural language description in the transcription output, which is a more structured transcript than a plain text dump.
The platform hosts a large library of user-uploaded voices — over 2,000,000 by the site's own count — described as suited to creative storytelling, advertisements and audiobooks. This is a genuine differentiator in breadth, and it is also the part of the product most entangled with the consent and rights questions discussed later, because a library of that size is populated by users rather than curated by the company.
The S2 model weights are published, and the site actively markets self-hosting as an option alongside the hosted API. The repository is substantial by community measures: 32.4k stars, 2.8k forks, 101 contributors and 14 releases at the time of writing, described as SOTA Open Source TTS. What the licence actually permits is covered in its own section below, and it is narrower than "open source" normally suggests.
Audiobook and long-form narration. The company's own framing is that five seconds of audio is easy and five minutes is the test — that TTS quality has been adequate for single sentences for a decade and breaks down over length. Long-form narration is precisely where emotion tags and pacing control earn their keep, because a flat read becomes unbearable well before chapter two.
Video voiceover at production speed. Turning scripts into scene-matched narration without booking a booth is the highest-volume use of tools like this. The ability to swap tone and add emotion without re-recording is what makes iteration cheap.
Character voices for games and animation. The platform explicitly targets this, and it is a natural fit: interactive media needs many distinct voices, often more than a project can afford to cast, and needs them regenerated whenever dialogue changes.
Conversational agents and customer support. This is the use case where sub-300ms streaming latency is not a vanity metric. A voice agent that pauses noticeably before every reply feels broken regardless of how good the voice sounds, so latency is a functional requirement rather than a nicety.
Multilingual localisation from a single performance. Zero-shot cross-lingual cloning means recording once and shipping the same voice identity across supported languages — useful for creators whose audience spans regions, though the language-count question below deserves reading first.
Self-hosted deployment for latency or data reasons. Because the weights are published, teams with strict data-residency requirements or their own inference infrastructure can run the model themselves rather than sending audio to a third party. The licence determines whether that is permitted for your particular use, which is the crux of the section below.
Step 1 — Try the homepage demo before signing up. The homepage carries a working text-to-speech box preloaded with a script demonstrating the tag system, with a 30,000-character field and the emotion and special tag palettes visible. This is the fastest way to hear whether the expressive control is real for your use case.
Step 2 — Create a free account. The free tier requires no credit card. It provides 8,000 credits monthly, generation up to seven minutes, up to 500 characters per generation, and three public voice slots.
Step 3 — Decide between library voices and cloning. For many jobs, the existing library is the faster path and sidesteps the consent question entirely. Cloning is for when you need a specific voice identity — ideally your own, or one you have documented permission to use.
Step 4 — Record or select a clean ten-second reference. Ten seconds is the stated floor. Reference quality matters more than length: the cleaner the source, the better the clone, and the product page acknowledges longer samples can help with very expressive voices.
Step 5 — Write the script with tags, not just words. This is the step most new users under-use. Placing [emphasis] on the specific word that carries the sentence, or a [long pause] where a human would breathe, is what separates output that sounds generated from output that sounds performed.
Step 6 — Check your commercial rights before publishing. This is not optional housekeeping. As documented below, the free tier's commercial status is stated inconsistently across the site, and getting it wrong means publishing monetised content without the rights to do so.
Step 7 — Move to the API when volume justifies it. Developer documentation and an API reference are hosted on the docs subdomain, with paid plans providing pay-as-you-go API access.
Tag sparingly and deliberately. Emotion tags are a directing tool, not a decoration. Tagging every clause produces performance that sounds unstable; tagging the two or three moments that actually carry the emotional weight of a passage produces something that sounds intentional.
Invest in the reference recording, not the retries. A clone inherits the acoustics of its source. Ten seconds recorded on a decent microphone in a quiet room will outperform a minute of noisy phone audio, and no amount of regeneration fixes a compromised reference.
Budget in credits, not minutes. The published conversion is roughly 600 to 625 credits per minute of generation. Against the free tier's 8,000 monthly credits, that works out to roughly the seven minutes the plan advertises. Knowing the conversion makes plan comparison arithmetic rather than guesswork.
Remember credits do not roll over. Monthly quotas reset at the start of each billing cycle and unused minutes are not carried forward, so a plan sized for your peak month is wasted in your quiet ones.
Test your specific language before committing. As documented below, the platform's language support is stated four different ways across its own pages, and the open repository carries active bug reports about specific scripts. Generating a representative sample in your target language costs a few credits and answers the question definitively.
Use private voice slots for anything sensitive. Public submissions carry markedly broader licence terms than private ones, and public content may remain available after account deletion. Slot choice is a rights decision, not a filing preference.
Content creators producing at volume. YouTubers, podcasters and course producers who narrate regularly get the clearest return, because the cost saved scales with output.
Game and animation studios needing many voices. Projects requiring a large cast, or frequently changing dialogue, benefit more than projects with a small fixed script.
Developers building voice agents. The combination of a documented API, sub-300ms streaming, and published weights covers both the hosted and the self-hosted architectural choice.
Localisation teams. Cross-lingual cloning addresses a real and expensive problem in multi-market content.
Researchers and hobbyists. The published weights are genuinely available for research and non-commercial work, which is exactly the case the licence was written to permit.
Who should be cautious. Anyone planning commercial self-hosted deployment needs to read the licence section below first, because the default answer there is no. Anyone whose compliance posture requires explicit vendor commitments on biometric data handling should note what the privacy policy does and does not say. And anyone intending to clone a voice that is not their own needs the consent section below more than they need the feature list.
Fish Audio is primarily a web application, usable from any modern browser with no installation. Developer access is through a documented REST API with SDKs, hosted on the docs subdomain alongside an API reference.
Mobile applications exist: the terms of service include a mobile application section acknowledging that availability depends on third-party stores, naming Apple's App Store and the Android app market, and requiring users to comply with those stores' own terms.
Self-hosting is an officially promoted third path. The product page presents the choice directly — self-host the model, use the sub-300ms streaming endpoint, or ship voices into agents and apps — with the licence determining which of those is available to you.
Enterprise deployment options are listed separately and include on-premise deployment, zero data retention, and SOC2 compliance, quoted as custom volume pricing billed annually.
One access note: the site's robots.txt allows Googlebot, Applebot and Bingbot full access while disallowing general crawlers from /auth and /text-to-speech.
Pricing was verified directly on the plans page. Note that promotional discounts were stacked at the time of checking — a three-months-off annual offer plus an anniversary 50% discount — so the monthly-equivalent figures below reflect those promotions and the list prices are quoted alongside.
Free Tier. $0, no credit card needed. Includes 8,000 credits monthly, up to seven minutes of generation, up to 500 characters per generation, three public voice slots, standard generation speed and enhanced voice cloning.
Plus. $5.5/month equivalent, billed $66 annually, against a $15 monthly list price. Includes 250,000 credits monthly, up to 200 minutes of generation, up to 15,000 characters per generation, unlimited public plus 10 private voice slots, priority generation, access to Voice Design and one professional voice slot.
Pro. $37.5/month equivalent, billed $450 annually, against a $100 monthly list price. Includes 2,000,000 credits monthly, up to 1,620 minutes, three team seats, up to 30,000 characters per generation, unlimited voice slots, five professional voice slots and a seven-day money-back guarantee.
Max. $749/month equivalent, billed $8,988 annually, against a $999 monthly list price. Includes 25,000,000 credits monthly, up to 6,250 minutes, ten team seats and fifteen professional voice slots.
Enterprise. Custom volume pricing billed annually, aimed at organisations with compliance needs, listing pay-as-you-go with organisation-level controls, zero data retention, on-premise deployment and SOC2 compliance.
The published conversion is that each minute of generation costs roughly 600 to 625 credits. Monthly quotas reset at the beginning of each billing cycle and unused minutes do not roll over.
This deserves flagging plainly, because it directly affects whether you may publish what you generate, and the site contradicts itself on the same page.
The pricing page's Free Tier feature list includes "Commercial use" as a bullet point. The FAQ further down that same pricing page states the opposite: "Premium subscribers can use verified voices (that you own) for commercial purposes. Free plan users can only use generated content for personal, non-commercial projects." The homepage FAQ agrees with the restrictive reading and is more explicit still: "Fish Audio's free plan is for personal use only. To monetize content or use voices commercially (YouTube, podcasts, business), upgrade to our paid plans for full commercial rights."
So two official sources say free is personal-use-only, and one official feature list says otherwise. Both readings are reported here rather than reconciled, because reconciling them would mean guessing. The prudent course is to treat the restrictive reading as operative and confirm with the vendor before monetising free-tier output — the downside of being wrong falls entirely on the publisher.
The most-named comparison in this category is ElevenLabs, and Fish Audio engages with that comparison directly: the site maintains a dedicated comparison section and an ElevenLabs-alternative page appears in its own sitemap. Positioning itself explicitly against that incumbent is a deliberate choice, and prospective users should note that comparison pages published by a vendor about its competitor are marketing material, not independent benchmarking.
The company's own claims in this area are self-reported and should be read as such: its funding post states S2.1 Pro was preferred by 66% of listeners over leading competitors in blind listening tests. No methodology, sample size or listener pool is published alongside that figure in the announcement, so it cannot be independently verified from the source provided.
The genuine alternative axis, though, is not another hosted vendor — it is the build-versus-buy decision that the published weights make available. For research and non-commercial work, running the model yourself is a real option that most competitors do not offer. For commercial work, that option requires a separate licence, which collapses the comparison back to a straightforward hosted-service evaluation.
This is the question that matters most, and the honest answer is that responsibility sits almost entirely with the user, with the platform performing no advance check.
The clearest statement is in the voice cloning page's FAQ, and it is worth quoting exactly: "You are responsible for confirming you have the rights, consents, and disclosures required for any voice you clone, and for complying with applicable laws — including those covering name, likeness, and AI-generated content in your region." The same answer sets out the enforcement posture: "Fish Audio does not pre-clear individual use cases and may remove content or accounts that violate our terms or applicable law." So the mechanism is after-the-fact takedown, not a gate before cloning.
What is notably absent is any consent requirement in the binding legal documents. The terms of service section covering permissions and restrictions runs to seventeen enumerated prohibitions, from (a) through (q), and none of them addresses voice cloning consent specifically. The closest is the generic clause (a), which prohibits contributing anything that "infringes or violates the intellectual property rights or any other rights of anyone else". The user submissions warranty is similarly generic, prohibiting submissions that "infringe any third party's copyrights or other rights (e.g., trademark, privacy rights, etc.)". Neither imposes a specific obligation to hold a documented consent for a cloned voice.
There is also a tension between the consent language and the marketing on the very same page. The voice cloning page's own headline copy invites exactly the use the consent language warns about: "Have a sitting president narrate your dating app, run a tech-billionaire launch for your worst idea, or build a fake-panel podcast — no booth, no impressionist on retainer." Its FAQ reinforces the expectation by noting that "most public-figure clips, podcast cuts, or phone-quality recordings work on the first try." Marketing that foregrounds cloning identifiable public figures, sitting beside a disclaimer telling you to secure their consent, is a genuine inconsistency and readers should weigh both.
The terms do impose one related duty: users must not distribute content in a misleading way, "including, without limitation, representing that the Content is entirely human generated", and the platform encourages proactive disclosure that content was created using AI — though it frames disclosure as encouragement rather than obligation. An abuse reporting route exists via a Report Abuse link in the footer.
The practical conclusion is straightforward. If you are cloning your own voice, none of this constrains you. If you are cloning anyone else's, the platform will not stop you and will not check, but it also disclaims responsibility and places the legal exposure squarely on you — and depending on your jurisdiction, that exposure can be substantial.
The site markets openness repeatedly: "Open-source S2 model" appears as a headline feature, self-hosting is presented as a supported path, and the repository describes itself as SOTA Open Source TTS with 32.4k stars behind it. The weights genuinely are published, and that is real. But the licence is not a permissive one, and the difference matters commercially.
The repository states plainly that "This codebase and its associated model weights are released under FISH AUDIO RESEARCH LICENSE", warning that it will act against violations. That licence, last updated 7 March 2026, states its intent in its introduction: "This Agreement is intended to allow research and non-commercial uses of the Materials free of charge. Any Commercial use of the Materials requires a separate license from Fish Audio."
Section III removes any ambiguity: "Any use of the Fish Audio Materials or Derivative Works for a Commercial Purpose requires a separate written license agreement from Fish Audio. No commercial rights are granted under this Agreement." Commercial licences are handled through a business email address given in the licence itself.
The definition of "Commercial Purpose" is broad enough to catch most business use: it covers "(i) creating, modifying, or distributing Your product or service, including via a hosted service or application programming interface, (ii) Your business's or organization's internal operations, and (iii) any use in connection with a product or service for which You charge a fee or generate revenue, whether directly or indirectly." Internal operations at a company is enough to qualify — you do not have to be reselling the model.
Two further differences from permissive licensing deserve attention. The research grant is described as "non-exclusive, worldwide, non-transferable, non-sublicensable, revocable and royalty-free" — revocable is the operative word, and it has no equivalent in MIT or Apache terms. And distribution carries attribution duties: a notice file crediting "Copyright © 39 AI, INC." plus prominent display of "Built with Fish Audio" on a related website or product documentation. The licence also forbids using the materials or their outputs to create or improve any foundational generative AI model.
One gap is worth recording. The licence incorporates an Acceptable Use Policy by reference, making compliance with it a licence condition — but that document could not be located. Candidate paths including /acceptable-use, /aup and /policies/acceptable-use all return 404, and no such page appears in the site's own static sitemap, which lists only the homepage, discovery, AI voice generator, customers, enterprise, ElevenLabs-alternative, terms, privacy and Japan legal information pages. A binding condition that points to a document users cannot find is a documentation problem worth raising with the vendor before relying on the licence.
The short version: open weights, yes; open source in the permissive sense, no. Research and hobby use are free and genuinely permitted. Anything commercial — including internal business use — needs a separate paid agreement, regardless of what the marketing page suggests.
The privacy policy carries an effective date of 28 August 2024. A full-text check of the policy body found zero occurrences of the words "voice", "biometric" and "train". For a product whose core function is capturing and reproducing human voices, the absence of any voice-specific or biometric-data provision is a substantive gap, and it also means the policy makes no statement whatsoever about whether user content is used to train or improve models. Per the evidence available, the honest characterisation is that the policy is silent on these points — not that any particular protection exists or does not.
Uploaded audio falls under the generic "User Content" category, which the policy defines as personal information included in input, file uploads or feedback. The third parties listed for that category include service providers, advertising partners, analytics partners and business partners.
The terms are considerably more specific about the rights you give than about the rights you retain. Licences granted to the platform over user submissions are "royalty-free, perpetual, sublicensable, irrevocable, and worldwide."
Deletion is bounded accordingly. On account deletion the platform stops displaying your submissions to other users, "(other than Public User Submissions, which may remain fully available)", and the terms acknowledge "it may not be possible to completely delete that content from Fish.Audio's records."
Public submissions carry the broadest terms of all, including the right to "use, display, perform, and distribute your Public User Submission for Fish.Audio's marketing and promotional purposes" and a licence to all other users to access and exercise rights in it. This is the concrete reason the private-versus-public slot choice is a rights decision. It is also worth understanding as context for the 2M+ community voice library — that library exists because public uploads carry these terms.
Official pages give four mutually inconsistent counts, and all four are reported here rather than averaged. The voice cloning page states zero-shot cross-lingual across thirteen languages. The homepage's API section says voices speak 30+ languages. The TTS page's FAQ enumerates just eight — English, Japanese, Korean, Chinese, French, German, Arabic and Spanish — and separately advertises "Automatic support for 8 languages with native accents". The funding announcement claims 83+ languages with native cadence.
These may describe different things: cloning versus synthesis, fully-supported versus best-effort. But the site does not explain the distinction, so verify your specific language empirically rather than trusting any single number.
Community reports substantiate that caution. The open repository carries active issues reporting that "Punjabi Gurmukhi input appears to mishandle addak and nasalization markers", alongside a related enhancement request to improve Punjabi pronunciation through diacritic-aware transliteration or grapheme-to-phoneme handling. At the time of writing the repository showed 7 open and 684 closed issues — a healthy maintenance ratio, but evidence that per-language quality is uneven.
No pre-publication content check. The repository's own legal disclaimer states: "We do not hold any responsibility for any illegal usage of the codebase. Please refer to your local laws about DMCA and other related laws." Combined with the no-pre-clearance stance, responsibility is consistently placed downstream.
Reverse engineering and competing models are prohibited. The terms forbid attempts to obtain source code, underlying model components or algorithms, and separately forbid using output from the services to develop competing models.
Age requirements. Users must be at least 13, under-18s require parental permission, and the platform states it does not knowingly collect personally identifiable information from children under 16.
Directory metadata drifts fast. This listing's stored category was "fashion" with tags including "fashion" and "templates" — plainly wrong for a speech synthesis product — and its stored description cited 1,000+ voices where the site now claims over 2,000,000. Fast-moving products outrun their directory entries; check the site for current figures.
There is a genuine free tier requiring no credit card, providing 8,000 credits monthly, up to seven minutes of generation, a 500-character-per-generation cap and three public voice slots. Whether you may use that output commercially is genuinely unclear from the site: the pricing page's free-tier feature list includes "Commercial use", while the FAQ on that same page and the homepage FAQ both state free-tier output is for personal, non-commercial use only. Treat the restrictive reading as operative and confirm with the vendor before monetising anything made on the free plan.
The platform places that responsibility entirely on you and performs no advance check. Its cloning FAQ states you are responsible for confirming you have the rights, consents and disclosures required for any voice you clone, and for complying with laws covering name, likeness and AI-generated content in your region. It also states it does not pre-clear individual use cases and may remove content or accounts after the fact. Note that the binding terms of service contain no clause specifically requiring voice-cloning consent — only generic prohibitions on infringing others' rights — and that the same product page's marketing actively suggests cloning public figures. Legally, the exposure is yours.
The weights are genuinely published, but the licence is not permissive. Both the codebase and model weights are released under the Fish Audio Research License, which states in its introduction that it is intended to allow research and non-commercial use free of charge and that any commercial use requires a separate licence. Section III grants no commercial rights at all. Free for research, study and hobby use; a separate paid agreement for anything else.
Not under the public licence. The licence's definition of Commercial Purpose explicitly includes creating, modifying or distributing your product or service including via a hosted service or API, your organisation's internal operations, and any use connected to a product or service that generates revenue directly or indirectly. Internal business use alone triggers it. A separate written agreement is required, obtainable through the business contact address given in the licence. Note also that the research grant is revocable, unlike MIT or Apache grants.
Ten seconds of clean speech is the stated minimum, with the product page noting longer samples can help with very expressive voices. Reference quality matters more than duration — a clean short sample beats a long noisy one. The platform additionally states cloning is zero-shot cross-lingual across thirteen languages, so one recording can carry across supported languages without retraining.
The site gives four different answers and this article will not pick one for you. The cloning page says thirteen for cross-lingual cloning; the homepage says 30+; the TTS page's FAQ lists eight specific languages and advertises automatic support for eight with native accents; the funding announcement claims 83+. These likely describe different capabilities, but the site does not say which. Generate a test sample in your target language before committing — the open repository carries active bug reports about specific scripts, including Punjabi Gurmukhi handling of addak and nasalization markers.
Four paid tiers were listed at the time of checking, with promotional discounts stacked. Plus works out to $5.5/month billed at $66 annually against a $15 list price, with 250,000 credits monthly. Pro is $37.5/month billed at $450 annually against $100 list, with 2,000,000 credits and three team seats. Max is $749/month billed at $8,988 annually against $999 list, with 25,000,000 credits and ten team seats. Enterprise is custom, billed annually. Roughly 600 to 625 credits buys a minute of generation, and unused credits do not roll over.
The privacy policy does not say. A full-text check found no occurrence of "train", "voice" or "biometric" anywhere in the policy body, which carries an effective date of 28 August 2024. That means there is no published commitment either way on this point, and it would be wrong to assert a protection the document does not contain. Uploaded audio falls under the generic User Content category, whose listed sharing partners include service providers, advertising, analytics and business partners. If this matters to your compliance posture, ask the vendor directly — and note the enterprise tier advertises zero data retention as a distinct feature.
Deletion is partial rather than complete. The terms state that on account deletion the platform stops displaying your submissions to other users, but explicitly except public user submissions, which may remain fully available, and acknowledge it may not be possible to completely delete content from their records. Because licences granted are perpetual, sublicensable and irrevocable, and public submissions additionally grant marketing and promotional rights plus access rights to all other users, choosing a private rather than public voice slot is a meaningful decision rather than a filing preference.
The site footer names Hanabi AI Inc., while the open-source licence requires attribution to "39 AI, INC." — both names are documented, and the relationship between them is not explained on the site. The company announced $52 million in seed funding on its first anniversary, led by Coreline Ventures and Capital Today. Its own announcement reports a team of twenty-two, $21M in annual recurring revenue, 8M+ platform users and 2M+ community voice models, and names CEO Rissa Cao and chief scientist Shijia Liao. Those business figures are self-reported and not independently audited.