Toolso.AI
Toolso.AI
All ToolsCategoriesTrendingLatest ToolsPricingBlog
Toolso.AI
Toolso.AI
Toolso.AI
Toolso.AI

Discover the best AI tools to boost your productivity

GitHubGitHubTwitterX (Twitter)YouTubeYouTubeTikTokEmail

Popular Categories

  • AI Writing
  • AI Image
  • AI Video
  • AI Coding
  • More Categories

Explore

  • Latest Tools
  • Popular Tools
  • More Tools
  • Submit Tool
  • Pricing

About

  • About Us
  • Contact
  • Blog
  • Changelog

Legal

  • Cookie Policy
  • Privacy Policy
  • Terms of Service
  • Refund Policy
© 2026 Toolso.AI All Rights Reserved
Limited timeLimited-time offerFeatured Listing24h priority review · No backlink · 30 days featured$29.90then $59.90Price rises to $59.90 after Oct 31Ends in--:--:--Submit now
  1. Home
  2. All Tools
  3. AI Tool
  4. Typecast
Typecast interface previewVisit Website
Typecast logo

Typecast

Typecast is a Korean-built AI voice platform from Neosapience that turns text into expressive speech using 700+ licensed voice-actor models, Smart Emotion context analysis, voice cloning and a multilingual TTS API for creators, developers and enterprises.

AI ToolVoice GenerationAudio Generation#Text To Speech#Multilingual#Voice Cloning
Try for Free
Saves
Visits
Views
Pricing
Freemium
Published
Aug 25, 2026
Domain
typecast.ai
Community rating

Used this tool? Rate it

Rate this tool

Typecast Product Information

Try for Free
Tool Information
Saves
Visits
Views
Pricing
Freemium
Published
Aug 25, 2026
Domain
typecast.ai
Community rating

Used this tool? Rate it

Rate this tool

Featured Tools

Related Tools

Try for Free

What is Typecast?

Typecast is an AI voice generation platform that converts written text into speech with controllable emotional delivery. The Typecast AI voice proposition is not raw voice count or price but expressiveness: the company positions itself as "The world's most expressive AI voice generator" and describes the product as "The AI voice generator that sounds natural and full of emotion — for creators, developers, and enterprises." In practice this means three connected products under one subscription — a browser-based voice editor, a voice cloning system, and a text-to-speech API — plus a video editor and a talking-avatar feature layered on top.

The operator is Neosapience, Inc., a South Korean company. This matters more than a footnote for two reasons. First, it determines the legal frame: the terms of use state plainly that they "form an agreement between you and Neosapience, Inc." and that they are "governed by and construed in accordance with the laws of the Republic of Korea," with disputes falling under Korean civil procedure. Second, it explains the product's linguistic centre of gravity — Korean, English, Japanese and Chinese are first-class languages here in a way they are not for every Western competitor. The Korean-language site discloses the operating entity in full: 네오사피엔스 주식회사, business registration number 883-86-00767, representative Taesu Kim, with a registered address in the Samseong-dong area of Gangnam-gu, Seoul. The same individual is named as Chief Privacy Officer in the privacy policy.

What separates Typecast from a generic text-to-speech box is the emotion layer. Alongside conventional controls for speed and pitch, the platform offers what it calls Smart Emotion — a system that reads the surrounding text to infer how a line should be delivered rather than requiring you to tag each sentence by hand. The company frames its voice library as "Exclusive voices from real voice actors," and its usage policy makes the licensing arrangement behind that phrase unusually explicit for this category. Because voice licensing and consent are the questions that most often decide whether an AI voice tool is safe to build a business on, this page treats them as central rather than as an afterthought.

Core Features

Emotional text-to-speech with a large actor-sourced voice library

The headline capability is emotional text to speech: synthesis with adjustable emotional delivery. The site advertises "700+ AI voices with adjustable emotion, speed, and dynamics," and the on-page demonstrations expose the specific emotional states available — happy, sad, angry, whisper and low tone are all shown as selectable modes applied to the same line of text. Voices are organised into functional categories including announcer, narrator, podcast, news reporter, robot and character voices, alongside age and gender filters. The important qualifier is that these are not synthetic constructs assembled from scratch: the company describes them as "Exclusive voices from real voice actors," which is what makes the licensing terms discussed later meaningful.

Smart Emotion: context-inferred delivery

Smart Emotion is the feature the company leads with, and it is a genuine departure from tag-based emotional control. Rather than marking each sentence with a mood, the system analyses the text around a line to decide how it should be spoken. The company's own description of the underlying research is that it "analyzes scripts line by line to automatically generate context-appropriate tone." The API exposes this directly: a request can carry a previous_text and next_text alongside the line being synthesised, so that a sentence like "Everything is going to be okay" is delivered differently depending on whether it follows good news or bad. This is a meaningful design choice for narrative work, where the same words carry different weight depending on what precedes them.

Voice cloning in two tiers

The platform offers voice cloning at two distinct quality levels. Instant Cloning creates a custom voice from a short audio sample and becomes available on the entry paid tier; Professional Cloning, a higher-fidelity path, appears from the mid tier upward. Cloning slots are metered rather than unlimited — one slot on the entry and mid tiers, two on the popular tier, and ten on the business tier. The company states that a cloned voice can then "Speak multilingual," meaning a voice captured in one language can be used to generate speech in others, which is the practical draw for creators publishing across markets.

A text-to-speech API with streaming and timestamps

The developer product is not a thin wrapper around the web editor. The documentation describes an API providing speech in 37 languages via the ssfm-v30 model and access to more than 500 voices, with four distinct generation modes: standard text-to-speech producing complete WAV or MP3 files; streaming TTS that plays audio as chunks arrive, aimed at voice agents and low-latency interactive applications; timestamp TTS that returns word- or character-level alignment data for subtitles, karaoke highlighting and lip-sync; and instant cloning exposed programmatically. Official SDK examples are published for Python, JavaScript, C#, Java, Kotlin and Rust, and the docs additionally publish an llms.txt file explicitly intended for AI coding agents to read.

A video editor and talking avatars

Beyond audio, the subscription includes a video editor with export quality tiered by plan — 720p on the free tier, 1080p on entry, and 4K from the mid tier upward — with watermark-free export gated behind any paid plan. A separate AI Talking Avatar feature generates a speaking on-screen presenter, metered by generation count rather than duration: up to 6 generations on the entry tier rising to 40 on business, with zero available on the free tier. Note that avatars carry a stricter age gate than the rest of the platform, discussed below.

Mobile applications with cross-device sync

Typecast publishes mobile applications for iOS and Android that generate voiceovers directly from a phone and synchronise projects across devices. The iOS listing is published by Neosapience, Inc., is free to download, requires iOS 16.4 or later, carries a 12+ content rating, and was updated in August 2026 — an actively maintained build rather than an abandoned companion app.

Nine years of speech research behind the model

The company publishes a model lineage rather than treating its synthesis engine as a black box: a first deep-learning speech synthesis model in 2018, the CATS model improving pronunciation and naturalness in 2021, SSFM 1.0 in 2024 introducing a new architecture trained on large-scale datasets, SSFM 2.0 in 2025 adding multilingual support and advanced voice controls, and SSFM 3.0 in 2025 bringing context-aware emotion generation. This lineage is worth weighing when comparing against newer entrants, though the claims are the company's own and are not independently audited.

Use Cases

Narrative and character work where emotional range is the point

The clearest fit is content where flat delivery would fail: audio drama, animation, game dialogue, audiobook fiction and story-driven video. One testimonial on the site, from a film director, captures the intended workflow — the ability to try different voices, adjust timing and play with emotion until a scene lands, described as "casting, directing, and producing all at once." The combination of a large character-voice catalogue with per-line emotional control is what makes this category viable rather than merely possible.

Multilingual publishing from a single recorded voice

For creators publishing the same content across markets, the cloning-plus-multilingual combination is the substantive draw: capture a voice once, then generate that voice speaking other languages. With 37 languages available through the API and Korean, English, Japanese, Chinese, Spanish and Vietnamese named explicitly in the documentation, the coverage is genuinely broad — and notably strong in East Asian languages where some Western-built competitors are weaker.

Voice agents and real-time applications

Streaming TTS exists specifically for cases where waiting for full synthesis is unacceptable. The company has stated its direction is "evolving from simple voice generation into true conversational AI that listens, thinks, and speaks in real time," and the enterprise tier explicitly references real-time conversational agents. For developers building assistants or interactive voice products, the streaming endpoint plus low-latency playback is the relevant surface, not the web editor.

Subtitles, dubbing and lip-sync pipelines

Timestamp TTS returns word- or character-level alignment data, which turns the output into something a production pipeline can consume rather than a bare audio file. This is what makes automated subtitle generation, karaoke-style highlighting and lip-sync animation practical without a separate forced-alignment step.

E-learning, corporate narration and product content

The company describes a user base spanning broadcasting, telecom, e-commerce and education, and the plan structure reflects institutional use: the business tier adds team members, ten cloning slots and additional credit purchasing. One licensing detail matters enormously here and is easy to miss — the usage policy states that "separate branches or companies must each have their own subscriptions to ensure compliance," so a single seat cannot lawfully serve a whole corporate group.

Rapid iteration before committing to human recording

Because generation and playback are unlimited on every plan including the free one, and credits are consumed only on download, Typecast supports a workflow that most competitors do not: audition an entire script across many voices and emotional readings at zero cost, and pay only for the takes you keep. Used deliberately, this makes it a casting and previsualisation tool as much as a production one.

How to use Typecast

Step 1: Audition freely before spending anything

Understand the credit model first, because it shapes everything. The pricing page states that "Voice generation and playback are free on every plan" and that "Credits are used when you download audio." Nothing is spent while you experiment in the editor. Generate the same line across a dozen voices, try each emotional mode, adjust pacing — none of it consumes credits until you export. Most users who feel the free tier is stingy have simply not internalised that the limit is on downloads, not on trying things.

Step 2: Choose a plan against the specific control you actually need

The tiers are not a simple quality ladder; each unlocks a specific control. The free tier gives 3,000 lifetime download credits (roughly five minutes) but restricts you to trial voices only, 16 kHz audio and three projects. The entry tier at $5 per month adds access to all voices, 44.1 kHz audio, one instant-cloning slot and a commercial licence. The next tier at $19 adds speed control and professional cloning. The popular tier at $29 is where emotion controls live, including Smart Emotion, custom emotion, intonation and pitch — so if emotional delivery is the reason you came to Typecast, that is your entry price, not $5. The business tier at $69 adds ten cloning slots, team members and additional credit purchases. Annual billing saves ten per cent.

Step 3: Write for Smart Emotion rather than fighting it

If you are on a tier with Smart Emotion, feed it context. Because the system infers delivery from surrounding text, splitting a scene into isolated one-line requests throws away the very signal it uses. Through the API this is explicit — supply previous_text and next_text — but the same principle applies in the editor: keep narrative passages intact so the model can read the arc rather than treating each line as an orphan.

Step 4: Verify the licence tier before publishing commercially

This is the step people skip and later regret. Commercial licensing is not uniform across tiers. On the free plan the pricing table marks the commercial licence as "Attribution Required," and the plan description states that "Attribution is required for all content downloaded on the Free plan." From the entry tier upward, attribution becomes "Optional." If you are publishing monetised content, confirm which tier your downloads were made under before distribution, not after.

Step 5: Download and archive everything you intend to keep

The retention rule is the single most consequential operational detail on this page. Output is kept on Typecast's servers for one year from creation, after which it is "automatically deleted from our servers without further notice." The terms state that "You are responsible for downloading and saving any Output you wish to retain before the Retention Period expires" and that the company has "no obligation to recover or restore" deleted output. Crucially, this "applies equally to all users, regardless of whether you are on a paid or free plan." Treat the platform as a generation service, not as your archive.

Step 6: Maintain a subscription if you need continued production rights

The usage policy sets a boundary that catches people after cancellation: "Generated audio can be downloaded and used for personal or commercial purposes solely during your subscription period. After expiration, you may continue using previously downloaded audio, but no new modifications or downloads can be made unless your subscription is extended." Already-downloaded audio remains usable, but the ability to re-download or derive new content stops with the subscription.

Step 7: Integrate via the API when the editor stops scaling

For programmatic use, obtain an API key and pick the mode that matches the workload: standard TTS for finished files, streaming for interactive latency-sensitive applications, timestamp TTS when you need alignment data. Official SDKs cover Python, JavaScript, C#, Java, Kotlin and Rust. Note that the usage policy requires separate agreements for API services, custom voice generation and integration into automated systems, so an API deployment is not simply a bigger version of a consumer subscription.

Tips & Best Practices

Treat generation as free and downloads as the real budget

Build your working habits around the credit model. Because playback costs nothing, do all comparison, direction and revision in the editor and export only final takes. A rough sense of scale helps: 30,000 credits corresponds to about 35 minutes of audio, 75,000 to about 90 minutes, and 200,000 to about 250 minutes — so credits map to finished output length, not to effort spent getting there.

Buy the tier that contains the one feature you need most

Because each tier gates a different control, the cheapest adequate plan is often not the cheapest plan. Speed control does not appear until $19; emotion controls including Smart Emotion do not appear until $29. Someone who subscribes at $5 expecting the expressive control the marketing showcases will be disappointed — not because the feature does not exist, but because it sits two tiers up.

Check the store you subscribe through

The pricing page carries a warning worth heeding: "Prices on the App Store and Google Play Store may differ from the amount shown above, based on each store's pricing policies," and "Subscriptions purchased via the mobile app (iOS/Android) cannot be modified on the web. Please manage your plan through the respective app store." Subscribing on mobile locks your plan management into that store.

Understand the refund conditions before your first download

The refund terms are strict and tie directly to usage. For monthly plans, "A full refund is available only if there is no history of credit usage, no history of Premium Cloning creation, and no pending tasks within 7 days of the billing date." Annual plans follow the same conditions within seven days of payment, after which a refund deducts the service fee for months used plus a processing fee. In effect, the first download or clone forfeits the clean refund. EU users have an additional statutory 14-day withdrawal right, though the terms note this is waived once services have been partially performed.

Never clone a voice you do not have rights to

The usage policy is explicit that "Unauthorized use of a person's identity, voice, or likeness for non-satirical, malicious, or commercial purposes without their consent" is prohibited, as is unauthorised impersonation of political figures. Beyond the platform's own rules, the terms require that you warrant you hold "all rights, licenses, consents, permissions, and/or authority" for content you upload. Cloning consent is your legal responsibility, not the platform's.

Keep political and regulated content off the platform unless cleared

Political advertising is prohibited "unless express written permission is granted," and election misinformation and impersonation of political figures or their family members are banned outright. Adult content, and use of AI voices in connection with pornography or adult services, is prohibited. If your work touches these areas, seek written clearance first rather than discovering the restriction after publication.

Plan for team structure before scaling

Because "separate branches or companies must each have their own subscriptions," an agency serving multiple client entities or a group with several legal subsidiaries cannot compliantly consolidate onto one seat. Work out the licensing structure at procurement rather than at audit.

Who is Typecast for?

A good fit

Typecast fits creators whose content depends on emotional delivery rather than mere intelligibility — audio fiction, animation, character work, story-driven video — because the emotion layer is where the product invests. It fits teams publishing across East Asian and Western markets, since Korean, Japanese and Chinese are core rather than afterthoughts. It fits developers building voice agents or subtitle pipelines, given streaming and timestamp endpoints and SDKs across six languages. It fits anyone who wants to audition extensively before paying, given unlimited free generation and playback. And it fits users who want an unusually clear answer on voice licensing, since the actor-consent position is published rather than implied.

A poor fit

It is a poor fit for anyone who needs the platform to serve as long-term storage: the one-year auto-deletion applies regardless of plan. It is a poor fit for organisations that want a single licence to cover multiple legal entities, given the separate-subscription requirement. It is a poor fit for adult or political content, both restricted by policy. It is a poor fit for users under 13 anywhere, and under 17 for the Talking Avatar feature specifically. It is a poor fit for buyers who need a large, mature body of independent reviews before committing, since the third-party review footprint is thin relative to the company's claimed user base. And it may be a poor fit for those uncomfortable with Korean governing law and a class-action waiver, discussed in the limitations.

The decision that actually matters

For most prospective users the real question is narrower than "is Typecast good." It is whether emotional control is worth the tier premium. If flat, clear narration suffices, cheaper options abound and the entry tier is arguably overpriced relative to competitors. If delivery is the product — if the difference between a whisper and a shout carries your content — then the $29 tier is the honest comparison point, and Smart Emotion is a genuinely differentiated capability rather than a repackaged pitch control.

Platforms

Typecast runs primarily in the browser: the voice editor, video editor and cloning interface are all web-based, with no desktop application. Mobile applications exist for both iOS and Android and are positioned for generating voiceovers on the go with project synchronisation across devices. The iOS application, published by Neosapience, Inc., is free to download, requires iOS 16.4 or later, carries a 12+ content rating, is listed under Photo & Video and Entertainment, and supports English and Korean interface languages. Its most recent version was released in August 2026, indicating active maintenance.

For developers, the platform surface is the TTS API rather than any installed client, with official SDKs for Python, JavaScript, C#, Java, Kotlin and Rust, plus a command-line tool and documentation designed to be consumed by AI coding agents. The web presence itself is bilingual, with separate English and Korean sites, and the Korean site carries the full statutory business disclosure that Korean e-commerce regulation requires.

One platform caveat deserves emphasis: because subscriptions bought through the mobile app stores cannot be managed on the web, the platform you subscribe on determines where you must later cancel or change plans.

Pricing & Plans

Typecast prices by monthly subscription with a ten per cent discount for annual billing, and the structure is built around download credits rather than generation time. Every tier including the free one offers unlimited voice generation and playback; credits are consumed only when audio is downloaded.

The free tier costs nothing and provides 3,000 lifetime download credits, equivalent to roughly five minutes of audio. It restricts downloads to trial voices only, caps audio at standard 16 kHz quality, allows three projects, provides 720p video export with 1 GB of media storage, retains download history for seven days, and requires attribution for all downloaded content. It includes no cloning slots and no avatar generations.

The Basic tier at $5 per month, or $54 billed yearly, provides 30,000 monthly credits (around 35 minutes), unlocks all AI voices, raises audio to 44.1 kHz, adds one instant-cloning slot, enables 1080p watermark-free video export with 5 GB storage, and makes attribution optional under a commercial licence.

The Plus tier at $19 per month, or $204 yearly, provides 40,000 credits (around 50 minutes), adds speed control and professional cloning, and raises video export to 4K with 30 GB of storage.

The Pro tier at $29 per month, or $312 yearly, is where the platform's signature capability lives. It provides 75,000 credits (around 90 minutes) and adds emotion controls including Smart Emotion's one-click adjustment, custom emotion, intonation and pitch, plus a second cloning slot and 50 GB of storage.

The Business tier at $69 per month, or $744 yearly, provides 200,000 credits (around 250 minutes), permits additional credit purchases, and adds ten cloning slots, team members and 100 GB of storage. Above this sits an Enterprise tier quoted on inquiry, covering a dedicated API and security package, real-time conversational agents and a dedicated account manager.

Two practical notes accompany the table. Mobile store pricing may differ from the web figures, and mobile subscriptions must be managed through the originating store. And refunds are conditional: a full refund requires no credit usage, no Premium Cloning creation and no pending tasks within seven days of billing, with annual plans past that window refunded net of months used and a processing fee.

Alternatives

ElevenLabs is the most frequently compared alternative and the strongest competitor on raw voice realism and ecosystem breadth. The honest distinction is emphasis: Typecast's investment is concentrated in context-inferred emotional delivery and an explicitly documented actor-licensing position, while ElevenLabs competes on model quality, scale and a wider surrounding toolset. Buyers should compare on the specific axis they care about rather than on general reputation.

Murf and WellSaid Labs target corporate and e-learning narration, where consistency and pronunciation control matter more than dramatic range. For product explainers, training modules and internal communications, these are often the more economical fit, and the emotional sophistication Typecast charges for goes largely unused.

Play.ht and Speechify occupy the volume-and-price end, appealing to publishers converting large text libraries to audio. Their advantage is throughput economics; Typecast's counter-argument is per-line direction, which matters little when the task is bulk conversion.

The major cloud providers — Google, Amazon and Microsoft — offer speech synthesis with deep infrastructure integration and enterprise procurement paths. They are the default when the requirement is embedded, high-volume synthesis inside an existing cloud estate. What they generally do not offer is a curated catalogue of licensed actor voices with published consent commitments, which is precisely Typecast's differentiator.

For Korean, Japanese and Chinese content specifically, Typecast's regional origin is a genuine advantage worth testing directly. Rather than accepting any comparison at face value, generate the same script in your target language on two platforms and listen — for East Asian languages the gap frequently runs the opposite way to the general English-language reputation ranking.

Limitations & Considerations

Output is deleted after one year, on every plan

The most consequential limitation is not a feature gap but a retention rule. Output generated through the service is retained for one year from creation, after which it is "automatically deleted from our servers without further notice." The terms make three points explicit: you are responsible for downloading anything you wish to keep; the company has no obligation to recover deleted output; and the policy "applies equally to all users, regardless of whether you are on a paid or free plan." Re-downloading previously generated output is only possible while it remains on the servers and while you hold an active subscription.

Commercial rights are tied to an active subscription

Downloaded audio remains usable after cancellation, but the usage policy states that generated audio may be used commercially "solely during your subscription period," and that after expiration "no new modifications or downloads can be made unless your subscription is extended." Combined with the retention rule, lapsing a subscription can permanently strand work you had not yet exported.

The free tier requires attribution and restricts voices

Free-tier output is not unconditionally usable: attribution is required for everything downloaded on that plan, downloads are limited to trial voices, and audio quality is capped at 16 kHz. The free tier is a genuine evaluation environment, not a free production tool.

Emotion controls — the headline feature — sit behind the $29 tier

The marketing leads with expressiveness, but basic emotion controls, Smart Emotion, custom emotion, intonation and pitch all appear only from the Pro tier. Speed control requires at least the $19 tier. A buyer drawn by the emotional demonstrations who subscribes at $5 will not receive the capability that attracted them.

You grant the company a broad, irrevocable licence over uploaded content

While the terms confirm that "you retain your ownership rights in Your Content and own the Output," uploading content grants Neosapience "a non-exclusive, irrevocable, worldwide, fully-paid, royalty-free license to store, reproduce, modify, transmit, display, publish, and distribute Your Content to operate, improve, promote and provide the Services." Read that scope carefully before uploading sensitive voice samples — it includes promotion, and it is irrevocable.

Korean governing law, a class-action waiver and a one-year claim window

Disputes are governed by the laws of the Republic of Korea and fall under Korean civil procedure. The terms further state that any dispute "SHALL BE RESOLVED INDIVIDUALLY, WITHOUT RESORTING TO ANY FORM OF CLASS ACTION," and that any cause of action "MUST COMMENCE WITHIN ONE (1) YEAR AFTER THE CAUSE OF ACTION ACCRUES." For non-Korean business users these are material terms, not boilerplate.

Refunds are effectively forfeited on first use

Because a full refund requires no credit usage, no Premium Cloning creation and no pending tasks, downloading a single file or creating one premium clone ends the clean-refund path. Given that generation and playback are free, this is navigable — but only if you evaluate before downloading.

One subscription cannot cover multiple entities

The requirement that "separate branches or companies must each have their own subscriptions" limits how agencies and corporate groups can deploy the platform, and separate agreements are additionally required for API services, custom voice generation and integration into automated systems.

Age restrictions and content prohibitions are broad

Users must be at least 13, or older where local law requires, and the Talking Avatar feature is restricted to users 17 and over. Prohibited content spans sexual material, violence and hate speech, political advertising and election misinformation, fraud and impersonation, child safety violations, and intellectual property or privacy infringement. The company reserves the right to remove generated content "for any reason (or no reason)" and to terminate access without prior notice.

Independent review evidence is thin relative to the claimed user base

This is a genuine evidentiary gap rather than a criticism of the product. The Trustpilot profile for the domain is unclaimed and carried only a single review with a 3.2 score when checked — a sample far too small to support any quality conclusion, and Trustpilot itself notes the company "hasn't invited their customers, so reviews may not be representative." The iOS application showed a 3.33 average from just three ratings, likewise below any threshold at which an average means anything. Attempts to verify ratings on G2 and Capterra were blocked on every available channel, so figures for those platforms circulating in secondary sources are not reproduced here. Anyone evaluating Typecast should test it directly rather than relying on aggregate scores.

Company-reported figures are not independently audited

Claims of 700+ voices, 35+ languages on the marketing site (against 37 in the documentation), and user numbers reported in the press are the company's own. Note that the user figure has moved substantially over time in the company's own communications: press coverage of the 2022 Series B reported "more than 1 million users," while coverage of the December 2025 pre-IPO round reported "more than 2 million registered users worldwide." Registered users are not active users, and neither figure has been independently verified.

Funding and corporate trajectory

For context on stability rather than as an endorsement: the company raised a $21.5 million Series B in February 2022 led by BRV Capital Management, bringing total funding to roughly $26.7 million at that point, and raised a further $11.5 million in a pre-IPO round announced in December 2025 with participation from HB Investment and K2 Investment Partners. The company was founded in 2017 by former Qualcomm engineers. A pre-IPO round signals intent toward a public listing, which typically brings more disclosure but also strategic change; neither is guaranteed.

FAQ

Q1. Is Typecast free to use, and what does the free plan actually allow?

There is a genuinely usable free plan, but its limits sit in a specific place. Voice generation and playback are unlimited on every plan including the free one — you can experiment endlessly at no cost. What is limited is downloading: the free tier provides 3,000 lifetime download credits, roughly five minutes of audio, and those credits are lifetime rather than monthly. Free downloads are restricted to trial voices, capped at 16 kHz audio, limited to three projects and 720p video with 1 GB of storage, and require attribution on all downloaded content. It is an evaluation tier rather than a free production tier.

Q2. Whose voices are these, and did the voice actors consent?

This is the clearest area of the platform's policy and a genuine differentiator. The usage policy states that "Voice actors retain ownership of their voice models" and that "Unauthorized replication, distribution, or alteration of a voice without permission is strictly prohibited." On consent specifically it states that the company ensures "every voice actor whose voice model is available on the platform has explicitly agreed to the use cases for their voice." It further provides that generated audio "may not be altered to misrepresent the voice actor or be used for content they have not approved," with misuse subject to removal and potential legal action. The company markets these as "Exclusive voices from real voice actors." Note that these are the company's published commitments; the underlying contracts with individual actors are not public, and compensation terms are not disclosed in any source examined.

Q3. Can I use Typecast output commercially?

Yes, with conditions that vary by tier and by subscription status. A commercial licence is included from the Basic tier upward with attribution optional; on the free plan attribution is required for all downloaded content. The usage policy grants "a non-exclusive, non-transferable right to use the audio content generated through our services for personal or commercial purposes," but ties it to time: use is permitted "solely during your subscription period," and after expiration you may keep using already-downloaded audio while no new downloads or modifications are possible. Separate branches or companies must each hold their own subscription.

Q4. What happens to my generated audio over time?

It is deleted. Output is retained on the servers for one year from creation and then "automatically deleted from our servers without further notice." This applies identically to paid and free users. Deletion affects only server-side copies — files you already downloaded remain yours — but the terms state clearly that you are responsible for saving what you want to keep, and that the company has no obligation to restore anything deleted. Re-downloading is only possible while output is still retained and while you hold an active subscription.

Q5. How much does Typecast cost and which tier do I actually need?

Plans run from free to $69 per month, with ten per cent off annual billing: Basic $5 (30,000 credits, all voices, one instant-cloning slot, commercial licence), Plus $19 (40,000 credits, speed control, professional cloning, 4K video), Pro $29 (75,000 credits, emotion controls including Smart Emotion, intonation and pitch, two cloning slots), Business $69 (200,000 credits, ten cloning slots, team members, extra credit purchases), plus Enterprise on inquiry. The tier you need depends on the control you want: emotion controls begin at $29 and speed control at $19, so the expressiveness the product is marketed on is a Pro-tier feature.

Q6. What is Smart Emotion and how is it different from picking a mood?

Smart Emotion infers delivery from context rather than requiring you to label each line. The system reads the surrounding text — through the API you pass the preceding and following lines explicitly — and generates a reading appropriate to that context, so identical words are delivered differently depending on what surrounds them. The company describes the underlying approach as analysing scripts line by line to produce context-appropriate tone. Conventional per-line emotion tags remain available as custom emotion, alongside intonation and pitch controls, from the Pro tier.

Q7. Which languages are supported and how many voices are there?

The marketing site advertises 700+ AI voices and 35+ languages; the API documentation states 37 languages via the ssfm-v30 model and more than 500 voices available programmatically. The difference is best explained by the web catalogue and the API catalogue not being identical, so verify availability for your specific language and voice through the tier and interface you plan to use. Languages named explicitly in the documentation include Korean, English, Japanese, Chinese, Spanish and Vietnamese.

Q8. Can I clone my own voice, and can the clone speak other languages?

Yes to both, from paid tiers. Instant Cloning creates a voice from a short sample and is available from Basic with one slot; Professional Cloning, a higher-fidelity option, appears from Plus. Slot counts rise to two on Pro and ten on Business. The company states that a cloned voice can speak multiple languages, which is the main practical reason to clone rather than pick from the catalogue. You are responsible for holding the rights to any voice you clone — the terms require you to warrant that you have all necessary rights, licences and consents, and cloning someone else's voice without permission is prohibited.

Q9. Who operates Typecast and under what law?

Typecast is operated by Neosapience, Inc., a South Korean company founded in 2017 by former Qualcomm engineers. The Korean-language site discloses the entity as 네오사피엔스 주식회사 with business registration number 883-86-00767, representative Taesu Kim, and an address in Gangnam-gu, Seoul; the same person is named Chief Privacy Officer. The terms are governed by the laws of the Republic of Korea with disputes under Korean civil procedure, and include a class-action waiver and a one-year limitation period for bringing claims. The privacy policy states compliance with Korean data protection law alongside the EU and UK data protection regimes, Californian privacy law and children's online privacy rules.

Q10. Are there content restrictions I should know about before building on it?

Yes, and several are broader than users expect. Sexual and adult content is prohibited, including using AI voices in connection with adult services. Political advertising is prohibited unless express written permission is granted, as is election misinformation and impersonation of political figures or their family members. Fraud, deception, unauthorised impersonation of any individual, hate speech, harassment, and intellectual property or privacy infringement are all banned. Users must be at least 13, and the Talking Avatar feature requires users to be 17 or older. The company reserves the right to review and remove generated content at its sole discretion and to terminate accounts without prior notice.

Know a Similar Tool?
If you know other great AI tools, feel free to submit them to us