Used this tool? Rate it
Used this tool? Rate it
Wan is Alibaba's family of video generation models, and it exists in two forms that are easy to confuse. One is a set of open model weights published on GitHub and Hugging Face under the Apache 2.0 licence, which anyone can download and run on their own hardware. The other is a hosted creative studio at wan.video, where you pay a subscription and generate in the browser without owning a GPU. Most tools in this category are one or the other. Wan is both, and which one you should care about depends entirely on whether you have a graphics card with enough memory.
The attribution is worth establishing up front, because domains bearing the names of well-known open models are frequently squatted by third parties. This one is not. The footer of wan.video reads "Wan © 2026 Powered by Alibaba Cloud" and lists wan_ai@service.alibaba.com as the contact address. More conclusively, the site's front-end assets are served from Alibaba's own content delivery network under the Tongyi Lab namespace — the stylesheet and application bundle both load from g.alicdn.com/tongyi-lab-fe/wan/ — its analytics identifier is set to tongyiwanxiang, the Chinese brand name for Wan, and its tracking script is an internal Alibaba npm package. No third party can deploy code into a company's own CDN namespace. This is the first-party site.
On the open side, the official GitHub organisation is Wan-Video and the official Hugging Face account is Wan-AI. The Wan2.2 repository describes itself as "Wan: Open and Advanced Large-Scale Video Generative Models" and holds 17,280 stars against 2,191 forks; the earlier Wan2.1 repository holds 16,888 stars and 3,450 forks. Both are licensed Apache-2.0, as are the organisation's derivative repositories including Wan-Animate-2, Wan-Dancer and Wan-skills, the last of which is described as agent skills that let an AI agent invoke Wan's generation capabilities.
The hosted product has moved faster than most third-party listings have kept up with. The site's current headline offering is Wan3.0 — the primary call to action reads "Try Wan3.0 Now" and the feature section is titled "Introducing Wan3.0" — while the open weights most widely deployed remain the 2.1 and 2.2 generations. If you have read a description of Wan as simply a text-to-image and text-to-video platform, that description predates the current release.
The open weights and the hosted studio are not the same product with different packaging. They differ in licence, in capability, and in what you are allowed to do with the output.
The weights give you Apache 2.0 terms, which permit commercial use, modification and redistribution provided you retain the copyright notice and licence text. The Hugging Face model card goes further and states plainly: "We claim no rights over the your generated contents, granting you the freedom to use them while ensuring that your usage complies with the provisions of this license." That is an unusually clean position on output ownership, and the typo in the original is reproduced here as published.
The hosted studio operates under its own subscription terms, and those terms could not be retrieved during this research — a limitation discussed in full below. Do not assume the Apache 2.0 output position automatically carries across to content generated on create.wan.video.
The site lists five video-side capabilities for Wan3.0, each with a specific claim rather than a vague one. Omni-Creation accepts up to 20 reference assets and can parse documents and web pages as input. Native 30s Duration generates at thirty seconds natively rather than by stitching shorter clips, which the company frames as enabling more complete storytelling. Pixel-Perfect Consistency is positioned around delivery certainty for production workflows — notably a production framing rather than a creative one. Immersive Experience covers realism, texture and sound design. Precision Video Editing covers controllable editing.
Five further capabilities address images. Portrait Customization works down to bone structure and eye detail. Advanced Text Rendering supports long-form text across twelve languages and can generate charts, formulas and infographics — text rendering being a long-standing weakness of image models, this is a pointed claim. Precise Color Control governs colour distribution. Multi-Image Editing fuses up to nine images into a single output. Sequential Storytelling generates up to twelve sequential images with consistent style and subject. Interactive Editing lets you box-select regions to align intent at the pixel level.
The creative console at create.wan.video carries eight navigation modules: Explore, Generate, Canvas, Skills, Playground, TimeLine, Assets and Favorites. The presence of Canvas and TimeLine in particular indicates a composition and sequencing environment rather than a single generate-and-download interaction. There are also dedicated CLI and API entry points in the console header.
Wan2.2 ships five variants, and the parameter counts matter because they determine what hardware you need. T2V-A14B is the text-to-video model, built as a mixture-of-experts with 27 billion total parameters and roughly 14 billion active per inference step. I2V-A14B is the image-to-video model on the same architecture. TI2V-5B is a 5-billion-parameter dense model that unifies both text and image-to-video in one checkpoint. S2V-14B handles speech-to-video. Animate-14B handles character animation and replacement. The models generate at 480P and 720P at 24fps.
The mixture-of-experts design is the architectural point of interest: rather than one model handling the whole denoising process, there are two specialised experts, one for high-noise early stages and one for low-noise refinement. ComfyUI's official documentation, which maintains a native workflow for the model, summarises the resulting strengths as cinematic-level aesthetic control with multi-dimensional command over lighting, colour and composition; large-scale complex motion with improved smoothness and controllability; and precise semantic compliance in complex multi-object scenes. It also notes the 5B version uses a high compression ratio VAE for memory optimisation.
The repository is direct about what it takes to run these models. For the A14B models, the documented command "can run on a GPU with at least 80GB VRAM." For TI2V-5B, the model "can run on consumer-grade graphics cards like 4090" and needs "at least 24GB VRAM (e.g, RTX 4090 GPU)."
That is a wide gap, and it has visibly shaped real-world adoption. Hugging Face's own API shows Wan2.1-T2V-1.3B-Diffusers leading downloads at 204,432 over thirty days, with Wan2.2-TI2V-5B-Diffusers close behind at 198,975. The flagship A14B models trail: 134,807 for image-to-video and 132,815 for text-to-video. In other words, the smallest and second-smallest models are downloaded roughly one and a half to three times as often as the largest ones. The binding constraint on this technology is not model quality; it is the graphics card sitting in the user's machine.
The site treats the API as a distinct third product line, given equal billing with Features and Open Source in the top navigation and its own footer group containing Platform Overview and API Access. One important commercial detail is documented on the pricing page: a studio subscription covers model creation on create.wan.video only, and the API and other products require separate purchases. The specific API pricing and call specifications were not retrievable during this research.
Self-hosted video generation. The clearest case for the open weights. If you have a 24GB card, TI2V-5B is reachable; if you have data-centre hardware, the A14B models are. Apache 2.0 licensing means the output and the deployment are yours, subject to the stated content restrictions.
Building on top of the models. Because the weights are downloadable and the licence permits modification and redistribution, the models can be fine-tuned, quantised, wrapped in a product, or integrated into a proprietary system. The fork counts — 3,450 on Wan2.1 alone — indicate this is happening at scale.
ComfyUI and node-based workflows. ComfyUI maintains an official native workflow for Wan2.2, which makes the models accessible to the large population of creators who work in node graphs rather than command lines.
Evaluation without hardware. The hosted studio's free tier exists precisely for this. You can assess output quality before deciding whether to invest in a GPU or a subscription.
Production video work at modest volume. The paid tiers unlock 1080p output, videos in the ten-to-thirty-second range, and watermark-free downloads — the three things that separate a demo from something you can deliver.
Multilingual text-in-image work. Advanced Text Rendering across twelve languages, including charts and infographics, targets a use case that most image models handle poorly.
Agent-driven generation. The Wan-skills repository packages the generation capabilities as agent skills, and the studio exposes a CLI, which together suit programmatic and automated pipelines.
Start with the 5B model even if you have the hardware for more. The download figures suggest the community has converged on it for good reason: it unifies text and image-to-video in one checkpoint and runs on a single consumer card. Establish your workflow there before scaling up.
Prompt in Chinese or English deliberately. The models were built with bilingual prompting in both languages. If your subject matter has stronger representation in one language than the other, that is worth testing rather than assuming.
Use image-to-video for control. Starting from an image you already control removes an entire axis of variability compared with text-to-video, which is why I2V-A14B carries the highest like count of any model in the account.
Do not rely on the free tier for anything you will publish. Watermark-free downloads begin at the paid tiers. Output generated to evaluate quality is not output you can ship.
Be deliberate about the first billing period. The pricing page states that subscriptions auto-renew and can be cancelled at any time, but that the first month or year is non-refundable once activated. Choose monthly for a first trial unless you are already confident.
Do not assume the studio inherits the open licence. The Apache 2.0 terms and the "no rights over your generated contents" statement come from the model repositories. The hosted service has its own terms, which were not retrievable here.
Retain the licence text if you redistribute. Apache 2.0 permits commercial use, modification and redistribution, but conditions this on retaining the original copyright notice and licence text. This is easy to overlook when packaging a model into a product.
Read the content restrictions before deploying at scale. The model card prohibits content that violates applicable laws, causes harm to individuals or groups, disseminates personal information intended for harm, spreads misinformation, or targets vulnerable populations, and holds users fully accountable for compliance.
Developers and teams with GPU access get the most distinctive value here. Open weights under a permissive licence with an explicit disclaimer of output rights is a combination few video models offer.
Researchers and fine-tuners are well served by the variant spread — five task-specific models plus separate animation and dancing repositories give plenty of starting points.
ComfyUI users have an officially supported path, which is not true of every model that gets ported to the ecosystem.
Creators without hardware can use the hosted studio, though they should evaluate it against dedicated hosted competitors on price and output rather than assuming the open-source pedigree translates into a better hosted product.
Agent and automation builders have the CLI, the API line and the Wan-skills repository to work with.
Who should look elsewhere: anyone expecting to self-host the flagship models on consumer hardware, since 80GB is a data-centre requirement; anyone who needs the hosted service's terms of use reviewed before committing, because those documents were not reachable; anyone in a region where the mobile app is unavailable and who requires a native app rather than mobile web; and anyone who needs a single vendor relationship covering studio, API and everything else, since these are separately purchased.
Self-hosted. Any machine meeting the VRAM threshold, via GitHub and Hugging Face. This is the most capable and least constrained way to use Wan.
Web studio. create.wan.video, with the eight-module console described above. No local hardware requirement.
Mobile. The site offers App Store and Google Play links but attaches an explicit caveat: "App not available in your region? Try our mobile web version." That regional limitation is real — searching the United States App Store across three keyword variations returns no Alibaba-published Wan application. What comes back instead are unrelated products and a couple of third-party apps using the name, each with one or two ratings, which are far too small a sample to say anything about and should not be mistaken for the official product.
API. Available as a separate product line with its own purchase requirement.
ComfyUI. Officially documented native workflow support.
Two things are being priced here, and they should not be conflated.
The open weights are free. Apache 2.0, downloadable, commercially usable. Your cost is hardware and electricity.
The hosted studio runs three membership tiers. Prices are shown per month with annual billing advertised at 50% off:
| Tier | Yearly (per month) | Monthly | Credits/month | Concurrent video | Concurrent image |
|---|---|---|---|---|---|
| Free | $0 | $0 | — | 1 | 1 |
| Pro | $5 | $10 | 300 | 3 | 3 |
| Premium | $20 | $40 | 1,200 | 8 | 5 |
The company provides conversion guidance so credits can be reasoned about: Pro's 300 credits work out to up to 1,200 images or 60 videos, and Premium's 1,200 credits to up to 4,800 images or 240 videos. Separate tabs on the pricing page sell Credits and Gift Cards outside the membership structure.
What free excludes is as important as what the paid tiers include. Free is limited to six image styles rather than the full set, cannot download watermark-free images or videos, cannot create 1080p video, and cannot create the longer ten-to-thirty-second videos. Paid tiers add all of those plus image upscaling. Every tier can earn credits through daily check-in.
Two payment terms deserve attention. Subscriptions auto-renew and can be cancelled at any time, but the first month or year is non-refundable once activated. And the subscription covers model creation on create.wan.video only — the API, DealDance and other products require separate purchases.
Hosted-only video generators are the natural comparison for the studio. If you have no intention of running weights locally, evaluate Wan's studio purely on output quality, price and rights terms against other hosted services; its open-source lineage confers no advantage in that comparison.
Other open-weight video models are the comparison that matters if you do self-host. The questions to ask are licence terms, VRAM requirements at each quality level, and ecosystem support. Wan scores well on the first — Apache 2.0 with an explicit output-rights disclaimer is at the permissive end — and its ComfyUI support is officially maintained.
Running the smaller Wan variants is itself the alternative to running the flagship ones. Given that TI2V-5B is downloaded nearly as often as the 1.3B model and considerably more than the A14B pair, the community's implicit verdict is that the quality-to-hardware trade at 5B is the sensible default.
Alibaba's other video models should not be confused with Wan. The company runs more than one video generation effort out of different teams; coverage of an Alibaba video model is not automatically coverage of Wan. This trips up secondary sources regularly.
The hosted service's terms could not be retrieved. This is the most significant gap in this write-up and it is stated plainly rather than papered over. The footer links for Usage Policy, Terms of Service, Privacy Policy and Training Data Summary are all JavaScript handlers carrying no href. Requesting the corresponding paths directly returns only a 258-character single-page-application shell. Attempting to render create.wan.video/terms in a headless browser redirects to an error page. Consequently nothing in this document asserts anything about the hosted service's terms, privacy handling or training data disclosures. Anyone with a commercial dependency should obtain those documents from the company before committing.
Output rights are clear for the weights, unclear for the studio. The model card's disclaimer of rights over generated content is specific and verifiable. It applies to the open models. Whether create.wan.video applies the same position is not something this research could establish.
Flagship models require data-centre hardware. 80GB of VRAM puts A14B out of reach for individual users. The 5B model at 24GB is the realistic consumer ceiling, and the download distribution shows the community has adapted accordingly.
The free tier cannot produce publishable output. No watermark-free downloads, no 1080p, no longer durations, six image styles only.
The first billing period is non-refundable. Auto-renewal with cancellation at any time, but no refund on the first month or year after activation.
Purchases do not transfer between surfaces. A studio subscription does not cover the API or other products. Budget accordingly.
Mobile availability is regionally restricted. The company says so on its own site, and the absence of an official application in the United States App Store corroborates it.
Content restrictions apply and liability sits with you. The model card prohibits several categories of content and states users are fully accountable for compliance.
Site metadata lags the product. The current release is Wan3.0, while much of the descriptive material circulating — including the metadata this listing inherited — describes an earlier and narrower product. Check the site itself for current capability rather than relying on third-party descriptions.
Category placement needed correcting. This listing previously carried video editing as its primary category. The product's own framing is generation-first: three of four headline capabilities are generative, all five open model variants are generative or driving models, and the company's own title calls it a video generation model. Editing is real but secondary.
Beware conflation in secondary coverage. During this research, a search summary attributed a well-known video model ranking to Wan when the underlying article was about a different Alibaba model from a different team. Performance and ranking claims about "Alibaba's video model" should be checked against the article's actual subject.
Both yes and no, depending on which product you mean. The model weights are genuinely free under Apache 2.0 — download them, run them, use them commercially, subject to retaining the licence text and observing the stated content restrictions. Your only cost is hardware. The hosted studio at wan.video has a free tier, but a constrained one: one concurrent video and one concurrent image, six image styles, no watermark-free downloads, no 1080p, no longer videos. Paid tiers start at $5 per month billed annually.
Alibaba, through its Tongyi Lab, and yes. The footer reads "Wan © 2026 Powered by Alibaba Cloud" with wan_ai@service.alibaba.com as the contact. The stronger evidence is infrastructural: the site's stylesheet and application bundle are served from Alibaba's own CDN under the tongyi-lab-fe namespace, its analytics product identifier is tongyiwanxiang, and it loads an internal Alibaba npm package for tracking. A third party cannot deploy into another company's CDN namespace. The official GitHub organisation is Wan-Video and the official Hugging Face account is Wan-AI.
For the open models, the position is unusually clear. Apache 2.0 permits commercial use, and the model card states the team claims no rights over your generated content, granting freedom to use it provided your usage complies with the licence. For the hosted studio, this could not be determined — its terms of service were not retrievable during this research. If your commercial plan depends on studio output rather than self-hosted output, get the position in writing first.
It depends entirely on the variant. TI2V-5B needs at least 24GB of VRAM and the documentation names the RTX 4090 as a workable consumer card. The A14B models — T2V-A14B, I2V-A14B — need at least 80GB, which means data-centre hardware. This gap is the single most important practical fact about self-hosting Wan, and the download statistics show it: the 1.3B and 5B models are downloaded far more than the 14B ones.
Wan2.2 provides five. T2V-A14B generates video from text and uses a mixture-of-experts design with 27 billion total parameters and about 14 billion active per step. I2V-A14B generates video from an image on the same architecture. TI2V-5B is a smaller dense model doing both tasks in one checkpoint, and it is the one that runs on consumer hardware. S2V-14B generates video from speech. Animate-14B handles character animation and replacement. All generate at 480P and 720P at 24fps.
Wan3.0 is the current headline release on the hosted site — the primary button reads "Try Wan3.0 Now" and the feature section is titled "Introducing Wan3.0." It brings native 30-second generation, up to 20 reference assets with document and web page parsing, pixel-level consistency aimed at production delivery, and a set of image capabilities including twelve-language text rendering and nine-image fusion. The open weights most widely deployed remain the 2.1 and 2.2 generations, so the hosted studio and the downloadable models are not at the same version.
Credits are the metering unit on the hosted tiers. Pro provides 300 per month and Premium 1,200. The company gives conversion guidance rather than leaving you to guess: 300 credits translate to up to 1,200 images or 60 videos, and 1,200 credits to up to 4,800 images or 240 videos. Credits can also be bought separately from the membership tiers, and every tier can earn additional credits through daily check-in.
There are App Store and Google Play links on the site, but with a caveat the company states itself: the app is not available in all regions, and affected users are directed to a mobile web version. Checking the United States App Store across several search variations returns no Alibaba-published Wan app — only unrelated products and a couple of third-party apps borrowing the name, each with one or two ratings and far too small a sample to be meaningful.
Yes, with official support. ComfyUI's documentation carries a native workflow example for Wan2.2, describing it as the official usage guide for the model in ComfyUI, and confirms the Apache 2.0 licensing including commercial use. For most people who want to run these models without building infrastructure, this is the practical entry point.
No, and the pricing page is explicit about it. A subscription is for model creation on create.wan.video only; the API, DealDance and other products require separate purchases. If your plan involves programmatic generation, budget for that separately — and note that this research was not able to retrieve API pricing or call specifications.