WaveSpeedAI is a hosted AI inference API and model hub that sells per-request access to generative models. The company describes the platform as offering 1,000+ image, video, audio and 3D generation models alongside 290+ LLMs, split between a media inference API and a separate OpenAI-compatible LLM API.
WaveSpeedAI was founded in 2025 by CEO Zeyi Cheng and is headquartered in Singapore, according to the company. The Terms of Service name two contracting entities: WaveSpeedAI PTE. LTD., incorporated in Singapore, and WaveSpeedAI LIMITED, incorporated in Hong Kong.
The pitch is consolidation: one endpoint format, one key and one prepaid balance instead of separate accounts with each model owner. Because WaveSpeedAI is a serving layer, output quality, content rules and commercial-use terms still depend on the specific model you call, and throughput depends on your account level rather than a monthly plan.
The media API is task-based. By default, WaveSpeedAI uses asynchronous mode: a submission returns a task ID, and your application later polls the result endpoint or receives a webhook. That design suits an AI video generation API, where long video jobs can outlast an HTTP request.
enable_sync_mode flag waits for the result, but the API can return a timeout body while the task keeps running; the option is API-only and supported only by some models.The LLM side runs on a separate base URL. The service exposes OpenAI-compatible Chat Completions and Responses endpoints plus an Anthropic-compatible Messages endpoint. One WaveSpeedAI API key covers every listed LLM provider, and switching models means changing the model value.
LoRA training lets you fine-tune supported models on your own images to get personalized styles or consistent characters. The guide recommends a dataset of 10-20 diverse images. Typical estimates are about 8 minutes for 1000 steps and about 25 minutes for 3000 steps, and a timeout or system error during training is refunded automatically.
Async mode is recommended for production integrations, long-running tasks, video generation, unstable networks, and workflows that need reliable result recovery. Sync mode cannot be combined with a webhook on the same request.
The CLI signs in through the browser or with an API key on headless machines. Its documented order of work is models, schema, price, run, then download or history, which puts a price quote before every paid run.
These scenarios come from WaveSpeedAI's own documentation and show intended use, not verified outcomes.
It is a weaker fit for individuals who want monthly invoicing: individual users without a product or project cannot get monthly credit lines and must prepay.
WaveSpeedAI uses pay-per-use pricing with no monthly fees or commitments: you top up credits and each run is charged at that model's rate. Eligible new accounts receive $1 in free credits without a credit card, although some premium models may not be available with trial credits.
Starting prices and units below are from the pricing page as captured on September 28, 2026; "Output per $1" is the page's own figure.
| Image model | Starting price | Output per $1 |
|---|---|---|
| Seedream 5.0 Pro | $0.045/image | 22 images |
| Nano Banana Pro | $0.14/image | 7 images |
| GPT Image 2.5 | $0.01/image | 100 images |
| Z-Image Turbo | $0.005/image | 200 images |
| Video model | Starting price | Output per $1 |
|---|---|---|
| Seedance 2.5 | $0.18/second | 5.5 seconds |
| Wan 3.0 | $0.05/second | 20 seconds |
| Veo 3.1 Fast | $0.10/second | 10 seconds |
| InfiniteTalk | $0.03/second | 33.3 seconds |
Starting prices are list prices at the lowest resolution or quality. Actual cost depends on resolution, duration and other parameters, and discounts may apply; the exact price for a given job is shown on the model page or returned by the pricing API.
| LLM | Context window | Input per 1M tokens | Output per 1M tokens |
|---|---|---|---|
| Claude Opus 5.5 | 1M | $4 | $20 |
| Gemini 3.8 Flash | 1M | $1.5 | $7.5 |
| GPT-6 Luna | 1.05M | $0.1 | $0.5 |
| DeepSeek V4.1 Flash | 1M | $0.15 | $0.6 |
Some LLMs include tiered pricing, cache pricing or other billing rules beyond these base rates. For open-source models, WaveSpeedAI says its pricing matches the original providers.
| Level | Predictions/min | Max concurrency | How it is reached |
|---|---|---|---|
| Bronze | 5 | 2 | Default for new users |
| Silver | 500 | 300 | Any successful single top-up below $1,000 |
| Gold | 3,000 | 3,000 | A single top-up of $1,000–$4,999 |
| Ultra | 5,000 | 10,000 | A single top-up of $5,000 or more |
Upgrades are based on a single top-up amount and are upward-only.
A software-alternatives directory lists fal, Replicate.com and KLING AI as the top competitors of WaveSpeedAI. KLING AI's video models are themselves hosted on WaveSpeedAI, so the closer comparisons are fal and Replicate.
Yes, with limits. Eligible new accounts get $1 in trial credit without a credit card, but some premium models are excluded and registering a new email does not guarantee eligibility.
No. The Acceptable Use Policy bans pornographic and sexually explicit content, the web Safety Checker is on by default, and the same rules apply to API use.
New accounts start at Bronze, which allows 5 predictions per minute and 2 concurrent tasks; any successful single top-up moves an unrestricted account to at least Silver.