fal is inference infrastructure for generative media. The distinction matters: fal does not train or own the models it serves, and it is not an application you open to make an image. It is the layer developers build against when their product needs to run image, video, audio or 3D generation in production, and they would rather not operate a GPU fleet to do it. The official positioning is unambiguous — a generative media platform for developers, offering the world's best generative image, video, and audio models, all in one place.
The company behind it has grown unusually fast. TechCrunch reported in December 2025 that fal raised a $140 million Series D led by Sequoia, with participation from Kleiner Perkins and Nvidia, at a $4.5 billion valuation, and put revenue at more than $200 million as of October. The founders are Burkay Gur, a former Coinbase machine learning leader, and Gorkem Yurtseven, previously an engineer at Amazon. Named customers include Adobe, Shopify, Canva and Quora. Two caveats belong with those numbers: this was the company's third fundraise of 2025, with the valuation tripling from roughly $1.5 billion in July, and the $140 million figure combines new capital with a secondary sale in which existing investors sold shares — it is not all fresh money into the business.
What you actually get is three product lines rather than one API. Model APIs let you call models that already exist on the platform. Serverless lets you deploy your own models onto the same engine, billed per second of execution with automatic scaling. Compute gives you dedicated GPU instances with full SSH access, billed at a fixed hourly rate, for training and fine-tuning. Choosing correctly between them is most of the work of adopting fal, and the sections below lay out where each one fits.
fal run boots the app on a temporary cloud GPU for testing on the same hardware you will use in production; fal deploy promotes it to a persistent authenticated endpoint with autoscaling and built-in retries, and every deploy creates a new revision for instant rollbacks.fal run before deploying. It boots your app on a temporary worker running setup() and your endpoints exactly as production would, so errors surface there instead of as a production crashloop.fal deploy, then tune min_concurrency, max_concurrency and concurrency_buffer against observed traffic. Watch the dashboard's request-level analytics, and export to Prometheus or an HTTPS log drain if you have an existing observability stack.def run(self, prompt: str) is interpreted as a query parameter, so callers sending a JSON body — which is what the clients and every example do — get an HTTP 422 response.fal run and fal deploy drive the development and deployment lifecycle from the terminal.Pricing follows the product split. Model APIs bill by output unit rather than by GPU time, which is the platform's main pricing distinction: video models are billed by output unit — per second or per video — depending on the model, with published examples including Wan 2.5 at $0.05 per second, Kling 2.5 Turbo Pro at $0.07 per second, Veo 3 at $0.4 per second and Ovi at $0.2 per video. Image models bill by image count or by megapixel, with Seedream V4 at $0.03 per image, Flux Kontext Pro at $0.04, Nanobanana at $0.039 and Qwen at $0.02 per megapixel. A third-party comparison notes that this is more predictable than per-GPU-second billing, where cost varies with how long processing takes.
Compute bills GPU instances by the hour, with list prices of $8.50 for a B300 (288GB), $6.25 for a B200 (180GB), $4.50 for an H200 (141GB), $4.50 for an H100 (80GB) and $2.99 for an RTX PRO 6000 (96GB), each with a lower "as low as" rate available through sales — down to $1.89 per hour for H100. Serverless bills per second of execution. Read the official caveats alongside the headline numbers: the per-dollar output comparisons assume an estimated average video of five seconds at 720p and vary with model, resolution and prompt complexity; image prices are normalized to 1MP with higher resolutions priced proportionally; and some models use GPU-based pricing rather than output-based pricing depending on architecture. Enterprise terms are quoted directly.
No. fal is inference infrastructure that developers call from their own applications. It runs generative image, video, audio and 3D models built by others, through an API. If you want to create an image directly without writing code, fal is the layer underneath such tools rather than the tool itself.
It depends on the product line. Model APIs bill by output unit — per second or per video for video models, per image or per megapixel for image models. Serverless bills per second of execution. Compute bills dedicated GPU instances at a fixed hourly rate.
A third-party comparison reports that new Model API accounts begin with 2 concurrent requests, rising based on paid invoices over the previous four weeks and self-service up to 40, with excess requests queued. Plan for this before a launch rather than during one.
Yes, through Serverless. You write a fal.App Python class where setup() loads weights and @fal.endpoint methods serve requests, declare hardware alongside the code, validate with fal run, then fal deploy to a persistent endpoint with autoscaling, retries and revision-based rollbacks.
Python and JavaScript SDKs, plus a REST API. Every model supports synchronous and asynchronous queue calls, and many also support streaming and real-time WebSocket connections. Note that timeout parameters differ between SDKs: Python uses client_timeout in seconds, JavaScript uses timeout in milliseconds.
For enterprise customers the site states plainly that your data stays yours and that fal never trains its models on enterprise customers' data. The enterprise offering also advertises SOC2 certification, SSO and private model hosting. Confirm the terms applicable to your specific plan.
There are three, with different owners. start_timeout is enforced server-side before processing begins and returns 504 while stopping retries. client_timeout or timeout is client-side only and does not stop server-side execution. request_timeout is set by the app developer as a per-attempt cap that kills the runner and triggers a retry.
Yes. fal retries queue requests that fail due to server errors, timeouts or rate limits by default. Send the X-Fal-No-Retry header on submission to disable this for a specific request, which matters for expensive or non-idempotent generations.
An independent comparison summarizes it as fal winning on speed and FLUX-family economics, and Replicate winning on model variety outside image and video and on community-contributed custom models. Output-based pricing also makes fal's per-call cost knowable in advance, whereas per-GPU-second billing varies with processing time.
It publishes a public status page, which at the time of writing showed all systems operational with 100% uptime across the 90-day window and no notices in the prior seven days, and it reports enterprise customers including Adobe, Shopify and Canva. That is a single-point snapshot rather than a long-run guarantee; evaluate against your own availability requirements and check the status history yourself.