Used this tool? Rate it
Used this tool? Rate it
Modal is a serverless cloud that runs Python code on GPUs and CPUs billed by the second, described by the company as cloud infrastructure for teams that develop, train, and serve AI applications at scale.
Modal puts your code in a container, executes it in the cloud and adds containers automatically when traffic grows, pooling capacity across major clouds to pick where each job runs. Container images and GPU requests are written as Python code, so a Modal app has no YAML deployment files.
Modal Labs operates modal.com. Independent tech press reporting says Modal was founded in 2021 by CEO Erik Bernhardsson and CTO Akshat Bubna; the company says it has raised over $466M from investors including General Catalyst and Redpoint Ventures. On September 28, 2026, a tech news report citing an unnamed source said Modal Labs, which it called an AI inference infrastructure provider, was nearing a $750 million round led by Accel at a $15.75 billion valuation; Modal Labs declined to comment.
Modal covers inference with what it calls sub-second cold starts, parallel batch jobs, training and fine-tuning, isolated Sandboxes for AI-generated code, and GPU-backed collaborative Notebooks.
The gpu argument accepts T4, L4, A10, L40S, A100 (generic, 40 GB or 80 GB), RTX PRO 6000, H100 (or H100!), H200, B200 (or B200+) and B300. B300, B200, H200, H100, A100, L4, T4 and L40S allow up to 8 GPUs per container (up to 2,304 GB of GPU RAM); A10 stops at 4 GPUs and 96 GB.
Two automatic upgrades matter for benchmarking: an H100 request may run on an H200 and an A100 request on an 80 GB A100, at the original price, while requesting H100! opts out of the H200 upgrade. A function can also list GPU types in order of preference, and Modal falls back down the list.
Modal Sandboxes are secure containers, defined at runtime, for executing untrusted user or agent code. Modal says millions can be created programmatically in under a minute, with snapshot and restore for fast startup.
Hosted Jupyter notebooks with serverless pricing, automatic idle shutdown, Modal GPUs and real-time collaborative editing.
Modal announced on October 1, 2026 that Modal Clusters are generally available through the @modal.clustered decorator: multi-node GPU groups with RDMA, gang scheduled from a shared capacity pool and billed by the second.
pip install modal, then modal setup to authenticate (or python -m modal setup if the first form fails).@app.function(...), naming the container image and GPU, and run the file with modal run path/to/file.py.Modal says containers boot in about one second, but a container is only warm once imports and other global-scope setup have run, so readiness can take from seconds to minutes. Three settings trade latency against idle spend:
scaledown_window: keeps idle containers alive for two seconds to twenty minutes; longer windows mean fewer cold starts, but idle resources are billed.min_containers and buffer_containers: the first keeps a floor of warm containers so the function never scales to zero, the second adds spare containers while it is active.us offers a larger resource pool than a narrow one, improving cold-start time and availability.Modal can deploy any containerizable Python application and targets compute-heavy work that scales horizontally: ML inference at scale, model training and fine-tuning, data pipelines, scientific computing and job queues.
The three customer statements are homepage testimonials, not independent case studies.
Independent review signals found this round are limited: on a launch-community review platform Modal showed 5.0 based on 61 reviews as captured on October 5, 2026, and that platform's AI summary says reviewers are mostly founders and the reviews contain almost no critical feedback.
routing_region also accepts us-west (Oregon), eu-west (Dublin) and ap-south (Mumbai).Modal combines a monthly plan with per-second GPU pricing and metered CPU and memory; list prices below were captured on October 5, 2026.
| Plan | Platform fee | Included compute | Seats | Containers / GPU concurrency | Log retention | Included egress |
|---|---|---|---|---|---|---|
| Starter | $0 + compute / month | $30 / month | 3 | 100 / 10 | 1 day | 1 TiB / month |
| Team | $250 + compute / month | $100 / month | Unlimited | 5000 / 50 | 30 days | 10 TiB / month |
| Enterprise | Custom | Custom | Unlimited | Higher GPU concurrency | Custom | 100 TiB / month |
Enterprise adds volume-based discounts, embedded ML engineering services, private Slack support, audit logs, Okta SSO and HIPAA. Starter allows 200 deployed apps and 5 deployed cron jobs and excludes custom domains, deployment rollbacks, environment-level budgets, the static IP proxy and RBAC, and Team keeps 3 rollback versions.
| Resource | Price per second |
|---|---|
| Nvidia B300 | $0.001972 |
| Nvidia B200 | $0.001736 |
| Nvidia H200 SXM | $0.001261 |
| Nvidia H100 SXM5 | $0.001097 |
| Nvidia RTX PRO 6000 | $0.000842 |
| Nvidia A100, 80 GB | $0.000694 |
| Nvidia A100, 40 GB | $0.000583 |
| Nvidia L40S | $0.000542 |
| Nvidia A10 | $0.000306 |
| Nvidia L4 | $0.000222 |
| Nvidia T4 | $0.000164 |
| CPU, per physical core (2 vCPU) | $0.0000131 |
| Memory, per GiB | $0.00000222 |
CPU has a minimum of 0.125 cores per container. Sandboxes and Notebooks use higher CPU and memory rates, $0.00003942 per core and $0.00000667 per GiB per second, with GPUs at standard prices. Volumes cost $0.09 per GiB per month after 1 TiB free, and egress beyond the plan allowance costs $0.04 per GiB. Modal's worked example puts Stable Diffusion on an A10G at about $0.491 per 1,000 images.
nonpreemptible=True triples CPU and memory list prices and is not available for GPU functions.Each serverless GPU platform below differs from Modal mainly in how idle time is billed and how wide the GPU menu is.
| Option | GPU choice | Scale to zero | Idle and billing rule |
|---|---|---|---|
| Google Cloud Run (GPU) | RTX PRO 6000 Blackwell (96 GB) or L4 (24 GB) | Yes | Instance-based billing; minimum instances charged at full rate even when idle |
| RunPod Serverless | Not covered in this comparison | Flex workers yes; active workers run 24/7 | Billed per second from worker start until it fully stops, rounded up |
| Replicate | Not covered in this comparison | Public models yes | Most private models run on dedicated hardware and are billed for setup, idle and active time |
RunPod's flex-versus-active split mirrors Modal's min_containers choice. Replicate bills most public models by run time and some by input and output. Modal's differentiators are Python-defined infrastructure and Sandboxes in the same account, plus a longer GPU list than Cloud Run.
routing_region can only be set at a function's first deployment, and inputs and outputs larger than 2 MiB still go to object storage in us-east.Modal's agreement prohibits cryptocurrency mining or related blockchain activities, denial-of-service attacks, peer-to-peer file sharing, and general file-hosting or media-serving platforms.
Starter has no platform fee and includes $30 of compute a month, but the billing guide requires a payment method on file and Shared Endpoint tokens are billed from the first request.
Not after it scales to zero. The default 60-second keep-alive and idle time inside a longer scaledown_window are billed, and min_containers keeps the function from ever reaching zero.
Functions and Sandboxes can be pinned to EU regions at a 1.15x or 1.75x multiplier, with routing through eu-west, but application logs stay in the United States, large inputs and outputs pass through us-east, and Shared Endpoints cannot be pinned.