Toolso.AI
Toolso.AI
All ToolsCategoriesTrendingLatest ToolsPricingBlog
Toolso.AI
Toolso.AI
Toolso.AI
Toolso.AI

Discover the best AI tools to boost your productivity

GitHubGitHubTwitterX (Twitter)YouTubeYouTubeTikTokEmail

Popular Categories

  • AI Writing
  • AI Image
  • AI Video
  • AI Coding
  • More Categories

Explore

  • Latest Tools
  • Popular Tools
  • More Tools
  • Submit Tool
  • Pricing

About

  • About Us
  • Contact
  • Blog
  • Changelog

Legal

  • Cookie Policy
  • Privacy Policy
  • Terms of Service
  • Refund Policy
© 2026 Toolso.AI All Rights Reserved
Limited timeLimited-time offerFeatured Listing24h priority review · No backlink · 30 days featured$29.90then $59.90Price rises to $59.90 after Oct 31Ends in--:--:--Submit now
  1. Home
  2. All Tools
  3. Developer Tools
  4. Modal
Modal interface preview
Modal logo

Modal

Modal runs Python functions, inference endpoints, isolated sandboxes and notebooks on GPUs and CPUs billed by the second, from a $0 Starter plan with $30 of monthly compute to custom Enterprise contracts.

Developer ToolsAI DevelopmentAI Training Platform#Machine Learning#Batch Processing#OpenAI Compatible API
Try for Free
Saves
Visits
Views
Pricing
Free
Published
Oct 5, 2026
Domain
modal.com
Community rating

Used this tool? Rate it

Rate this tool

Modal Product Information

Try for Free
Tool Information
Saves
Visits
Views
Pricing
Free
Published
Oct 5, 2026
Domain
modal.com
Community rating

Used this tool? Rate it

Rate this tool

Featured Tools

Related Tools

Try for Free

What is Modal?

Modal is a serverless cloud that runs Python code on GPUs and CPUs billed by the second, described by the company as cloud infrastructure for teams that develop, train, and serve AI applications at scale.

Modal puts your code in a container, executes it in the cloud and adds containers automatically when traffic grows, pooling capacity across major clouds to pick where each job runs. Container images and GPU requests are written as Python code, so a Modal app has no YAML deployment files.

Modal Labs operates modal.com. Independent tech press reporting says Modal was founded in 2021 by CEO Erik Bernhardsson and CTO Akshat Bubna; the company says it has raised over $466M from investors including General Catalyst and Redpoint Ventures. On September 28, 2026, a tech news report citing an unnamed source said Modal Labs, which it called an AI inference infrastructure provider, was nearing a $750 million round led by Accel at a $15.75 billion valuation; Modal Labs declined to comment.

Core features

Modal covers inference with what it calls sub-second cold starts, parallel batch jobs, training and fine-tuning, isolated Sandboxes for AI-generated code, and GPU-backed collaborative Notebooks.

Functions, endpoints and batch jobs

  • HTTPS endpoints: any function can become an HTTPS endpoint that scales with traffic and idles at zero.
  • Fan-out and background work: a job can fan out over thousands of GPUs without an orchestration layer, or spawn asynchronous batch work.
  • Storage: Volumes act as a distributed filesystem for images, model weights and datasets.
  • Observability: per-container logs and metrics, with export to OpenTelemetry providers such as Datadog.

GPU selection

The gpu argument accepts T4, L4, A10, L40S, A100 (generic, 40 GB or 80 GB), RTX PRO 6000, H100 (or H100!), H200, B200 (or B200+) and B300. B300, B200, H200, H100, A100, L4, T4 and L40S allow up to 8 GPUs per container (up to 2,304 GB of GPU RAM); A10 stops at 4 GPUs and 96 GB.

Two automatic upgrades matter for benchmarking: an H100 request may run on an H200 and an A100 request on an 80 GB A100, at the original price, while requesting H100! opts out of the H200 upgrade. A function can also list GPU types in order of preference, and Modal falls back down the list.

Inference endpoints

  • Shared Endpoints: a subset of Modal Library models served from Modal-managed pools and billed per token; each can also run on a Dedicated Endpoint for isolated or configurable capacity.
  • OpenAI-compatible serving: Modal markets OpenAI-compatible endpoints for open models.

Modal Sandboxes

Modal Sandboxes are secure containers, defined at runtime, for executing untrusted user or agent code. Modal says millions can be created programmatically in under a minute, with snapshot and restore for fast startup.

Notebooks

Hosted Jupyter notebooks with serverless pricing, automatic idle shutdown, Modal GPUs and real-time collaborative editing.

Training and multi-node clusters

Modal announced on October 1, 2026 that Modal Clusters are generally available through the @modal.clustered decorator: multi-node GPU groups with RDMA, gang scheduled from a shared capacity pool and billed by the second.

Guide

Getting started

  1. Create an account at modal.com.
  2. Run pip install modal, then modal setup to authenticate (or python -m modal setup if the first form fails).
  3. Write a Python function decorated with @app.function(...), naming the container image and GPU, and run the file with modal run path/to/file.py.

Tuning cold starts and idle cost

Modal says containers boot in about one second, but a container is only warm once imports and other global-scope setup have run, so readiness can take from seconds to minutes. Three settings trade latency against idle spend:

  • scaledown_window: keeps idle containers alive for two seconds to twenty minutes; longer windows mean fewer cold starts, but idle resources are billed.
  • min_containers and buffer_containers: the first keeps a floor of warm containers so the function never scales to zero, the second adds spare containers while it is active.
  • Broad regions: a broad region such as us offers a larger resource pool than a narrow one, improving cold-start time and availability.

Use cases and examples

Modal can deploy any containerizable Python application and targets compute-heavy work that scales horizontally: ML inference at scale, model training and fine-tuning, data pipelines, scientific computing and job queues.

  • Company-wide AI platforms: DoorDash co-founder Andy Fang says its AI platform uses Modal for inference, sandboxes for agents and compute for custom code.
  • Reinforcement learning plus serving: Cognition CEO Scott Wu says Modal runs both its RL infrastructure and its production inference.
  • Agent evaluation: Ramp's head of applied AI says its Inspect tool would have been hard to build without fast Modal sandboxes.
  • Reference projects: examples cover real-time transcription with Kyutai STT, Flux fine-tuning, a coding agent on Sandboxes and LangGraph, training a small language model, and parallel Parquet processing on S3.

The three customer statements are homepage testimonials, not independent case studies.

Who is it for

  • Python developers and ML engineers: functions are defined in Python; JavaScript/TypeScript and Go can call functions, run Sandboxes and manage resources.
  • Teams with bursty demand: Modal argues that its per-hour rates may look higher while total cost is lower for spiky or unpredictable request volumes.
  • Small teams to large organizations: Starter targets small teams and independent developers, Team startups and larger organizations, Enterprise organizations prioritizing security and support.
  • Startups and researchers: early-stage startups can apply for free compute credits, and academics can get up to $10k in credits.

Independent review signals found this round are limited: on a launch-community review platform Modal showed 5.0 based on 61 reviews as captured on October 5, 2026, and that platform's AI summary says reviewers are mostly founders and the reviews contain almost no critical feedback.

Platforms

  • Python SDK and CLI: most interactions go through the open-source modal command-line tool and Python client library.
  • JavaScript/TypeScript and Go SDKs: in Beta; defining functions will likely remain Python-only.
  • Browser: Notebooks open at modal.com/notebooks, and .ipynb files can be uploaded.
  • Capacity regions: the US, EU, Asia Pacific, Japan, Australia, the United Kingdom, Canada, the Middle East, South America, Africa and Mexico, across more than 20 clouds.
  • Request routing: function inputs route through us-east (Virginia) by default; routing_region also accepts us-west (Oregon), eu-west (Dublin) and ap-south (Mumbai).
  • Cloud marketplaces: Enterprise customers can buy through the AWS and GCP marketplaces to use committed spend.

Pricing

Modal combines a monthly plan with per-second GPU pricing and metered CPU and memory; list prices below were captured on October 5, 2026.

Plans

PlanPlatform feeIncluded computeSeatsContainers / GPU concurrencyLog retentionIncluded egress
Starter$0 + compute / month$30 / month3100 / 101 day1 TiB / month
Team$250 + compute / month$100 / monthUnlimited5000 / 5030 days10 TiB / month
EnterpriseCustomCustomUnlimitedHigher GPU concurrencyCustom100 TiB / month

Enterprise adds volume-based discounts, embedded ML engineering services, private Slack support, audit logs, Okta SSO and HIPAA. Starter allows 200 deployed apps and 5 deployed cron jobs and excludes custom domains, deployment rollbacks, environment-level budgets, the static IP proxy and RBAC, and Team keeps 3 rollback versions.

Per-second GPU pricing

ResourcePrice per second
Nvidia B300$0.001972
Nvidia B200$0.001736
Nvidia H200 SXM$0.001261
Nvidia H100 SXM5$0.001097
Nvidia RTX PRO 6000$0.000842
Nvidia A100, 80 GB$0.000694
Nvidia A100, 40 GB$0.000583
Nvidia L40S$0.000542
Nvidia A10$0.000306
Nvidia L4$0.000222
Nvidia T4$0.000164
CPU, per physical core (2 vCPU)$0.0000131
Memory, per GiB$0.00000222

CPU has a minimum of 0.125 cores per container. Sandboxes and Notebooks use higher CPU and memory rates, $0.00003942 per core and $0.00000667 per GiB per second, with GPUs at standard prices. Volumes cost $0.09 per GiB per month after 1 TiB free, and egress beyond the plan allowance costs $0.04 per GiB. Modal's worked example puts Stable Diffusion on an A10G at about $0.491 per 1,000 images.

What counts toward the bill

  • Load and idle time: Modal bills load and processing time plus a default, configurable 60-second keep-alive after the last input; an app scaled to zero is not charged.
  • Request or usage: CPU and memory are charged on whichever is higher, requested or used, with a minimum of 128 MiB and 0.125 cores per container.
  • Region pinning: a pinned container region costs 1.15x base prices for a broad region and 1.75x for a narrow one; mixing both applies the smaller multiplier.
  • Non-preemptible execution: nonpreemptible=True triples CPU and memory list prices and is not available for GPU functions.
  • Shared Endpoint tokens: included compute does not cover them; since September 1, 2026, tokens are billed from the first request, even on Starter.
  • Payment method: the billing guide requires one on file to use Modal at all, while the GPU guide ties it to GPU use; the pages differ in scope.
  • Billing cycle and caps: billing is monthly, plus auto-charges when usage first crosses certain thresholds; a reached spend limit stops workloads that would add out-of-pocket charges.
  • Cloud credits and refunds: AWS, GCP or Azure credits cannot be applied, and Modal's Software as a Service Agreement, effective May 2026, says fees are not refundable except as described in its section 3.2.

Modal alternatives

Each serverless GPU platform below differs from Modal mainly in how idle time is billed and how wide the GPU menu is.

OptionGPU choiceScale to zeroIdle and billing rule
Google Cloud Run (GPU)RTX PRO 6000 Blackwell (96 GB) or L4 (24 GB)YesInstance-based billing; minimum instances charged at full rate even when idle
RunPod ServerlessNot covered in this comparisonFlex workers yes; active workers run 24/7Billed per second from worker start until it fully stops, rounded up
ReplicateNot covered in this comparisonPublic models yesMost private models run on dedicated hardware and are billed for setup, idle and active time

RunPod's flex-versus-active split mirrors Modal's min_containers choice. Replicate bills most public models by run time and some by input and output. Modal's differentiators are Python-defined infrastructure and Sandboxes in the same account, plus a longer GPU list than Cloud Run.

Limitations

Runtime limits

  • Preemption by default: any function can be preempted and is restarted on the same input; interruption becomes more likely as run duration grows.
  • Timeouts: function executions default to 300 seconds, configurable from 1 second to 24 hours; Sandboxes default to a 5-minute lifetime, extendable to 24 hours.
  • Concurrency: workspace limits on concurrent containers and GPUs follow the plan, and a single function has a hard limit of 4,000 concurrent containers.
  • Large GPU requests: more than 2 GPUs per container usually means longer waits.
  • CUDA floor: B300 requires CUDA 13.1 or later.
  • Shared Endpoint throttling: model-specific concurrency limits apply across the workspace, and requests above the limit receive HTTP 429.
  • Inconsistent multi-node status: the GPU guide still says multi-node training is in private Beta, while the October 1, 2026 announcement calls Modal Clusters generally available.

Regions and data residency

  • Strict pinning: region-pinned workloads are never moved elsewhere, even if that region runs out of capacity.
  • Routing restrictions: routing_region can only be set at a function's first deployment, and inputs and outputs larger than 2 MiB still go to object storage in us-east.
  • Shared Endpoints: no selectable container or routing region; traffic routes through us-west across the global compute pool.
  • Logs: application logs from Functions, Sandboxes and Servers are stored in the United States.

Acceptable use

Modal's agreement prohibits cryptocurrency mining or related blockchain activities, denial-of-service attacks, peer-to-peer file sharing, and general file-hosting or media-serving platforms.

Privacy, security and data rights

  • Data access: Modal says it will never access or use your source code, function inputs or outputs, or data stored in Images or Volumes; app logs and metadata are accessed only with your permission for troubleshooting.
  • Retention: function inputs and outputs are encrypted at rest and deleted within a maximum TTL of 7 days; Modal Inference endpoints are zero data retention, with payloads never written to disk.
  • Ownership: customers keep all rights in Customer Data and grant Modal a limited license to provide the service; Modal may use aggregated or permanently anonymized Service Metrics for its own business purposes during and after the agreement.
  • Processing locations: customers can restrict processing to specified regions, and Modal will not process Customer Data elsewhere without prior written consent.
  • Certifications: Modal has completed a SOC 2 Type 2 audit; the report is available on request through its Security Portal.
  • HIPAA: a Business Associate Agreement, available on Enterprise, must be in place before any PHI is submitted; Volumes v1, Images (except Filesystem and Directory Snapshots), Memory Snapshots and user code fall outside it, while Volumes v2 are covered.
  • Isolation and encryption: compute jobs are virtualized with gVisor, Sandboxes may also run on a secure VM runtime, user data is encrypted in transit and at rest, and public APIs use TLS 1.3.
  • Shared responsibility: customers are responsible for backups of their own data.
  • Privacy policy: last updated May 17th, 2023; it says the services do not address anyone under 13.

FAQ

Q1. Does Modal have a free tier?

Starter has no platform fee and includes $30 of compute a month, but the billing guide requires a payment method on file and Shared Endpoint tokens are billed from the first request.

Q2. Am I charged when my app is idle?

Not after it scales to zero. The default 60-second keep-alive and idle time inside a longer scaledown_window are billed, and min_containers keeps the function from ever reaching zero.

Q3. Can I keep workloads inside the EU?

Functions and Sandboxes can be pinned to EU regions at a 1.15x or 1.75x multiplier, with routing through eu-west, but application logs stay in the United States, large inputs and outputs pass through us-east, and Shared Endpoints cannot be pinned.

Know a Similar Tool?
If you know other great AI tools, feel free to submit them to us