Nebius is an AI cloud provider that rents NVIDIA GPU virtual machines and multi-node clusters, together with managed Kubernetes, managed Slurm, storage and a hosted inference service for open models. Nebius describes itself as an AI cloud company whose platform spans data, model training and tuning, production runtime and deployment.
Nebius is listed on Nasdaq under the ticker NBIS and is headquartered in Amsterdam. The offer splits into two products that are billed differently. Nebius AI Cloud is infrastructure: you create VMs, clusters, disks and buckets, and pay for the time and capacity they use. Nebius Token Factory is a model API: you send prompts to hosted open models and pay for tokens instead of servers.
Three points matter before choosing: GPU list prices changed on October 1, 2026, on-demand capacity is not contractually guaranteed, and Token Factory keeps prompts for speculative decoding unless Zero Data Retention is on.
Serverless AI has three services. Devlabs gives interactive environments with Jupyter and VS Code, Jobs run containerized workloads that start, run and finish, and Endpoints serve custom models over HTTP. Serverless AI is billed only while workloads are running, with no idle GPU costs.
Nebius storage includes a shared filesystem, a WEKA filesystem, block volumes, local SSD disks and object storage with Standard, Intelligent and Enhanced tiers. Object storage speaks the Amazon S3 API, so existing S3 tools can be reused.
The examples below come from customer stories that Nebius publishes on its own site; they are vendor-selected results, not independent audits.
Nebius fits less well for buyers who need guaranteed on-demand capacity without signing a commitment, or for users who do not want to track and delete storage after each experiment.
Nebius GPU cloud pricing is pay-as-you-go by default. GPUs on running VMs are billed per second and priced per hour, so 30 minutes costs half the hourly rate. Prices are in USD for all customers except companies from Israel, which are billed in ILS. Listed prices exclude applicable taxes, including VAT.
| Instance | vCPUs | RAM, GB | On-demand, GPU-hour | GPU-hour (Effective October 1, 2026) |
|---|---|---|---|---|
| NVIDIA GB300 NVL72 | 112 | 800 | Contact us | Contact us |
| NVIDIA HGX B300 | 24 | 346 | $7.85 | $9.50 |
| NVIDIA GB200 NVL72 | 112 | 800 | Contact us | Contact us |
| NVIDIA HGX B200 | 20 | 224 | $7.15 | $8.50 |
| NVIDIA HGX H200 | 16 | 200 | $4.50 | $5.40 |
| NVIDIA HGX H100 | 16 | 200 | $3.85 | $4.50 |
| NVIDIA RTX PRO 6000 | 24 | 218 | $1.80 | $1.80 |
| NVIDIA L40S with Intel CPU | 8-40 | 32-160 | from $1.55 | from $1.55 |
| NVIDIA L40S with AMD CPU | 16-192 | 96-1152 | from $1.82 | from $1.82 |
As captured on October 5, 2026, the AI Cloud price page shows both GPU-hour columns side by side, labeled "On-demand" and "Effective October 1, 2026". The Compute pricing documentation lists the H100 NVLink at $4.50 per GPU hour from October 1, 2026 and $3.85 before that date. The same documentation says the update covers VMs with B300, B200, H200 and H100 GPUs.
Token Factory gives $1 in trial credit on first sign-up, valid for 30 days. The card is charged automatically at the start of a month with a negative balance or when the billing threshold is reached, and a failed card charge suspends the account.
An independent GPU-cloud rating report from September 2026 said Nebius still trails CoreWeave on many technical aspects and relationships with NVIDIA and frontier labs.
An independent rating report published in September 2026 called Nebius an industry leader with strong offerings in every category and the ability to command a significant price premium. It also described, as anecdote, aggressive negotiations with high prices and prepayments reaching 100% on one-year commitments. In the report's own tests, automatic return of a failed GB300 node to service worked but took 8h40m. The testers also first saw slow 4 KiB writes and EEXIST errors on a cluster data filesystem until Nebius retuned it after their feedback.
Yes. For GPUs on running VMs the billing unit is one second and the pricing unit is one hour.
No. Commitment discounts require an account registered as a company; individual accounts stay on pay-as-you-go billing.
Nebius states that Token Factory customer content is not used to train any models. Without Zero Data Retention, prompts and outputs are still stored and used for speculative decoding.
Token Factory adds $1 of trial credit valid for 30 days. On AI Cloud, free credits come through promo codes or, for startups, through partner venture funds.