Toolso.AI
Toolso.AI
All ToolsCategoriesTrendingLatest ToolsPricingBlog
Toolso.AI
Toolso.AI
Toolso.AI
Toolso.AI

Discover the best AI tools to boost your productivity

GitHubGitHubTwitterX (Twitter)YouTubeYouTubeTikTokEmail

Popular Categories

  • AI Writing
  • AI Image
  • AI Video
  • AI Coding
  • More Categories

Explore

  • Latest Tools
  • Popular Tools
  • More Tools
  • Submit Tool
  • Pricing

About

  • About Us
  • Contact
  • Blog
  • Changelog

Legal

  • Cookie Policy
  • Privacy Policy
  • Terms of Service
  • Refund Policy
© 2026 Toolso.AI All Rights Reserved
Limited timeLimited-time offerFeatured Listing24h priority review · No backlink · 30 days featured$29.90then $59.90Price rises to $59.90 after Oct 31Ends in--:--:--Submit now
  1. Home
  2. All Tools
  3. Developer Tools
  4. Nebius
Nebius interface preview
Nebius logo

Nebius

Nebius rents NVIDIA GPU virtual machines and clusters with managed Kubernetes and Slurm, and serves open models through Token Factory's OpenAI-compatible API.

Developer ToolsAI DevelopmentAI Training Platform#Machine Learning#LLM#OpenAI Compatible API
View Pricing
Saves
Visits
Views
Pricing
Paid
Published
Oct 6, 2026
Domain
nebius.com
Community rating

Used this tool? Rate it

Rate this tool

Nebius Product Information

View Pricing
Tool Information
Saves
Visits
Views
Pricing
Paid
Published
Oct 6, 2026
Domain
nebius.com
Community rating

Used this tool? Rate it

Rate this tool

Featured Tools

Related Tools

View Pricing

What is Nebius?

Nebius is an AI cloud provider that rents NVIDIA GPU virtual machines and multi-node clusters, together with managed Kubernetes, managed Slurm, storage and a hosted inference service for open models. Nebius describes itself as an AI cloud company whose platform spans data, model training and tuning, production runtime and deployment.

Nebius is listed on Nasdaq under the ticker NBIS and is headquartered in Amsterdam. The offer splits into two products that are billed differently. Nebius AI Cloud is infrastructure: you create VMs, clusters, disks and buckets, and pay for the time and capacity they use. Nebius Token Factory is a model API: you send prompts to hosted open models and pay for tokens instead of servers.

Three points matter before choosing: GPU list prices changed on October 1, 2026, on-demand capacity is not contractually guaranteed, and Token Factory keeps prompts for speculative decoding unless Zero Data Retention is on.

Core features

GPU compute

  • Non-virtualized GPUs: Nebius virtual instances do not virtualize GPUs or network interfaces, which the company says delivers performance on par with the best industry benchmarks.
  • Cluster scale: Nebius compute scales from single-node instances to thousand-GPU clusters on a non-blocking NVIDIA Quantum-2 InfiniBand fabric.
  • GPU line-up: rack-scale GB300 NVL72 and GB200 NVL72 systems, HGX B300, B200, H200 and H100 servers, RTX PRO 6000 and L40S instances, plus CPU-only AMD EPYC Genoa and Intel Ice Lake VMs. The HGX H200 carries 141 GB of memory per GPU.
  • On-demand and preemptible VMs: preemptible VMs can run on all GPU VM platforms; they are cheaper but can be reclaimed.

Orchestration

  • Managed Kubernetes: a fully managed container orchestration layer that Nebius describes as pre-optimized for AI workloads.
  • Soperator (managed Slurm): Nebius says a production-scale Slurm cluster is ready in 20–30 minutes from a console configuration, with nodes, dependencies and NVIDIA drivers handled automatically.
  • Open-source code: Soperator is built in-house, open-sourced on GitHub, and can also be self-deployed on any cloud.
  • SkyPilot: Nebius hosts a SkyPilot API server, so team configs and job history stay in the cloud environment instead of on a local server.

Serverless AI

Serverless AI has three services. Devlabs gives interactive environments with Jupyter and VS Code, Jobs run containerized workloads that start, run and finish, and Endpoints serve custom models over HTTP. Serverless AI is billed only while workloads are running, with no idle GPU costs.

Storage

Nebius storage includes a shared filesystem, a WEKA filesystem, block volumes, local SSD disks and object storage with Standard, Intelligent and Enhanced tiers. Object storage speaks the Amazon S3 API, so existing S3 tools can be reused.

Nebius Token Factory

  • API access: Nebius Token Factory offers a playground and an OpenAI-compatible API for inference and fine-tuning. It supports text-to-text, embedding and vision model types.
  • Two deployment modes: serverless inference for evaluation and variable traffic, or dedicated endpoints on isolated capacity for control over performance, scaling, model configuration or region.
  • Base and Fast flavors: both return identical outputs but differ in token pricing and latency; you select Fast by appending -fast to the model name.
  • Post-training: you can adapt open models with your own data and then deploy the resulting model on Token Factory.

Guide

Getting started on Nebius AI Cloud

  1. Open the web console and sign in with a Google, GitHub or Microsoft account.
  2. Enter your name and email and accept the Terms of Use; Nebius then creates a project and a tenant automatically.
  3. In Billing, choose whether you pay for resources as a company or as an individual. Adding a bank card triggers a $25 charge that goes into your account balance.
  4. Create a VM under Compute → Virtual machines: on the Compute step select With GPUs, set the VM type to Preemptible if you want spot capacity, and pick a platform and preset.
  5. For a preemptible VM, either select a pricing policy that caps the hourly spot price or choose Follow spot price.
  6. When the work is finished, delete the VM and its volumes, because a stopped VM still accrues storage charges.

Getting started on Token Factory

  1. Sign up for Token Factory and create a billing account during onboarding; it cannot be skipped and needs a bank card.
  2. Create an API key and point any OpenAI-compatible client at the Token Factory base URL.
  3. Pick a model; add -fast to its name if latency matters more than token cost.
  4. To stop prompt retention, enable Zero Data Retention on your account profile page.

Nebius use cases and customer examples

The examples below come from customer stories that Nebius publishes on its own site; they are vendor-selected results, not independent audits.

  • Fraud detection and support automation: Revolut runs FinCrime agents, chat orchestration and PRAGMA, its event-based transaction model, on more than 200 NVIDIA H100 GPUs plus Token Factory.
  • Physical AI and robotics: RoboForce runs its physical AI workflow on Nebius AI Cloud with NVIDIA Blackwell infrastructure; the story credits the platform with a 70% cut in AI setup pipeline time.
  • Large-scale training on Slurm: Photoroom used Slurm clusters on Nebius, and its story reports multi-month training runs executed without interruption.
  • High-volume open-model serving: Prosus runs high-volume open-model workloads through Token Factory Dedicated Endpoints.

Who is it for

  • AI builders and enterprises: Nebius says it serves customers in healthcare and life sciences, robotics and physical AI, financial services, media and entertainment, retail and other industries.
  • Companies with long, stable workloads: commitment discounts are only available to accounts registered as a company, so individuals stay on pay-as-you-go.
  • ML teams without dedicated DevOps staff: the orchestration tools provision and configure clusters so ML engineers can schedule jobs without DevOps expertise.
  • Interruption-tolerant jobs: preemptible VMs suit workloads that can survive being stopped, because they do not guarantee availability.
  • Developers who want an API, not servers: Token Factory serverless inference lets you start through a familiar API without planning infrastructure upfront, aimed at evaluation, prototyping and variable workloads.
  • Startups: credit offerings are currently available only through Nebius's venture capital partners.

Nebius fits less well for buyers who need guaranteed on-demand capacity without signing a commitment, or for users who do not want to track and delete storage after each experiment.

Platforms

  • Interfaces: a web console, a CLI, a Terraform provider and an API. Object Storage is compatible with the Amazon S3 API, while the Nebius API, CLI and Terraform provider offer the most complete feature coverage.
  • Public regions: eu-north1 (Finland), eu-south1 (Madrid, Spain), eu-west1 and eu-west2 (France), me-west1 (Israel), us-central1 (Kansas City, Missouri), us-north1 (Woodbury, Minnesota), and uk-south1 and uk-south2 (United Kingdom).
  • Private region: eu-north2 in Iceland is available only to users who already have deployments there.
  • GPU placement: B300 VMs are only available in uk-south1, eu-west2 and us-north1, H200 VMs in eu-north1, eu-north2, eu-west1 and us-central1, and H100 VMs only in eu-north1, so the GPU you want decides your region.
  • Service coverage: service availability depends on the project region, and you can ask support to add a missing service to a region.
  • Token Factory regions: for dedicated endpoints the data-center region is fixed in the endpoint configuration, while public endpoints show a Global region.
  • Support: Nebius AI Cloud provides free technical support, with a 30-minute first-response target for critical tickets.

Pricing

Nebius GPU cloud pricing is pay-as-you-go by default. GPUs on running VMs are billed per second and priced per hour, so 30 minutes costs half the hourly rate. Prices are in USD for all customers except companies from Israel, which are billed in ILS. Listed prices exclude applicable taxes, including VAT.

GPU instance prices

InstancevCPUsRAM, GBOn-demand, GPU-hourGPU-hour (Effective October 1, 2026)
NVIDIA GB300 NVL72112800Contact usContact us
NVIDIA HGX B30024346$7.85$9.50
NVIDIA GB200 NVL72112800Contact usContact us
NVIDIA HGX B20020224$7.15$8.50
NVIDIA HGX H20016200$4.50$5.40
NVIDIA HGX H10016200$3.85$4.50
NVIDIA RTX PRO 600024218$1.80$1.80
NVIDIA L40S with Intel CPU8-4032-160from $1.55from $1.55
NVIDIA L40S with AMD CPU16-19296-1152from $1.82from $1.82

As captured on October 5, 2026, the AI Cloud price page shows both GPU-hour columns side by side, labeled "On-demand" and "Effective October 1, 2026". The Compute pricing documentation lists the H100 NVLink at $4.50 per GPU hour from October 1, 2026 and $3.85 before that date. The same documentation says the update covers VMs with B300, B200, H200 and H100 GPUs.

Spot, commitments and other charges

  • Spot prices disagree between pages: the price page lists preemptible HGX H100 capacity from $0.79 per GPU-hour and calls spot prices dynamic. The Compute pricing documentation lists a preemptible H100 NVLink price of $2.15 per GPU hour.
  • Spot volatility: the spot price may change as often as every 15 minutes, and its upper bound stays below the regular pay-as-you-go price for the same platform.
  • Commitments: reserving large-scale clusters for multiple months can cost up to 35% less than on-demand rates. Commitment discounts are generally prepaid.
  • Free items: Managed Kubernetes, Managed Soperator (free software with consumption-based pricing), egress and ingress traffic and public IP addresses are listed as free.
  • Storage: the shared filesystem costs $0.0800 per GiB per month and Standard object storage $0.0147 per GiB per month.
  • Credits: promo codes from Nebius special offers add free credits for AI Cloud resources.

Token Factory billing

Token Factory gives $1 in trial credit on first sign-up, valid for 30 days. The card is charged automatically at the start of a month with a negative balance or when the billing threshold is reached, and a failed card charge suspends the account.

Nebius alternatives

  • CoreWeave: sells Bare Metal Compute with direct hardware access and no virtualization overhead, a contrast with Nebius's non-virtualized-GPU VMs. CoreWeave also runs a Zero Egress Migration program for moving data in without egress fees.
  • Lambda: offers 1-Click Clusters of HGX B200 and H100 GPUs for distributed training, plus instances that spin up in minutes for prototyping.
  • AWS and other hyperscalers: Nebius advertises 43% better TCO for fine-tuning and 112% better TCO for inference versus AWS, both marked with an asterisk; these are the vendor's own figures.

An independent GPU-cloud rating report from September 2026 said Nebius still trails CoreWeave on many technical aspects and relationships with NVIDIA and frontier labs.

Limitations

Capacity and interruptions

  • No capacity guarantee on demand: the services agreement lets Nebius refuse on-demand and preemptible services based on available capacity, price, cluster utilization or other operational considerations.
  • Preemptible stops: preemptible VMs may be stopped at any time; Compute sends a SIGTERM signal 60 seconds before stopping the VM.
  • Fixed VM type: an existing VM cannot be switched between regular and preemptible; you have to create a new one.
  • Quotas: quotas start at default values and are set separately for each region. Nebius may also reduce quotas that stay well below their limit for a long time, after a grace period that is usually two days.

Billing risks

  • Storage keeps billing: stopped VMs are not charged for compute, but their storage volumes are still charged until deleted.
  • Threshold charges: if a card cannot cover the billing-threshold charge, the account and access to resources may be suspended.
  • Commitment exits: shortening or cancelling a commitment is only possible if the addendum allows it, and Nebius can refuse the request.

Token Factory boundaries

  • Rate limits: limits grow automatically with sustained usage but stop at 20 times the base allocation; beyond that Nebius requires an Enterprise plan.
  • Optimized models: Token Factory applies techniques such as quantization and states that optimized models keep approximately 99% of the original model's quality.
  • Public endpoints: they carry no region commitment and are not intended for production use cases that require a stable, predictable region.

Independent testing

An independent rating report published in September 2026 called Nebius an industry leader with strong offerings in every category and the ability to command a significant price premium. It also described, as anecdote, aggressive negotiations with high prices and prepayments reaching 100% on one-year commitments. In the report's own tests, automatic return of a failed GB300 node to service worked but took 8h40m. The testers also first saw slow 4 KiB writes and EEXIST errors on a cluster data filesystem until Nebius retuned it after their feedback.

Privacy, data and rights

  • Token Factory retention: unless Zero Data Retention is enabled, inputs and outputs are kept for speculative decoding, and Nebius states that content is not used to train models in either mode.
  • Zero Data Retention: with ZDR on, prompts and responses are not stored after each request, but enabling it may reduce service levels such as inference speed.
  • Storage location: retained Token Factory inputs and outputs are stored in Finland. All fine-tuning datasets, artifacts and model outputs are stored in EU data centers.
  • Ownership: you keep ownership of inputs, outputs and fine-tuned models, though outputs may not be unique and other users may receive similar ones.
  • AI Cloud data: under the services agreement, Nebius uses uploaded customer data solely to perform the agreement and related terms. After termination, customer data is marked and deleted with platform resources within 72 hours unless law requires a different period.
  • Sensitive data: customers must not upload sensitive personal data unless the service expressly supports it and additional terms are in place.
  • Certifications: Nebius lists SOC 2 Type II with HIPAA, ISO 27001, ISO 42001 and other standards, and states that it implements GDPR- and CCPA-compliant policies.

FAQ

Q1. Does Nebius bill GPUs by the second?

Yes. For GPUs on running VMs the billing unit is one second and the pricing unit is one hour.

Q2. Can an individual get commitment discounts?

No. Commitment discounts require an account registered as a company; individual accounts stay on pay-as-you-go billing.

Q3. Does Token Factory train models on my prompts?

Nebius states that Token Factory customer content is not used to train any models. Without Zero Data Retention, prompts and outputs are still stored and used for speculative decoding.

Q4. Is there a free trial?

Token Factory adds $1 of trial credit valid for 30 days. On AI Cloud, free credits come through promo codes or, for startups, through partner venture funds.

Know a Similar Tool?
If you know other great AI tools, feel free to submit them to us