Unsloth is an open-source framework for running and training LLMs, and it lets you run and train AI models on your own local hardware through an open-source UI. It covers text, vision, diffusion image and video, audio and embedding models, for both local inference and LLM fine-tuning.
Unsloth can be installed in three distinct ways: Unsloth Desktop as the native app, Unsloth Studio as the browser-based web UI, or Unsloth Core as the original code-based Python package. Desktop and Studio are the no-code interfaces; Core is the library that developers call from their own training scripts and notebooks.
Unsloth Desktop is labelled Beta and described as a free, open-source app for running and training AI models on your own hardware, available for macOS, Windows and Linux. Unsloth Studio is also in Beta, positioned as an open-source, no-code web UI for training, running and exporting open models in one local interface. The homepage frames Unsloth Desktop with three words: open-source, free and 100% local. A 2024 business-press profile reported that the company's co-founders are brothers Michael and Daniel Han, who were then taking part in Y Combinator.
Unsloth's headline claim is that it trains LLMs, diffusion, TTS and embedding models 2× faster with 70% less VRAM and no accuracy loss. This is the vendor's own figure, and the test conditions behind it are set out below.
The published training benchmarks were run on H100 and Blackwell GPUs using the Alpaca dataset, a batch size of 2, gradient accumulation steps of 4, rank 32 and QLoRA on all linear layers. Against a Hugging Face + FA2 baseline, the table lists 2x speed, more than 75% VRAM reduction and 13x longer context for Llama 3.3 (70B), and 2x speed, more than 70% VRAM reduction and 12x longer context for Llama 3.1 (8B), both on 80GB GPUs.
Unsloth attributes its long-context results to its gradient checkpointing algorithm plus Apple's CCE algorithm, and says the more data there is, the less VRAM it uses. These are vendor benchmarks, not independent measurements.
Unsloth Studio exports models, including fine-tuned ones, to safetensors or GGUF for use with llama.cpp, vLLM, Ollama, LM Studio and other runtimes. That export step is what moves a model trained in Unsloth into other software.
In the other direction, Unsloth serves local models through an OpenAI-compatible API, and it can also connect a ChatGPT/Codex subscription and cloud providers. The unsloth run command serves a model and changes settings such as context size, GPU layers, threading, sampling, networking and tool configuration.
Studio can also import a YAML training config and pre-fill the relevant settings instead of starting from the presets.
By default unsloth studio binds to 127.0.0.1, so only the local machine can reach it. The --secure option serves it only through a free Cloudflare HTTPS link and does not start at all if the tunnel cannot come up, so the raw port is never exposed.
unsloth start command, so an agent workflow can run against a model on your own GPU.| Plan | Listed price | Speed boost | VRAM reduction | GPU support |
|---|---|---|---|---|
| Free | Free | 2x | 60% | Single |
| Unsloth Pro | Not displayed (Contact us) | 2.5x no. of GPUs | 80% | Multi |
| Unsloth Enterprise | Not displayed (Contact us) | 32x no. of GPUs | 90% | Multi + node |
Unsloth Pro is pitched at 2.5x faster training, 20% less VRAM and support for up to 8 GPUs, and Unsloth Enterprise at 30x faster training, multi-node support and 30% accuracy, but both show only a Contact us button, so no public price is displayed for either. Enterprise additionally lists up to +30% accuracy, 5x faster inference, full training, all Pro features, multi-node support and customer support.
The plan cards read differently from the documentation. The Free card describes itself as freeware of the standard version of Unsloth, marks it open-source, names Mistral, Gemma and Llama 1, 2, 3 support, and still lists MultiGPU as coming soon. The plan-difference table rates the Free tier at a 2x speed boost and 60% VRAM reduction, which differs from the 2× faster, 70% less VRAM figure used across the documentation.
Licensing is part of what "free" means here. Unsloth uses a dual-licensing model: the core Unsloth package remains Apache 2.0, while certain optional components such as the Unsloth Studio UI are AGPL-3.0. The company says this structure helps support ongoing development while keeping the project open source.
An independent 2026 comparison of open-source fine-tuning frameworks sets Unsloth against three other projects and assigns each a different job:
These choices are not mutually exclusive: LLaMA-Factory can run Unsloth as a backend, TRL ships an Unsloth integration, and Axolotl calls TRL trainers internally. Choosing LLaMA-Factory or TRL can therefore still mean running Unsloth underneath.
--disable-tools flag.studio/* and unsloth_cli/*).Unsloth Desktop is described as a free, open-source app, and the Free plan is the standard version. Unsloth Pro and Unsloth Enterprise exist, but the pricing page shows Contact us instead of a price.
Not for every task. Chat and Data Recipes work on CPU, and Studio training supports NVIDIA, AMD, MLX and Intel devices. Training on Windows lists an NVIDIA GPU as a requirement.
Yes. Models you already downloaded are found automatically, and you can specify custom folders if yours are not detected.
Yes. Unsloth exports to GGUF or safetensors for runtimes including llama.cpp, Ollama, vLLM and LM Studio, as long as the chat template matches the one used in training.
The core package is Apache 2.0, while optional components such as the Unsloth Studio UI are AGPL-3.0.