Unsloth logo

Unsloth

Free tier

Run and Train Models Locally

Free tier available·All audiences·API available·Open source

Key strengths

Local-first AI model running and trainingOpen-source and free desktop appOpenAI-compatible API for agent integrationSignificant VRAM reduction (up to 90% on Enterprise)Significant training speed improvements (up to 32x on Enterprise)Day Zero support for latest modelsImage and video generation supportSelf-healing tool calls for improved accuracyCross-platform (Mac, Windows, Linux)
Free tier + paid plans
Founded 2023
Self-hostable
No ratings yet

Unsloth — Developer Use Cases

  • Local LLM fine-tuning: Run 4-bit or 16-bit LoRA fine-tuning jobs on a single consumer GPU with 60% less VRAM than standard approaches, scaling to multi-GPU (Pro) or multi-node (Enterprise) as workloads grow.
  • Agentic pipelines with local models: Expose locally running models via the OpenAI-compatible API to connect AI agents (e.g., built with Claude Code or Codex) to self-hosted inference endpoints, with self-healing tool calls for improved function-calling reliability.
  • Quantized model deployment: Export and serve models in GGUF or MLX formats for efficient local inference across Mac (Apple Silicon via MLX) and other platforms.
  • Secure remote model access: Use the Cloudflare tunnel integration to expose a local Unsloth instance over a secure tunnel for remote development or team access without cloud hosting.
  • Image and video generation: Run diffusion models locally for image and video generation workloads alongside text-based LLMs within the same environment.
  • High-throughput enterprise training: Leverage the Enterprise tier for 32x faster training, 5x faster inference, multi-node scaling, and up to 30% accuracy gains for production fine-tuning pipelines.