Skip to content

Summary

Open-source desktop app to run, serve and fine-tune text, image, video and audio models entirely on your own machine.

Screenshots

Description

Unsloth started as the fine-tuning library that made training open-weight models fast enough to fit on a single consumer GPU. Unsloth Desktop packages that work into a graphical app that runs and trains models locally on Mac, Windows and Linux — open source, free, and with nothing leaving the machine.

The app covers more than chat. It generates images and video locally through models such as FLUX, Z-Image and MiniMax-H3 with LoRA adapter support, plus Wan and LTX for video; on an NVIDIA B200 the project reports MiniMax-H3 producing a 960x544, 124-frame clip in 13 seconds at 8 steps, down from over 70. A built-in model hub handles downloads and picks quantisations that actually fit the device you are on, with day-zero support for new releases across the Qwen, GLM, Gemma and NVIDIA families.

What makes it useful beyond a local chat window is that it serves. Running unsloth start exposes an OpenAI-compatible endpoint, so Claude Code, Codex and any existing script or SDK can point at a local model instead of a paid API. Tool calling is supported with self-healing retries, Bash and Python execute in a sandbox, and a free Cloudflare tunnel can publish a local or Colab model over HTTPS when you need to reach it from elsewhere. Private web search and a Deep Research mode that produces cited reports round it out.

The whole stack is open source and free, with no paid tier: the trade-off is that throughput depends entirely on your own hardware. Aimed at developers and researchers who want to fine-tune on their own data, avoid per-token pricing, or keep sensitive material off third-party servers.

Reviews

Similar App Suggestions

App: Roomote

Roo Code

New

Self-hostable cloud coding agent from the Roo Code team — investigates repos, verifies its own work and opens pull requests, driven from Slack, Teams, Discord or Telegram rather than an IDE.

Coding & DevelopmentFreemium

An Apache-2.0 TypeScript framework for building AI agents — workflows, memory, RAG and evals — with a local studio and an agentic software factory on top.

Coding & DevelopmentFreemium

AI design engineer that generates distinctive UI designs and production code inside your own repo, Figma and design system.

Coding & DevelopmentFreemium

App: oMLX

Jun Kim

New

Menu-bar LLM inference server for Apple Silicon, with continuous batching and tiered KV caching that keeps local models fast enough for real coding work.

Coding & DevelopmentFree

Open-source local inference server that profiles your hardware, picks models that fit, and points your coding agent at them — free, private and offline.

Coding & DevelopmentFree