Run large language models 100% locally on your Mac. Private by design, fast, and fully offline.
A native, lightweight macOS app built on Apple MLX — your conversations never leave your device.
All inference runs locally on your Mac. No cloud, no account, no telemetry. Your prompts and responses stay private.
Built with Swift and AppKit on Apple MLX, optimized for Apple Silicon. Lightweight, responsive, and energy-efficient.
Watch responses generate token by token, and interrupt generation at any moment with a single tap.
A live context ring shows how much of the model's context window is in use as your conversation grows.
Download, switch between, and manage multiple models. MTP speculative decoding accelerates supported models.
Run up to four separate chat sessions side by side, each with its own context. Conversations live in memory only and are never written to disk.
Models are downloaded directly from Hugging Face on first use. sMLX does not redistribute model weights.
Mac with Apple Silicon (M1 or later) · macOS 14 Sonoma or later · Sufficient free disk space and unified memory for the selected model.