sMLX app icon

sMLX

Run large language models 100% locally on your Mac. Private by design, fast, and fully offline.

Explore Models Get Support

Why sMLX

A native, lightweight macOS app built on Apple MLX — your conversations never leave your device.

🔒

100% On-Device

All inference runs locally on your Mac. No cloud, no account, no telemetry. Your prompts and responses stay private.

⚡

Native & Fast

Built with Swift and AppKit on Apple MLX, optimized for Apple Silicon. Lightweight, responsive, and energy-efficient.

💬

Streaming Output

Watch responses generate token by token, and interrupt generation at any moment with a single tap.

📊

Context Visualization

A live context ring shows how much of the model's context window is in use as your conversation grows.

📦

Multi-Model Management

Download, switch between, and manage multiple models. MTP speculative decoding accelerates supported models.

🗂

Multiple Conversations

Run up to four separate chat sessions side by side, each with its own context. Conversations live in memory only and are never written to disk.

Built-in Models

Models are downloaded directly from Hugging Face on first use. sMLX does not redistribute model weights.

Qwen 3.6-27B (MTP)
4-bit · MTP speculative decoding · by Alibaba Qwen · MLX build by samwang0041
Apache 2.0
Qwen 3-8B
4-bit · by Alibaba Qwen · MLX build by mlx-community
Apache 2.0

System Requirements

Mac with Apple Silicon (M1 or later) · macOS 14 Sonoma or later · Sufficient free disk space and unified memory for the selected model.