Python Apache-2.0

omlx

LLM inference server with continuous batching & SSD caching for Apple Silicon — managed from the macOS menu bar

J

jundot

Dernière activité 29 sept. 2026
jundot/omlx

22,4 k

étoiles

1,9 k

forks

1,5 k

issues ouvertes

apple-siliconinference-serverllmmacosmlxopenai-api

Ce README est souvent en anglais.

oMLX

oMLX

LLM inference, optimized for your Mac
Continuous batching and tiered KV caching, managed directly from your menu bar.

License Python 3.11-3.13 Apple Silicon

junkim.dot@gmail.com · https://omlx.ai/me

Install · Quickstart · Features · Models · CLI Configuration · Benchmarks · oMLX.ai

English · 中文 · 한국어 · 日本語


oMLX Admin Dashboard

Every LLM server I tried made me choose between convenience and control. I wanted to pin everyday models in memory, auto-swap heavier ones on demand, set context limits - and manage it all from a menu bar.

oMLX persists KV cache across a hot in-memory tier and cold SSD tier - even when context changes mid-conversation, all past context stays cached and reusable across requests, making local LLMs practical for real coding work with tools like Claude Code. That's why I built it.

Install

macOS App

Download the .dmg from Releases, drag to Applications, done. The app includes in-app auto-update, so future upgrades are just one click. The macOS app also installs a lightweight ~/.omlx/bin/omlx CLI shim so terminal commands and Apple Shortcuts can control the app-managed server.

Homebrew

brew tap jundot/omlx https://github.com/jundot/omlx
brew install jundot/omlx/omlx --with-custom-kernel   # With native custom kernels (needs full Xcode)
# No full Xcode? brew install jundot/omlx/omlx installs without the kernels

# Upgrade to the latest version
brew update && brew upgrade omlx

# Run as a background service (auto-restarts on crash)
omlx start

# Optional: MCP (Model Context Protocol) support
/opt/homebrew/opt/omlx/libexec/bin/pip install mcp

From Source

git clone https://github.com/jundot/omlx.git
cd omlx
make install                # Editable install with the web UI and native custom kernels
# No full Xcode? make install-no-kernels installs without the kernels
make mcp                    # Optional: MCP (Model Context Protocol) support

Requires macOS 15.0+ (Sequoia), Python 3.11–3.13, and Apple Silicon (M1/M2/M3/M4/M5).

Note on native custom kernels: GLM-5.2, MiniMax M3, and Qwen3.5 need them. Without them those families silently fall back to much slower generic paths -- for GLM-5.2 the fused DSA prefill is roughly 30x faster with the kernels (measured 845 vs ~29 tok/s on an M3 Ultra), and the fallback also uses more memory (#2137). make install builds them and checks that each one loads. That needs full Xcode with the Metal toolchain (xcodebuild -downloadComponent MetalToolchain); Command Line Tools alone do not provide it (xcrun: error: unable to find utility "metal"). Without Xcode, use the official DMG, which ships the kernels precompiled, or make install-no-kernels. A plain pip install -e . also skips them. make kernels rebuilds only the kernels from scratch. Homebrew builds them with --with-custom-kernel, which also needs full Xcode. To verify any install:

python -c "from omlx.custom_kernels import native_kernel_status; print(native_kernel_status())"

Quickstart

macOS App

Launch oMLX from your Applications folder. The Welcome screen guides you through three steps - model directory, server start, and first model download. That's it. To connect OpenClaw, OpenCode, Codex, Hermes Agent, Copilot, or DeepSeek Harness, see Integrations.

oMLX Welcome Screen oMLX Menubar

CLI

# Managed background server (macOS app or Homebrew install)
omlx start
omlx stop
omlx restart

# Foreground server attached to this terminal
omlx serve --model-dir ~/models

The server discovers LLMs, VLMs, embedding models, and rerankers from subdirectories automatically. Any OpenAI-compatible client can connect to http://localhost:8000/v1. A built-in chat UI is also available at http://localhost:8000/admin/chat.

Homebrew Service

If you installed via Homebrew, you can run oMLX as a managed background service:

omlx start                    # Start via brew services
omlx stop                     # Stop
omlx restart                  # Restart

brew services start omlx    # Start (auto-restarts on crash)
brew services stop omlx     # Stop
brew services restart omlx  # Restart
brew services info omlx     # Check status

The service runs omlx serve with zero-config defaults (~/.omlx/models, port 8000). omlx start, omlx stop, and omlx restart are the portable lifecycle commands; Homebrew installs delegate them to brew services. To customize, either set environment variables (OMLX_MODEL_DIR, OMLX_PORT, etc.) or run omlx serve --model-dir /your/path once to persist settings to ~/.omlx/settings.json.

Logs are written to two locations:

  • Service log: $(brew --prefix)/var/log/omlx.log (stdout/stderr)
  • Server log: ~/.omlx/logs/server.log (structured application log)

Features

Supports text LLMs, vision-language models (VLM), OCR models, embeddings, and rerankers on Apple Silicon.

Admin Dashboard

Web UI at /admin for real-time monitoring, model management, chat, benchmark, and per-model settings. Supports English, Korean, Japanese, Chinese, French, Russian, Spanish, and Brazilian Portuguese. All CDN dependencies are vendored for fully offline operation.

oMLX Admin Dashboard

Experimental Multi-Mac Inference

Source builds can split one downloaded language model across unequal-memory Macs using MLX pipeline ranks over Ring or Thunderbolt RDMA/JACCL. The Cluster dashboard handles read-only peer discovery, strict SSH/runtime verification, byte-aware unequal shard planning, measured compute/link rebalancing, headroom-aware execution tuning, activation, and a live shard/performance map on both Macs. Interactive, balanced, and throughput profiles expose coalesced batching, prompt-cache affinity, rotating-KV limits, Ring connection tuning, and a capability-gated experimental token-only output path. See Distributed inference across Macs for setup, security boundaries, current limitations, and the physical-hardware validation checklist.

Vision-Language Models

Run VLMs with the same continuous batching and tiered KV cache stack as text LLMs. Supports multi-image chat, base64/URL/file image inputs, and tool calling with vision context. MiMo V2.6 checkpoints with bundled sidecars also accept sampled-frame video and 24 kHz audio. Qwen3.5, Qwen3.6 and Qwen3.8 checkpoints (dense and MoE) accept native video input as base64 video_url / input_video data URIs; video needs OpenCV (opencv-python-headless). oQ conversion of official MiMo V2.6 checkpoints preserves image and audio support. OCR models (DeepSeek-OCR, DOTS-OCR, GLM-OCR) are auto-detected with optimized prompts.

Tiered KV Cache (Hot + Cold)

Block-based KV cache management inspired by vLLM, with prefix sharing and Copy-on-Write. The cache operates across two tiers:

  • Hot tier (RAM): Frequently accessed blocks stay in memory for fast access.
  • Cold tier (SSD): When the hot cache fills up, blocks are offloaded to SSD in safetensors format. On the next request with a matching prefix, they're restored from disk instead of recomputed from scratch - even after a server restart.

oMLX Hot & Cold Cache

Continuous Batching

Handles concurrent requests through mlx-lm's BatchGenerator. Max concurrent requests is configurable via CLI or admin panel.

Claude Code Optimization

Runs smaller context models with Claude Code by reporting the model's real context window to auto-compact instead of scaling token counts, and SSE keep-alive prevents read timeouts during long prefill.

Multi-Model Serving

Load LLMs, VLMs, embedding models, and rerankers within the same server. Models are managed through a combination of automatic and manual controls:

  • LRU eviction: Least-recently-used models are evicted automatically when memory runs low.
  • Manual load/unload: Interactive status badges in the admin panel let you load or unload models on demand.
  • Model pinning: Pin frequently used models to keep them always loaded.
  • Per-model TTL: Set an idle timeout per model to auto-unload after a period of inactivity.
  • Process memory enforcement: Total memory limit (default: system RAM - 8GB) prevents system-wide OOM.

Per-Model Settings

Configure sampling parameters, chat template kwargs, TTL, model alias, model type override, and more per model directly from the admin panel. Changes apply immediately without server restart.

  • Model alias: set a custom API-visible name. /v1/models returns the alias, and requests accept both the alias and directory name.
  • Model type override: manually set a model as LLM or VLM regardless of auto-detection.
  • Profiles: save named bundles of per-model settings and switch between them from the admin panel. A profile can optionally be exposed as its own model: /v1/models then also lists <model>:<profile> (e.g. qwen3-8b:thinking), which serves on the same engine as the base model with the profile's settings overlaid per request — no extra memory, no reload. When the base model has an alias, the exposed ID is advertised as <alias>:<profile>; the directory-name form keeps working, just like for the base model.

oMLX Chat Template Kwargs

Built-in Chat

Chat directly with any loaded model from the admin panel. Supports conversation history, model switching, dark mode, reasoning model output, and image upload for VLM/OCR models.

oMLX Chat

Model Downloader

Search and download MLX models from HuggingFace directly in the admin dashboard. Browse model cards, check file sizes, and download with one click.

oMLX Model Downloader

Integrations

Set up OpenClaw, OpenCode, Codex, Hermes Agent, Copilot, Pi, and DeepSeek Harness directly from the admin dashboard with a single click. No manual config editing required.

oMLX Integrations

Performance Benchmark

One-click benchmarking from the admin panel. Measures prefill (PP) and text generation (TG) tokens per second, with partial prefix cache hit testing for realistic performance numbers.

oMLX Benchmark Tool

macOS Menubar App

Native Swift / SwiftUI menubar app (not Electron). Start, stop, and monitor the server without opening a terminal. Includes local usage history with per-model totals and an hourly heatmap, persistent serving stats (survives restarts), auto-restart on crash, and built-in auto-update.

oMLX Menubar Stats

API Compatibility

Drop-in replacement for OpenAI and Anthropic APIs. Supports streaming usage stats (stream_options.include_usage), llama.cpp-style prefill progress (return_progress), Anthropic adaptive thinking, and vision inputs (images as base64 or URL, video as base64).

Endpoint Description
POST /v1/chat/completions Chat completions (streaming)
POST /v1/completions Text completions (streaming)
POST /v1/messages Anthropic Messages API
POST /v1/embeddings Text embeddings
POST /v1/rerank Document reranking
POST /v1/systemone Typed decisions with decision models (TypeSafe System One)
GET /v1/models List available models
POST /tokenize, POST /detokenize vLLM-compatible tokenizer API (also under /v1)

The admin API for settings, models, downloads, quantization, and benchmarks takes the main API key as a Bearer token and also works in headless mode. See Admin API.

Tool Calling & Structured Output

Supports all function calling formats available in mlx-lm, JSON schema validation, and MCP tool integration. Tool calling requires the model's chat template to support the tools parameter. The following model families are auto-detected:

Model Family Format
Llama, Qwen, DeepSeek, etc. JSON <tool_call>
Qwen3.5 Series XML <function=...>
Gemma <start_function_call>
GLM (4.7, 5) <arg_key>/<arg_value> XML
MiniMax Namespaced <minimax:tool_call>
Mistral [TOOL_CALLS]
IFM K2 Horizon XML or JSON inside <ifm|tool_calls>. Requires omlx[grammar]
Kimi K2 <|tool_calls_section_begin|>
Longcat <longcat_tool_call>

Models not listed above may still work if their chat template accepts tools and their output uses a recognized <tool_call> XML format. For tool-enabled streaming, assistant text is emitted incrementally while known tool-call control markup is suppressed from visible content; structured tool calls are emitted after parsing the completed turn.

A bare JSON, EBNF, or regex grammar constrains the answer from the first generated token, so thinking_budget is ignored. Configure a compatible reasoning_parser to combine constrained output with a separate, budgeted reasoning phase.

Models

Point --model-dir at a directory containing MLX-format model subdirectories. Two-level organization folders (e.g., mlx-community/model-name/) are also supported.

~/models/
├── Step-3.5-Flash-8bit/
├── Qwen3-Coder-Next-8bit/
├── gpt-oss-120b-MXFP4-Q8/
├── Qwen3.5-122B-A10B-4bit/
└── bge-m3/

Models are auto-detected by type. You can also download models directly from the admin dashboard.

Type Models
LLM Any model supported by mlx-lm
VLM Qwen3.5 Series, GLM-4V, Pixtral, and other mlx-vlm models
OCR DeepSeek-OCR, DOTS-OCR, GLM-OCR
Embedding BERT, BGE-M3, ModernBERT, EmbeddingGemma 2
Reranker ModernBERT, XLM-RoBERTa
Decision Clef, Clef-Flash, OpenJev

CLI Configuration

# Managed background server (macOS app or Homebrew install)
omlx start
omlx stop
omlx restart

# Start with default settings (memory guard tier = balanced, manage via admin UI)
omlx serve --model-dir ~/models

# Choose a memory guard tier at startup
omlx serve --model-dir ~/models --memory-guard safe

# Set a custom memory guard ceiling in GB
omlx serve --model-dir ~/models --memory-guard-gb 48

# Enable SSD cache for KV blocks
omlx serve --model-dir ~/models --paged-ssd-cache-dir ~/.omlx/cache

# Set in-memory hot cache size
omlx serve --model-dir ~/models --hot-cache-max-size 20%

# Adjust max concurrent requests (default: 8)
omlx serve --model-dir ~/models --max-concurrent-requests 16

# With MCP tools
omlx serve --model-dir ~/models --mcp-config mcp.json

# HuggingFace mirror endpoint (for restricted regions)
omlx serve --model-dir ~/models --hf-endpoint https://hf-mirror.com

# API key authentication
omlx serve --model-dir ~/models --api-key your-secret-key
# Localhost-only: skip verification via admin panel global settings

# Network access requires authentication
OMLX_API_KEY=your-secret-key omlx serve --model-dir ~/models --host 0.0.0.0

# Inference and admin APIs only, without the web UI (not saved to settings)
omlx serve --model-dir ~/models --headless

The default SSD cache limit, auto, uses 50% of the sum of free disk space and existing SSD cache files, including GDN sidecars. The budget is refreshed during use and does not shrink simply because the cache grows or the server restarts. Other disk usage can change the budget. Set --paged-ssd-cache-max-size 20GB for a fixed limit.

Most settings can also be configured from the web admin panel at /admin. Settings are persisted to ~/.omlx/settings.json, and CLI flags take precedence. Set the main API key before changing the server host to a LAN address or 0.0.0.0, or save both settings together. oMLX refuses to start on any non-loopback address without a main API key. The existing skip_api_key_verification option remains restricted to loopback-only binds.

For keyless inference, stop oMLX, manually set auth.allow_unauthenticated_inference to true in settings.json, and restart. It defaults to false and has no UI toggle. This allows anyone who can reach the server to use inference (including stored Responses and audio), MCP tools, and web search. On network binds, keep a main API key configured and skip_api_key_verification set to false; management endpoints still require authentication.

Architecture
FastAPI Server (OpenAI / Anthropic API)
    │
    ├── EnginePool (multi-model, LRU eviction, TTL, manual load/unload)
    │   ├── BatchedEngine (LLMs, continuous batching)
    │   ├── VLMEngine (vision-language models)
    │   ├── EmbeddingEngine
    │   ├── RerankerEngine
    │   └── DecisionEngine
    │
    ├── ProcessMemoryEnforcer (total memory limit, TTL checks)
    │
    ├── Scheduler (FCFS, configurable concurrency)
    │   └── mlx-lm BatchGenerator
    │
    └── Cache Stack
        ├── PagedCacheManager (GPU, block-based, CoW, prefix sharing)
        ├── Hot Cache (in-memory tier, write-back)
        └── PagedSSDCacheManager (SSD cold tier, safetensors format)

Development

Build commands

Command What it does
make install Editable install of the server and the web UI, then the native custom kernels
make install-no-kernels The same without the kernels, for machines without full Xcode
make mcp Add MCP (Model Context Protocol) support
make dev Same as make install, with dev tools (make dev-no-kernels skips the kernels)
make kernels Delete any built native custom kernels, compile them again in place, and check that each one loads
make web Rebuild the web UI CSS and normalize the translation files
make app Stage a runnable oMLX.app with freshly compiled native custom kernels

The native custom kernels need full Xcode with the Metal toolchain (xcodebuild -downloadComponent MetalToolchain).

CLI Server

git clone https://github.com/jundot/omlx.git
cd omlx
make dev
pytest

Web UI

The web admin UI lives in apps/omlx-web/ and ships in the same package, so make dev installs it too. After editing its templates, JavaScript, or translations, run make web to rebuild the CSS and normalize the translation files.

macOS App

The native SwiftUI app lives at apps/omlx-mac/. Requires Xcode 26.5+ and Python 3.11+. venvstacks is declared as a dev dependency so make dev (or uv sync --dev) brings the pinned version in. The build script also falls back to uvx venvstacks or pipx run venvstacks if you prefer a host-global tool runner.

# Stage a runnable oMLX.app (xcodebuild + venvstacks Python layers + native kernels + ad-hoc sign)
make app

# Result lands at apps/omlx-mac/build/Stage/oMLX.app
open apps/omlx-mac/build/Stage/oMLX.app

# Force a fresh venvstacks rebuild (otherwise it's cached by fingerprint)
apps/omlx-mac/Scripts/build.sh release --rebuild-donor

First cold build takes 10–20 minutes (venvstacks Python layer assembly). Subsequent builds reuse the cached packaging/_export/ and finish in about 4 minutes. See packaging/README.md for the layer configuration and apps/omlx-mac/ for the Swift sources.

Contributing

Contributions are welcome! See Contributing Guide for details.

  • Bug fixes and improvements
  • Performance optimizations
  • Documentation improvements

License

Apache 2.0

Acknowledgments

  • MLX and mlx-lm by Apple
  • mlx-vlm - Vision-language model inference on Apple Silicon
  • vllm-mlx - oMLX started from vllm-mlx v0.1.0 and evolved significantly with multi-model serving, tiered KV caching, VLM with full paged cache support, an admin panel, and a macOS menu bar app
  • venvstacks - Portable Python environment layering for the macOS app bundle
  • mlx-embeddings - Embedding model support for Apple Silicon
  • dflash-mlx - Block diffusion speculative decoding on Apple Silicon
  • MTPLX - Lightning MTP's verify-shape Metal kernels are powered by MTPLX by Youssof Altoukhi, which also inspired the depth-k pipeline
  • mlx-serve - The fused GDN verify prework kernel is adapted from mlx-serve's port of the mlxfast-challenge qwen35_packed_gdn_prework kernel, and Qwen4's fused GDN decode and prefill kernels are adapted from mlx-serve's MIT-licensed transformer.zig; Qwen4 QSA's 128-bit K/V staging is adapted from mlx-serve's MIT-licensed msv_attn_p256 kernel
  • Splash - The verify-shape linear kernels use Splash's bf16 0x4300 | q weight operand with per-group input sums (from its Apache-2.0 linear_q4_sgmatrix.metal), and the tensor-op verify attention adapts the tile design of Splash's paged_attention_tile.h
  • SiliconScope - The menu bar statistics take their design and rendering approach from SiliconScope by Kennt Kim, which also inspired the energy-efficient re-render gating

Projets similaires

Rapid-MLX is an open-source (Apache 2.0) OpenAI- and Anthropic-compatible LLM inference server and Mac app for Apple Silicon, built on MLX, focused on reliable tool calling for coding agents. Release-gated with Claude Code, Codex CLI, Aider, Hermes and DeepSeek Harness.

Pythonanthropic-apiapple-siliconclaude-code
Rraullenchai
3,9 k étoiles424

High-performance OpenAI and Anthropic compatible LLM inference server for Apple Silicon. Native MLX, continuous batching, multimodal models, MCP tool calling, and Claude Code support.

Pythonanthropicanthropic-apiapple-silicon
Wwaybarrios
1,6 k étoiles226

MLX Omni Server is a local inference server powered by Apple's MLX framework, specifically designed for Apple Silicon (M-series) chips. It implements OpenAI-compatible API endpoints, enabling seamless integration with existing OpenAI SDK clients while leveraging the power of local ML inference.

Pythonfunction-callinggenaimlx
Mmadroidmaq
745 étoiles88