Python MIT

mlx-omni-server

MLX Omni Server is a local inference server powered by Apple's MLX framework, specifically designed for Apple Silicon (M-series) chips. It implements OpenAI-compatible API endpoints, enabling seamless integration with existing OpenAI SDK clients while leveraging the power of local ML inference.

M

madroidmaq

Dernière activité 9 mai 2026
madroidmaq/mlx-omni-server

745

étoiles

88

forks

25

issues ouvertes

function-callinggenaimlxopenaiopenai-apistructured-outputstttoolstts

Ce README est souvent en anglais.

MLX Omni Server

Local AI inference server optimized for Apple Silicon

PyPI version Python 3.11+ License: MIT Ask DeepWiki

MLX Omni Server Banner

MLX Omni Server provides dual API compatibility with both OpenAI and Anthropic APIs, enabling seamless local inference on Apple Silicon using the MLX framework.

Installation • Quick Start • Documentation • Contributing

✨ Features

  • 🚀 Apple Silicon Optimized - Built on MLX framework for M1/M2/M3/M4 chips
  • 🔌 Dual API Support - Compatible with both OpenAI and Anthropic APIs
  • 🎯 Complete AI Suite - Chat, audio processing, image generation, embeddings
  • ⚡ High Performance - Local inference with hardware acceleration
  • 🔐 Privacy-First - All processing happens locally on your machine
  • 🛠 Drop-in Replacement - Works with existing OpenAI and Anthropic SDKs

🚀 Installation

pip install mlx-omni-server

⚡ Quick Start

  1. Start the server:

    mlx-omni-server
  2. Choose your preferred API:

    OpenAI API (Click to expand)
    from openai import OpenAI
    
    client = OpenAI(
        base_url="http://localhost:10240/v1",
        api_key="not-needed"
    )
    
    response = client.chat.completions.create(
        model="mlx-community/gemma-3-1b-it-4bit-DWQ",
        messages=[{"role": "user", "content": "Hello!"}]
    )
    print(response.choices[0].message.content)
    Anthropic API (Click to expand)
    import anthropic
    
    client = anthropic.Anthropic(
        base_url="http://localhost:10240/anthropic",
        api_key="not-needed"
    )
    
    message = client.messages.create(
        model="mlx-community/gemma-3-1b-it-4bit-DWQ",
        max_tokens=1000,
        messages=[{"role": "user", "content": "Hello!"}]
    )
    print(message.content[0].text)

🎉 That's it! You're now running AI locally on your Mac.

📋 API Support

OpenAI Compatible Endpoints (/v1/*)

Endpoint Feature Status
/v1/chat/completions Chat with tools, streaming, structured output ✅
/v1/audio/speech Text-to-Speech ✅
/v1/audio/transcriptions Speech-to-Text ✅
/v1/images/generations Image Generation ✅
/v1/embeddings Text Embeddings ✅
/v1/models Model Management ✅

Anthropic Compatible Endpoints (/anthropic/v1/*)

Endpoint Feature Status
/anthropic/v1/messages Messages with tools, streaming, thinking mode ✅
/anthropic/v1/models Model listing with pagination ✅

⚙️ Configuration

# Default (port 10240)
mlx-omni-server

# Custom options
mlx-omni-server --port 8000
MLX_OMNI_LOG_LEVEL=debug mlx-omni-server

# View all options
mlx-omni-server --help

🛠 Development

Development Setup
git clone https://github.com/madroidmaq/mlx-omni-server.git
cd mlx-omni-server
uv sync

# Start with hot-reload
uv run uvicorn mlx_omni_server.main:app --reload --host 0.0.0.0 --port 10240

Testing:

uv run pytest                    # All tests
uv run pytest tests/chat/openai/ # OpenAI tests
uv run pytest tests/chat/anthropic/ # Anthropic tests

Code Quality:

uv run black . && uv run isort . # Format code
uv run pre-commit run --all-files # Run hooks

🎯 Key Features

Model Management

  • Auto-discovery of MLX models in HuggingFace cache
  • On-demand loading and intelligent caching
  • Automatic model downloading when needed

Advanced Capabilities

  • Function calling with model-specific parsers
  • Real-time streaming for both APIs
  • JSON schema validation and structured output
  • Extended reasoning (thinking mode) for supported models

📚 Documentation

Resource Description
OpenAI API Guide Complete OpenAI API reference
Anthropic API Guide Complete Anthropic API reference
Examples Practical usage examples

🔍 Troubleshooting

Common Issues

Requirements:

  • Python 3.11+
  • Apple Silicon Mac (M1/M2/M3/M4)
  • MLX framework installed

Quick fixes:

# Check requirements
python --version  # Should be 3.11+
python -c "import mlx; print(mlx.__version__)"

# Pre-download models (if needed)
huggingface-cli download mlx-community/gemma-3-1b-it-4bit-DWQ

# Enable debug logging
MLX_OMNI_LOG_LEVEL=debug mlx-omni-server

🤝 Contributing

Quick contributor setup:

git clone https://github.com/madroidmaq/mlx-omni-server.git
cd mlx-omni-server
uv sync && uv run pytest

🙏 Acknowledgments

Built with MLX by Apple • FastAPI • MLX-LM

📄 License

MIT License • Not affiliated with OpenAI, Anthropic, or Apple

🌟 Star History

Star History Chart

Projets similaires

LLM inference server with continuous batching & SSD caching for Apple Silicon — managed from the macOS menu bar

Pythonapple-siliconinference-serverllm
Jjundot
22,4 k étoiles1,9 k

Rapid-MLX is an open-source (Apache 2.0) OpenAI- and Anthropic-compatible LLM inference server and Mac app for Apple Silicon, built on MLX, focused on reliable tool calling for coding agents. Release-gated with Claude Code, Codex CLI, Aider, Hermes and DeepSeek Harness.

Pythonanthropic-apiapple-siliconclaude-code
Rraullenchai
3,9 k étoiles424

A high-performance API server that provides OpenAI-compatible endpoints for MLX models. Developed using Python and powered by the FastAPI framework, it provides an efficient, scalable, and user-friendly solution for running MLX-based vision and language models locally with an OpenAI-compatible interface.

Pythonapple-siliconcontinuous-batchingfastapi
Ccubist38
362 étoiles69