Python

bit-inference

Running BitNet 1-bit LLM on Apple Silicon Mac (M4 Pro) with NEON fix

U

ukosoukoso

Dernière activité 30 janv. 2026
ukosoukoso/bit-inference

179

étoiles

17

forks

0

issues ouvertes

Ce README est souvent en anglais.

BitNet 1-Bit LLM on Apple Silicon

Running Microsoft's BitNet b1.58 2B model on Apple Silicon Mac.

BitNet Chat Demo

Choose Your Setup

Setup Platform Status
macos/ Mac (M1/M2/M3/M4) ✅ Ready
docker/ Linux Servers ✅ Ready
# Clone BitNet and apply the NEON fix
git clone --recursive https://github.com/microsoft/BitNet.git
cd BitNet/3rdparty/llama.cpp
curl -O https://raw.githubusercontent.com/ukosoukoso/bitnet-apple-silicon-demo/main/macos/neon-detection-fix.patch
patch -p1 < neon-detection-fix.patch
cd ../..

# Download model and build
huggingface-cli download microsoft/BitNet-b1.58-2B-4T-gguf --local-dir models/BitNet-b1.58-2B-4T
python setup_env.py -md models/BitNet-b1.58-2B-4T -q i2_s

# Run
./build/bin/llama-cli -m models/BitNet-b1.58-2B-4T/ggml-model-i2_s.gguf -p "Hello" -n 50 -t 8

The Problem We Fixed

BitNet's llama.cpp fork has a bug in NEON detection on macOS Apple Silicon:

// Original (broken)
sysctlbyname("hw.optional.AdvSIMD", ...)  // Returns 0 on M1/M2/M3/M4!

// Fixed
sysctlbyname("hw.optional.arm.AdvSIMD", ...)  // Returns 1 correctly

This caused NEON = 0 in system_info, leading to crashes or poor performance.

Benchmark (M4 Pro)

Metric Value
Model BitNet b1.58 2B (I2_S)
Size 1.10 GiB
NEON ✅ Enabled
Speed ~40-70 tokens/sec

Credits

License

MIT (same as BitNet)

Projets similaires

Official inference framework for 1-bit LLMs

C++
Mmicrosoft
40,4 k étoiles3,7 k

LLM inference server with continuous batching & SSD caching for Apple Silicon — managed from the macOS menu bar

Pythonapple-siliconinference-serverllm
Jjundot
22,4 k étoiles1,9 k

AirLLM 70B inference with single 4GB GPU

Jupyter Notebookchinese-llmchinese-nlpfinetune
Llyogavin
35,2 k étoiles3,7 k