C++ Apache-2.0

LiteRT-LM

LiteRT-LM is Google's production-ready, high-performance, open-source inference framework for deploying Large Language Models on edge devices.

G

google-ai-edge

Dernière activité 29 sept. 2026
google-ai-edge/LiteRT-LM

6,5 k

étoiles

734

forks

678

issues ouvertes

edge-aion-device-aion-device-llm

Ce README est souvent en anglais.

LiteRT-LM

LiteRT-LM is Google's production-ready orchestration layer to run LLMs with LiteRT, engineered for high-performance, cross-platform execution.

🔗 Product Website | 💬✨ Chat Demo | 🔎✨ Embedding Demo

🔥 What's New: v0.18.0

  • ✨ New Model Release (EmbeddingGemma 2): Shipped multimodal EmbeddingGemma 2 supporting text, vision, and audio embeddings with Matryoshka dimension truncation across Python, Kotlin, Swift, Web (JavaScript), and C++. Check out our DevSite documentation for more details!
  • 🛠️ CLI & Developer Experience: Added fast model imports (litert-lm import) and an OpenAI-compatible /v1/embeddings endpoint (litert-lm serve) for multimodal embedding models.
  • 🔍 Model Info API: Added ModelInfo and litert-lm describe for full model introspection—covering metadata, capabilities, and runtime requirements before loading.
  • ⚡ NPU & GPU Acceleration: Enabled dynamic on-demand KV cache growth on NPU and attention mask pruning optimizations on GPU.

👉 Try Gemma4-E4B with MTP on Linux, macOS, Windows or Raspberry Pi with the LiteRT-LM CLI:

litert-lm run  \
   --from-huggingface-repo=litert-community/gemma-4-E4B-it-litert-lm \
   gemma-4-E4B-it.litertlm \
   --backend=gpu \
   --enable-speculative-decoding=true \
   --prompt="What is the capital of France?"

🌟 Key Features

  • 📱 Cross-Platform Support: Android, iOS, Web, Desktop, and IoT (e.g. Raspberry Pi).
  • 🚀 Hardware Acceleration: Peak performance via GPU and NPU accelerators.
  • 👁️ Multi-Modality: Support for vision and audio inputs.
  • 🔧 Tool Use: Function calling support for agentic workflows.
  • 📚 Broad Model Support: Gemma, Llama, Phi-4, Qwen, and more.


🚀 Production-Ready for Google's Products

LiteRT-LM powers on-device GenAI experiences in Chrome, Chromebook Plus, Pixel Watch, and more.

You can also try the Google AI Edge Gallery app to run models immediately on your device.

Install the app today from Google Play Install the app today from App Store
Get it on Google Play Download on the App Store

📰 Blogs & Announcements

Link Description
Bringing Gemma 4 12B to your Laptop: Unlocking Local, Agentic Workflows with Google AI Edge Bring agentic, multimodal AI capabilities to everyday laptops, enabling local data processing and visual insight generation.
Blazing-fast on-device GenAI with LiteRT-LM Unlock Gemma 4's full potential with blazing speed and incredible efficiency using newly added Swift, JavaScript, and Flutter APIs.
Accelerating Gemma 4: faster inference with multi-token prediction drafters An overview of how Multi-Token Prediction (MTP) drafters are making Gemma 4 models up to 3x faster at inference.
Bring state-of-the-art agentic skills to the edge with Gemma 4 Deploy Gemma 4 in-app and across a broader range of devices with stellar performance and broad reach using LiteRT-LM.
On-device GenAI in Chrome, Chromebook Plus and Pixel Watch Deploy language models on wearables and browser-based platforms using LiteRT-LM at scale.
On-device Function Calling in Google AI Edge Gallery Explore how to fine-tune FunctionGemma and enable function calling capabilities powered by LiteRT-LM Tool Use APIs.
Google AI Edge small language models, multimodality, and function calling Latest insights on RAG, multimodality, and function calling for edge language models.

🏃 Quick Start

⚡ Quick Try (No Code)

Try LiteRT-LM immediately from your terminal without writing a single line of code using uv:

uv tool install litert-lm

litert-lm run \
  --from-huggingface-repo=google/gemma-3n-E2B-it-litert-lm \
  gemma-3n-E2B-it-int4 \
  --prompt="What is the capital of France?"

📚 Supported Language APIs

Ready to get started? Explore our language-specific guides and setup instructions.

Language Status Best For... Documentation
Python ✅ Stable Prototyping & Scripting Python Guide
Kotlin ✅ Stable Android apps & JVM Kotlin Guide
Swift 🚀 Early Preview Native iOS & macOS Swift Guide
JavaScript (web) 🚀 Early Preview Browser environments JavaScript Guide
Flutter 🚀 Community Cross-platform mobile Flutter Guide
C++ ✅ Stable High-performance native C++ Guide

🏗️ Build From Source

This guide shows how you can compile LiteRT-LM from source. If you want to build the program from source, you should checkout the stable Latest Release tag.


Projets similaires

LiteRT, successor to TensorFlow Lite. is Google's On-device framework for high-performance ML & GenAI deployment on edge platforms, via efficient conversion, runtime, and optimization

C++
Ggoogle-ai-edge
3,5 k étoiles464

High-performance, light-weight C++ LLM and VLM Inference Software for Physical AI

Python
NNVIDIA
573 étoiles137

LiteRT and LiteRT-LM sample apps, model recipes, agent skills and utilities.

Pythonedge-aigenerative-aion-device-ai
Ggoogle-ai-edge
447 étoiles121