Python Apache-2.0

TensorRT-Edge-LLM

High-performance, light-weight C++ LLM and VLM Inference Software for Physical AI

N

NVIDIA

Dernière activité 3 sept. 2026
NVIDIA/TensorRT-Edge-LLM

573

étoiles

137

forks

98

issues ouvertes

Ce README est souvent en anglais.

TensorRT Edge-LLM

High-Performance Large Language Model Inference Framework for NVIDIA Edge Platforms

Documentation version license

Overview   |   Support Matrix   |   Quick Start   |   Performance   |   Documentation   |   Roadmap


Latest News


Overview

TensorRT Edge-LLM is NVIDIA's C++ inference runtime for text, vision, audio, speech, and action models on NVIDIA Jetson, NVIDIA DRIVE, and NVIDIA DGX Spark. The supported frontend exports Hugging Face checkpoints to ONNX for C++ engine building; an experimental direct frontend builds engines from checkpoints without ONNX. Both paths use the same C++ deployment runtimes.


Getting Started

Check the Official Support Matrix, then follow the Quick Start Guide. Checkpoint IDs are listed in Supported Models.

For a supported target, install a published Python wheel without compiling Edge-LLM. Use tensorrt-edgellm[server] for the high-level Python API and HTTP serving; see the extras guide for export/tools dependencies and the minimal base workflow.

Install the base package from PyPI (requires a supported CUDA/TensorRT stack):

pip install tensorrt-edgellm==0.11.0

Documentation

Introduction

User Guide

Developer Guide

Software Design

Advanced Topics


Performance

See the Performance Benchmarks page for released benchmark results covering LLM and VLM prefill, generation throughput, memory usage, and EAGLE speculative decoding speedups.


Use Cases

🚗 Automotive

  • In-vehicle AI assistants
  • Voice-controlled interfaces
  • Scene understanding
  • Driver assistance systems

🤖 Robotics

  • Natural language interaction
  • Task planning and reasoning
  • Visual question answering
  • Human-robot collaboration

🏭 Industrial IoT

  • Equipment monitoring with NLP
  • Automated inspection
  • Predictive maintenance
  • Voice-controlled machinery

📱 Edge Devices

  • On-device chatbots
  • Offline language processing
  • Privacy-preserving AI
  • Low-latency inference

Follow our GitHub repository for the latest updates, releases, and announcements.


Support


License

Apache License 2.0


Contributing

We welcome contributions! Please see our Contributing Guidelines for details.


Projets similaires

TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtimes that orchestrate the inference execution in a performant way.

Pythonblackwellcudallm-serving
NNVIDIA
14,7 k étoiles2,8 k

LiteRT-LM is Google's production-ready, high-performance, open-source inference framework for deploying Large Language Models on edge devices.

C++edge-aion-device-aion-device-llm
Ggoogle-ai-edge
6,5 k étoiles734

LiteRT, successor to TensorFlow Lite. is Google's On-device framework for high-performance ML & GenAI deployment on edge platforms, via efficient conversion, runtime, and optimization

C++
Ggoogle-ai-edge
3,5 k étoiles464