Python Apache-2.0

rtp-llm

RTP-LLM: Alibaba's high-performance LLM inference engine for diverse applications.

A

alibaba

Dernière activité 29 sept. 2026
alibaba/rtp-llm

1,4 k

étoiles

278

forks

228

issues ouvertes

gptinferencellamallmllm-servingllmopsmodel-serving

Ce README est souvent en anglais.

logo

license issue resolution open issues


| Documentation | Contact Us |

News

  • [2025/09] 🔥 RTP-LLM 0.2.0 release with enhanced performance and new features
  • [2025/01] 🚀 RTP-LLM now supports Prefill/Decode separation with detailed technical report
  • [2025/01] 🌟 Qwen series model and bert embedding model now supported on Yitian ARM CPU
  • [2024/06] 🔄 Major refactor: Scheduling and batching framework rewritten in C++, complete GPU memory management, and new Device backend
  • [2024/06] 🏗️ Multi-hardware support in development: AMD ROCm, Intel CPU and ARM CPU support coming soon
More

About

RTP-LLM is a Large Language Model (LLM) inference acceleration engine developed by Alibaba's Foundation Model Inference Team. It is widely used within Alibaba Group, supporting LLM service across multiple business units including Taobao, Tmall, Idlefish, Cainiao, Amap, Ele.me, AE, and Lazada.

RTP-LLM is a sub-project of the havenask project.

Key Features

🏢 Production Proven

Trusted and deployed across numerous LLM scenarios:

⚡ High Performance

  • Utilizes high-performance CUDA kernels, including PagedAttention, FlashAttention, FlashDecoding, etc.
  • Implements WeightOnly INT8 Quantization with automatic quantization at load time
  • Support WeightOnly INT4 Quantization with GPTQ and AWQ
  • Adaptive KVCache Quantization
  • Detailed optimization of dynamic batching overhead at the framework level
  • Specially optimized for the V100 GPU

🔧 Flexibility and Ease of Use

  • Seamless integration with the HuggingFace models, supporting multiple weight formats such as SafeTensors, Pytorch, and Megatron
  • Deploys multiple LoRA services with a single model instance
  • Handles multimodal inputs (combining images and text)
  • Enables multi-machine/multi-GPU tensor parallelism
  • Supports P-tuning models

🚀 Advanced Acceleration Techniques

  • Loads pruned irregular models
  • Contextual Prefix Cache for multi-turn dialogues
  • System Prompt Cache
  • Speculative Decoding

Getting Started

Benchmark and Performance

Learn more about RTP-LLM's performance in our benchmark reports:

Acknowledgments

Our project is mainly based on FasterTransformer, and on this basis, we have integrated some kernel implementations from TensorRT-LLM. We also draw inspiration from vllm, transformers, llava, and qwen-vl. We thank these projects for their inspiration and help.

Citation

If you find RTP-LLM useful in your research or project, please consider citing:

@Misc{rtp-llm,
  author       = {Alibaba},
  title        = {RTP-LLM: A High-Performance LLM Inference Engine},
  howpublished = {\url{https://github.com/alibaba/rtp-llm}},
  year         = {2025},
}

Contact Us

DingTalk Group

WeChat Group

Projets similaires

A high-throughput and memory-efficient inference and serving engine for LLMs

Pythonamdblackwellcuda
Vvllm-project
92,9 k étoiles22,8 k

LightLLM is a Python-based LLM (Large Language Model) inference and serving framework, notable for its lightweight design, easy scalability, and high-speed performance.

Pythondeep-learninggptllama
MModelTC
4,3 k étoiles368

20+ high-performance LLMs with recipes to pretrain, finetune and deploy at scale.

Pythonaiartificial-intelligencedeep-learning
LLightning-AI
13,7 k étoiles1,5 k