Python Apache-2.0

model-optimization

A toolkit to optimize ML models for deployment for Keras and TensorFlow, including quantization and pruning.

T

tensorflow

Dernière activité 24 sept. 2026
tensorflow/model-optimization

1,6 k

étoiles

348

forks

251

issues ouvertes

compressiondeep-learningkerasmachine-learningmlmodel-compressionoptimizationpruningquantizationquantized-networksquantized-neural-networksquantized-trainingsparsitytensorflow

Ce README est souvent en anglais.

TensorFlow Model Optimization Toolkit

The TensorFlow Model Optimization Toolkit is a suite of tools that users, both novice and advanced, can use to optimize machine learning models for deployment and execution.

Supported techniques include quantization and pruning for sparse weights. There are APIs built specifically for Keras.

For an overview of this project and individual tools, the optimization gains, and our roadmap refer to tensorflow.org/model_optimization. The website also provides various tutorials and API docs.

The toolkit provides stable Python APIs.

Installation

For installation instructions, see tensorflow.org/model_optimization/guide/install.

Contribution guidelines

If you want to contribute to TensorFlow Model Optimization, be sure to review the contribution guidelines. This project adheres to TensorFlow's code of conduct. By participating, you are expected to uphold this code.

We use GitHub issues for tracking requests and bugs.

Maintainers

Subpackage Maintainers
tfmot.clustering Arm ML Tooling
tfmot.quantization TensorFlow Model Optimization
tfmot.sparsity TensorFlow Model Optimization

Community

As part of TensorFlow, we're committed to fostering an open and welcoming environment.

  • TensorFlow Blog: Stay up to date on content from the TensorFlow team and best articles from the community.

Projets similaires

A unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural architecture search, speculative decoding, etc. It compresses deep learning models for downstream deployment frameworks like TensorRT-LLM, TensorRT, vLLM, etc. to optimize inference speed.

Python
NNVIDIA
5,1 k étoiles703

Model Compression Toolkit (MCT) is an open source project for neural network model optimization under efficient, constrained hardware. This project provides researchers, developers, and engineers advanced quantization and compression tools for deploying state-of-the-art neural networks.

Pythondeep-learningdeep-neural-networksedge-ai
SSonySemiconductorSolutions
453 étoiles80

SOTA low-bit LLM quantization (INT8/FP8/MXFP8/INT4/MXFP4/NVFP4) & sparsity; leading model compression techniques on PyTorch, TensorFlow, and ONNX Runtime

Pythonauto-tuningawqfp4
Iintel
2,7 k étoiles325