Model Optimizer
Open SourceA unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural archi...
🐳 Self-Hostable🔓 No Sign-up Required⚡ Traction Signal: 88/100
★4,793 Stars / Upvotes
A unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural architecture search, speculative decoding, etc. It compresses deep learning models for downstream deployment frameworks like TensorRT-LLM, TensorRT, vLLM, etc. to optimize inference speed.
Compare other freshly discovered tools and open-source alternatives.