1 min read
Model Optimization Techniques: From Training to Deployment
Model optimization is both science and art. Start with the techniques that offer the best impact for your specific constraints.
5 articles
Model optimization is both science and art. Start with the techniques that offer the best impact for your specific constraints.
Tomorrow we'll explore model distillation techniques. LLM.int8() Paper PyTorch Quantization ONNX Runtime Quantization
Tomorrow we'll explore INT8 quantization in more detail. PyTorch Quantization bitsandbytes GPTQ Paper
Tomorrow we'll explore quantization basics in more detail. PyTorch Quantization Flash Attention torch.compile
ONNX Runtime is the inference engine I reach for when a Python-trained model needs to be deployed somewhere other than a Python service — a .NET…