Quantization FP32 to INT8, pruning, ONNX, TensorFlow Lite, latency vs accuracy, batch inference, caching, and monitoring.
AdvancedAbout 6 hours6 lessons
What you will learn
Quantization, FP32 to INT8
Pruning
ONNX and TensorFlow Lite
Latency vs accuracy, animated
Caching and monitoring
Curriculum
1
Make it small and fast
2 lessonsFree preview
Quantization, FP32 to INT8Preview9m
PruningPreview8m
2
Formats and serving
3 lessons
ONNX and TensorFlow Lite9m
Latency vs accuracy, animated9m
Caching and monitoring8m
3
Hands-on checkpoint
1 lessons
Checkpoint: quantize and benchmark a model20m
Hands-on checkpoint
Every course ends with a Colab task you complete yourself. Our Gemini-assisted review gives feedback and points you to what to fix, it coaches you, it does not do it for you.