The story of making neural networks smaller and faster (number formats, quantization, pruning, and so on) has moved and is now serialized on the company blog. The write-ups and resource notes all continue there.
Where to read it
- Model Efficiency series: blog.caveduck.io/page/efficient-ml-series
Its backbone is the free MIT 6.5940 course by Song Han. The details continue at the link above.