The story of making neural networks smaller and faster (number formats, quantization, pruning, and so on) has moved and is now serialized on the company blog. The write-ups and resource notes all continue there.

Where to read it

Its backbone is the free MIT 6.5940 course by Song Han. The details continue at the link above.