Neural Networks for Beginners. Part Zero. Overview

Neural Networks for the Little Ones
Every time you say “Thank you” to a neural network, you launch a pipeline that multiplies hundreds of matrices with billions of elements, and burn as much electricity as an LED lamp in a few seconds.
This is the first article in a short series dedicated to networks for AI/ML clusters and HPC.
In this series, we’ll touch on the principles of model operation and training, parallelization, DMA and RDMA technologies, network topologies, InfiniBand and RoCE, and we’ll also philosophize on the topic of general and specialized solutions.
In this particular article, we’ll figure out what a neural network is, how it works, how it’s trained, and most importantly, why it needs hundreds of expensive GPU cards and some kind of special network.
The refrain of today’s story: there’s no magic in neural networks—it’s just a multitude of simple operations on numbers, performed on computers with special chips. There’s no magic in how they work, nor in the infrastructure they run on.


















