AI & Data › Machine Learning Basics
Neural Network
Layers of weighted connections that learn from data.
Also known as: neural network, artificial neural network, ANN
A neural network is a model built from layers of simple units: each unit takes a weighted sum of its inputs, adds a bias, and passes the result through a non-linear function. Stacking layers lets the network represent complicated functions, and the weights are what training adjusts.
inputs → [weighted sum + bias → activation] × layers → output
The non-linearity is essential. Without it, any stack of layers collapses to a single linear transformation and cannot learn anything beyond a straight-line relationship. With it, the network can approximate very flexible mappings given enough data and capacity.
The classic mistakes:
- Assuming more neurons always helps. Extra capacity without enough data memorises noise and overfits (see overfitting).
- Ignoring input scaling. Features on wildly different scales train slowly or unstably; normalise inputs before training.
- Choosing architecture by fashion. The right shape depends on the data’s structure (spatial, sequential, tabular); mismatched architectures waste capacity.
- Reading meaning into individual weights. Behaviour is distributed across many parameters; single weights rarely map to human concepts.
- Forgetting the training loop is part of the design. Learning rate, batch size and initialisation affect results as much as the architecture does.
The core idea to keep: a neural network is a parameterised function trained by adjusting its weights to reduce error. Everything else (architectures, optimisers, regularisation) is a way of making that adjustment work better on a particular kind of data.
When a network underperforms, check the boring parts first: input scaling, label quality, batch size and a learning rate that is neither too large nor too small. Architecture changes rarely fix problems that originate in the data or the training loop.