Neural network recipes are ordered training procedures: understand the data and prove the pipeline on a single batch before regularizing and tuning, one change at a time. The best-known version is Andrej Karpathy's A Recipe for Training Neural Networks, and its premise has aged well. A broken network rarely crashes. It trains and quietly underperforms.
Ordinary code fails loudly. A neural network fails quietly, so you need a procedure that checks each step before the next one hides it.
Karpathy makes two claims that explain why. The first is that a network is a leaky abstraction: the library hides backpropagation behind a tidy call, yet the moment you step off a standard example, you need to know what happens underneath.
The second is that training fails silently. Shuffle the images and their labels out of step, or forget to normalise the inputs, and the code still runs. The loss goes down a little, and nothing tells you it should have gone down a lot.
A recipe answers that with order. Every stage proves the one before it worked, so a bug surfaces at the stage that introduced it. Skip ahead to a large model with every trick switched on and a bad result could come from any of twenty places.
A network with a bug behaves like the colleague who never admits to being stuck. It nods through every stand-up, and three weeks later you learn it has been sorting the inbox by font size.
This guide runs the recipe in the order you should follow it, with a small PyTorch model you can copy and the checks that catch a bug in minutes instead of weeks.
How to build a neural network, starting with the data
Before writing a layer, look at the examples. Then build the whole pipeline around the dumbest model that could work.
Karpathy's first stage is to become one with the data, meaning hours spent scrolling through thousands of examples. You are hunting for duplicates and corrupt files, and for imbalance, where one class swamps the others. Small scripts that sort examples by label or length find most of it. What you see also shapes the model: if a human needs distant context to label an example, the network needs a way to see that far.
The second stage is an end-to-end skeleton with a trivial baseline. A few habits make it trustworthy:
- Fix the random seed, so two runs of the same code give the same numbers.
- Switch off data augmentation, which randomly alters training examples, until the plain version works.
- Check the loss before any training. For a softmax over n classes it should sit near −log(1/n), which is 2.30 for ten classes.
- Record a baseline, such as always predicting the most common class, so every later number has something to beat.
- Plot the inputs at the last moment before they enter the model, after every transform has run.
The baseline needs an honest score, and how a classification report is calculated walks through reading one without being fooled by accuracy.
How to determine input and output layers in a neural network
The data decides the input layer and the task decides the output layer. Only the hidden layers in between are yours to choose.
The input layer has one unit per feature in a single example, after any reshaping. A 28 by 28 greyscale image flattened into a row gives 784 inputs. A table with 12 numeric columns gives 12. Categorical columns grow when one-hot encoded, meaning one 0-or-1 column per category, so a field with five possible values becomes five inputs.
The output layer follows the question you are asking:
| Task | Output units | Typical loss |
|---|---|---|
| Regression (predict a number) | 1, no activation | Mean squared error |
| Binary classification | 1 logit | Binary cross-entropy |
| Classification into n classes | n logits | Cross-entropy |
A logit is the raw score the network produces before any probability is calculated from it.
For hidden layers, borrow. Google's Deep Learning Tuning Playbook advises starting from a well-established architecture and building something custom later. Karpathy's version of the same advice is "don't be a hero". A layout that worked on a similar problem beats one designed from first principles on day one.
A minimal PyTorch neural network in Python
Here is the smallest useful version of the recipe: a classifier trained on one batch, with a loop that should drive the loss to almost zero.
The layer sizes come from the PyTorch model-building tutorial, which classifies 28 by 28 clothing images into ten categories:
import torch
from torch import nn
torch.manual_seed(0)
model = nn.Sequential(
nn.Flatten(),
nn.Linear(28 * 28, 512), nn.ReLU(),
nn.Linear(512, 512), nn.ReLU(),
nn.Linear(512, 10),
)
loss_fn = nn.CrossEntropyLoss()
opt = torch.optim.Adam(model.parameters(), lr=3e-4)
x, y = next(iter(train_loader)) # one batch, reused every step
x, y = x[:8], y[:8]
for step in range(500):
loss = loss_fn(model(x), y)
opt.zero_grad()
loss.backward()
opt.step()
if step == 0 or step % 100 == 99:
print(step, round(loss.item(), 4))
Two numbers tell you whether the pipeline is sound. The first printed loss should land near 2.30, the −log(1/10) from earlier. The last should be close to zero, because eight examples are easy to memorise. A loss stuck above zero points to a bug in the labels or the loss function, and finding it now costs a coffee instead of a week.
Adam at a learning rate of 3e-4 is Karpathy's safe starting point. That number has been passed around so long it qualifies as a family recipe: nobody remembers who wrote it down, and everyone is slightly afraid to change it.
Overfit the training set, then regularize
First get a model large enough to memorise the training data. Then pull it back until it generalises to examples it has never seen.
Overfitting on purpose proves the model has enough capacity. Train on the full training set and watch the training loss fall well below the baseline. Add complexity one change at a time, and leave learning-rate decay off until the end, since a schedule tuned for one setup can quietly hurt another.
Once training loss is low, validation loss is the number to watch. Karpathy ranks the fixes roughly in this order:
- Collect more real data. Nothing else works as reliably.
- Augment the data, for example with flips and crops on images.
- Start from pretrained weights when a similar model exists.
- Add dropout, which zeroes random activations during training.
- Increase weight decay, a penalty on large weights.
- Stop early, at the point where validation loss stops improving.
Batch size belongs in a different drawer. The tuning playbook says to treat it as a hardware decision and use the largest batch your accelerator holds, since it mainly changes training speed.
Deciding whether training your own model makes sense at all comes before any of this, and when to build an LLM and when to skip it covers that call for language models specifically.
Tune, and plot the classification boundary of a neural network
Tuning searches for better hyperparameters, the settings you choose before training. Plotting the decision boundary shows what the model learned, on problems small enough to draw.
Karpathy prefers random search over grid search, because a grid wastes most of its runs on settings that barely matter. The playbook adds a mindset: spend most of your budget exploring how each setting affects results, and only a small share on squeezing out the final point. Once a configuration is settled, ensembles of several models and longer training runs add the last gains.
For a model with two input features, the classification boundary is the line where its prediction flips from one class to another. Predict over a fine grid and colour it:
import numpy as np
import matplotlib.pyplot as plt
xx, yy = np.meshgrid(np.linspace(-3, 3, 300), np.linspace(-3, 3, 300))
grid = torch.tensor(np.c_[xx.ravel(), yy.ravel()], dtype=torch.float32)
with torch.no_grad():
zz = model(grid).argmax(dim=1).reshape(xx.shape)
plt.contourf(xx, yy, zz, alpha=0.3)
plt.scatter(X[:, 0], X[:, 1], c=y, s=8)
plt.show()
A smooth boundary that follows the shape of the classes is healthy. One that wraps individual training points like cling film is memorising them, and the regularization list above is where to go next. With more than two features, plot two at a time and hold the rest at their average.
A neural network recipe card to keep
Pin this next to your training script and work down it in order:
- Read hundreds of raw examples before writing a model.
- Fix the seed and record a trivial baseline.
- Confirm the starting loss matches −log(1/n).
- Overfit one small batch to near-zero loss.
- Copy a proven architecture and overfit the full training set.
- Regularize, starting with more data.
- Tune with random search, then train longer.
The same discipline carries up to language models, where one silent bug costs days of GPU time. LLM Systems Engineering applies it to training and fine-tuning transformers, and AI Engineering places model training inside a full production system.
Like any recipe, this one works best when followed in order. Nobody ices a cake before baking it, and nobody should tune a learning rate before checking the labels line up.