Neural Network Recipes: How to Train a Model That Learns

Cook weighing an ingredient on a digital kitchen scale, measuring each step the way a neural network recipe does
Photo by Anna Shvets on Pexels

Neural network recipes are ordered training procedures: understand the data and prove the pipeline on a single batch before regularizing and tuning, one change at a time. The best-known version is Andrej Karpathy's A Recipe for Training Neural Networks, and its premise has aged well. A broken network rarely crashes. It trains and quietly underperforms.

Ordinary code fails loudly. A neural network fails quietly, so you need a procedure that checks each step before the next one hides it.

Karpathy makes two claims that explain why. The first is that a network is a leaky abstraction: the library hides backpropagation behind a tidy call, yet the moment you step off a standard example, you need to know what happens underneath.

The second is that training fails silently. Shuffle the images and their labels out of step, or forget to normalise the inputs, and the code still runs. The loss goes down a little, and nothing tells you it should have gone down a lot.

A recipe answers that with order. Every stage proves the one before it worked, so a bug surfaces at the stage that introduced it. Skip ahead to a large model with every trick switched on and a bad result could come from any of twenty places.

A network with a bug behaves like the colleague who never admits to being stuck. It nods through every stand-up, and three weeks later you learn it has been sorting the inbox by font size.

This guide runs the recipe in the order you should follow it, with a small PyTorch model you can copy and the checks that catch a bug in minutes instead of weeks.

How to build a neural network, starting with the data

Before writing a layer, look at the examples. Then build the whole pipeline around the dumbest model that could work.

Karpathy's first stage is to become one with the data, meaning hours spent scrolling through thousands of examples. You are hunting for duplicates and corrupt files, and for imbalance, where one class swamps the others. Small scripts that sort examples by label or length find most of it. What you see also shapes the model: if a human needs distant context to label an example, the network needs a way to see that far.

The second stage is an end-to-end skeleton with a trivial baseline. A few habits make it trustworthy:

  • Fix the random seed, so two runs of the same code give the same numbers.
  • Switch off data augmentation, which randomly alters training examples, until the plain version works.
  • Check the loss before any training. For a softmax over n classes it should sit near −log(1/n), which is 2.30 for ten classes.
  • Record a baseline, such as always predicting the most common class, so every later number has something to beat.
  • Plot the inputs at the last moment before they enter the model, after every transform has run.

The baseline needs an honest score, and how a classification report is calculated walks through reading one without being fooled by accuracy.

How to determine input and output layers in a neural network

The data decides the input layer and the task decides the output layer. Only the hidden layers in between are yours to choose.

The input layer has one unit per feature in a single example, after any reshaping. A 28 by 28 greyscale image flattened into a row gives 784 inputs. A table with 12 numeric columns gives 12. Categorical columns grow when one-hot encoded, meaning one 0-or-1 column per category, so a field with five possible values becomes five inputs.

The output layer follows the question you are asking:

TaskOutput unitsTypical loss
Regression (predict a number)1, no activationMean squared error
Binary classification1 logitBinary cross-entropy
Classification into n classesn logitsCross-entropy

A logit is the raw score the network produces before any probability is calculated from it.

For hidden layers, borrow. Google's Deep Learning Tuning Playbook advises starting from a well-established architecture and building something custom later. Karpathy's version of the same advice is "don't be a hero". A layout that worked on a similar problem beats one designed from first principles on day one.

A minimal PyTorch neural network in Python

Here is the smallest useful version of the recipe: a classifier trained on one batch, with a loop that should drive the loss to almost zero.

The layer sizes come from the PyTorch model-building tutorial, which classifies 28 by 28 clothing images into ten categories:

import torch
from torch import nn

torch.manual_seed(0)

model = nn.Sequential(
    nn.Flatten(),
    nn.Linear(28 * 28, 512), nn.ReLU(),
    nn.Linear(512, 512), nn.ReLU(),
    nn.Linear(512, 10),
)
loss_fn = nn.CrossEntropyLoss()
opt = torch.optim.Adam(model.parameters(), lr=3e-4)

x, y = next(iter(train_loader))   # one batch, reused every step
x, y = x[:8], y[:8]

for step in range(500):
    loss = loss_fn(model(x), y)
    opt.zero_grad()
    loss.backward()
    opt.step()
    if step == 0 or step % 100 == 99:
        print(step, round(loss.item(), 4))

Two numbers tell you whether the pipeline is sound. The first printed loss should land near 2.30, the −log(1/10) from earlier. The last should be close to zero, because eight examples are easy to memorise. A loss stuck above zero points to a bug in the labels or the loss function, and finding it now costs a coffee instead of a week.

Adam at a learning rate of 3e-4 is Karpathy's safe starting point. That number has been passed around so long it qualifies as a family recipe: nobody remembers who wrote it down, and everyone is slightly afraid to change it.

Laptop showing Python code in a dim room with a coffee mug, the setting for running a PyTorch neural network recipe
Photo by Daniil Komov on Pexels

Overfit the training set, then regularize

First get a model large enough to memorise the training data. Then pull it back until it generalises to examples it has never seen.

Overfitting on purpose proves the model has enough capacity. Train on the full training set and watch the training loss fall well below the baseline. Add complexity one change at a time, and leave learning-rate decay off until the end, since a schedule tuned for one setup can quietly hurt another.

Once training loss is low, validation loss is the number to watch. Karpathy ranks the fixes roughly in this order:

  • Collect more real data. Nothing else works as reliably.
  • Augment the data, for example with flips and crops on images.
  • Start from pretrained weights when a similar model exists.
  • Add dropout, which zeroes random activations during training.
  • Increase weight decay, a penalty on large weights.
  • Stop early, at the point where validation loss stops improving.

Batch size belongs in a different drawer. The tuning playbook says to treat it as a hardware decision and use the largest batch your accelerator holds, since it mainly changes training speed.

Deciding whether training your own model makes sense at all comes before any of this, and when to build an LLM and when to skip it covers that call for language models specifically.

Tune, and plot the classification boundary of a neural network

Tuning searches for better hyperparameters, the settings you choose before training. Plotting the decision boundary shows what the model learned, on problems small enough to draw.

Karpathy prefers random search over grid search, because a grid wastes most of its runs on settings that barely matter. The playbook adds a mindset: spend most of your budget exploring how each setting affects results, and only a small share on squeezing out the final point. Once a configuration is settled, ensembles of several models and longer training runs add the last gains.

For a model with two input features, the classification boundary is the line where its prediction flips from one class to another. Predict over a fine grid and colour it:

import numpy as np
import matplotlib.pyplot as plt

xx, yy = np.meshgrid(np.linspace(-3, 3, 300), np.linspace(-3, 3, 300))
grid = torch.tensor(np.c_[xx.ravel(), yy.ravel()], dtype=torch.float32)
with torch.no_grad():
    zz = model(grid).argmax(dim=1).reshape(xx.shape)

plt.contourf(xx, yy, zz, alpha=0.3)
plt.scatter(X[:, 0], X[:, 1], c=y, s=8)
plt.show()

A smooth boundary that follows the shape of the classes is healthy. One that wraps individual training points like cling film is memorising them, and the regularization list above is where to go next. With more than two features, plot two at a time and hold the rest at their average.

A neural network recipe card to keep

Pin this next to your training script and work down it in order:

  1. Read hundreds of raw examples before writing a model.
  2. Fix the seed and record a trivial baseline.
  3. Confirm the starting loss matches −log(1/n).
  4. Overfit one small batch to near-zero loss.
  5. Copy a proven architecture and overfit the full training set.
  6. Regularize, starting with more data.
  7. Tune with random search, then train longer.

The same discipline carries up to language models, where one silent bug costs days of GPU time. LLM Systems Engineering applies it to training and fine-tuning transformers, and AI Engineering places model training inside a full production system.

Like any recipe, this one works best when followed in order. Nobody ices a cake before baking it, and nobody should tune a learning rate before checking the labels line up.

Frequently asked questions

What is the delta in neural networks?

In backpropagation, the delta of a neuron is its error signal: the gradient of the loss with respect to that neuron's input before activation. Deltas are calculated at the output layer first and passed backwards, layer by layer, to work out how much each weight should change. The name comes from the delta rule, an early learning rule for single-layer networks.

Are LLMs neural networks?

Yes. A large language model is a neural network, almost always a transformer, trained to predict the next token of text. It is built from the same parts as the small classifier above, such as linear layers and activations, at the scale of billions of parameters.

Is ChatGPT a neural network?

ChatGPT is a product built on neural networks. At its core is a GPT model, a transformer-based language model, wrapped in software that handles the conversation and safety filtering. The responses themselves come from the neural network predicting text one token at a time.

How do you create a neural network?

Define the layers in a framework such as PyTorch or Keras, choosing the input size from your data and the output size from your task. Pick a loss function and an optimiser, then train on batches of examples. Before training on everything, confirm the model can overfit one small batch, which proves the pipeline works.

Can a neural network write cooking recipes?

Yes. A language model trained on recipe text will generate new recipes, and general chat models do it on request. The results read convincingly, yet quantities and cooking times can be wrong, so treat a generated recipe as a draft and check it before you cook.

Sources