How Is a Classification Report in Deep Learning Calculated?

A dog and a tabby cat resting side by side in the sun, two classes an image classifier must tell apart in its classification report
Photo by MEHMET KAYNAR on Pexels

A classification report in deep learning is calculated from one table: the confusion matrix, which counts how often each true class was predicted as each class. For every class, precision is the share of its predictions that were right and recall is the share of its real examples the model found. F1 is the harmonic mean of the two, and the bottom rows average the per-class scores, either equally or weighted by class size.

The network never enters the calculation. The report sees two lists of labels, the true answers and the predictions, so the arithmetic is identical for a logistic regression and a transformer. Below, every number is worked by hand, then reproduced in code.

What a classification report shows

The report is a small table with one row per class and four columns. This is the one the article builds, from a three-class image classifier scored on 20 test images:

              precision    recall  f1-score   support

         cat      0.727     0.800     0.762        10
         dog      0.800     0.667     0.727         6
         fox      0.750     0.750     0.750         4

    accuracy                          0.750        20
   macro avg      0.759     0.739     0.746        20
weighted avg      0.754     0.750     0.749        20

The columns answer these questions:

ColumnWhat it answers
precisionOf everything the model labelled as this class, how much really was?
recallOf everything that really was this class, how much did the model find?
f1-scoreOne score that balances the two
supportHow many real examples of this class the test set contains

Support is the column people skip, and it tells you how far to trust the other three. A recall of 0.750 on four foxes means the model found three of them. Move one fox and the figure swings by 25 points.

The same report grades a fine-tuned language model used as a classifier, and LLM Systems Engineering covers fine-tuning a model for that kind of task.

It is also the most screenshotted artefact in machine learning. It gets pasted into a team channel, someone reacts with a rocket emoji, and nobody scrolls down to the averages.

The confusion matrix behind every classification report

A confusion matrix is a grid where rows are the true classes and columns are the predicted ones. Each cell counts the test examples with that combination. The diagonal holds the correct answers, and everything off it is a mistake, filed by type.

True \ PredictedcatdogfoxRow total
cat81110
dog2406
fox1034
Column total115420

Each class then gets three counts. Treat cat as the positive class and every other animal as negative:

  • True positives (TP): cats predicted as cat. The diagonal cell, 8.
  • False positives (FP): other animals predicted as cat. The rest of the cat column, 2 + 1 = 3.
  • False negatives (FN): cats predicted as something else. The rest of the cat row, 1 + 1 = 2.

This one-versus-rest step turns a three-class problem into three binary ones, which is why every row is calculated independently.

How precision and recall are calculated per class

The scikit-learn documentation defines precision as TP / (TP + FP) and recall as TP / (TP + FN). For the cat class:

  • Precision = 8 / (8 + 3) = 8 / 11 = 0.727
  • Recall = 8 / (8 + 2) = 8 / 10 = 0.800

The other two classes follow the same arithmetic:

ClassTPFPFNPrecisionRecall
cat8328 / 11 = 0.7278 / 10 = 0.800
dog4124 / 5 = 0.8004 / 6 = 0.667
fox3113 / 4 = 0.7503 / 4 = 0.750

Watch the denominators. Recall divides by the row total, which is the support. Precision divides by the column total, which is how often the model chose that label. A model that says "dog" rarely and cautiously earns high dog precision and low dog recall, and that is exactly the pattern here.

Which one matters depends on what each mistake costs. A spam filter favours precision, since a real email lost to the junk folder hurts more. A disease screen favours recall, because the missed case is the expensive error.

Close-up of a line chart on a laptop screen, the kind of metric trend a classification report feeds
Photo by Markus Winkler on Pexels

How the F1 score combines precision and recall

The F1 score is one number that stays high only when precision and recall are both high.

It is their harmonic mean: F1 = 2 × P × R / (P + R). For the cat class that is 2 × 0.727 × 0.800 / (0.727 + 0.800) = 0.762. Dog comes out at 0.727 and fox at 0.750.

The harmonic mean is used because it punishes imbalance. A class with precision 1.0 and recall 0.1 has an ordinary average of 0.55, which sounds passable. Its F1 is 0.18, which is what that model deserves. The harmonic mean is the friend who refuses to let you round your grade up because one exam went well.

Accuracy, macro average and weighted average

The last three rows summarise the whole model. They differ only in how much say the small classes get.

  • Accuracy: correct predictions over all predictions. The diagonal sums to 8 + 4 + 3 = 15, and 15 / 20 = 0.750. It is a single number, printed in the f1-score column.
  • Macro average: the plain mean of the per-class scores. Macro precision is (0.727 + 0.800 + 0.750) / 3 = 0.759. Every class counts equally, whatever its size.
  • Weighted average: the mean weighted by support. Weighted precision is (0.727 × 10 + 0.800 × 6 + 0.750 × 4) / 20 = 0.754. Large classes dominate.

Look at weighted recall: 0.750, identical to accuracy. That is an identity. Weighting each class's recall by its support multiplies TP / support by support, which leaves total true positives over total examples, the definition of accuracy.

Macro average is the honest number when classes are imbalanced. Google's Machine Learning Crash Course gives the classic case: when one class appears 1% of the time, a model that always predicts the other class scores 99% accuracy while being useless. Its weighted average looks nearly as good. Its macro average collapses, because the rare class scores zero and still gets a full vote.

That model is the most confident employee you will ever manage. It is never wrong about the easy case and has never once done the job it was hired for.

Getting a classification report from a deep learning model

A neural network outputs scores and the report needs labels. Converting one into the other is the only step specific to deep learning.

A multi-class network ends in logits, meaning one raw score per class. Taking the argmax picks the highest-scoring class as the prediction. In PyTorch:

import torch
from sklearn.metrics import classification_report

model.eval()
y_true, y_pred = [], []
with torch.no_grad():
    for x, y in test_loader:
        logits = model(x)
        y_pred.extend(logits.argmax(dim=1).tolist())
        y_true.extend(y.tolist())

print(classification_report(y_true, y_pred,
      target_names=["cat", "dog", "fox"], digits=3))

In Keras the conversion is one line, model.predict(x_test).argmax(axis=1). If your labels are one-hot encoded, apply argmax to them too, or scikit-learn raises an error about mixing target types.

Binary classifiers with a sigmoid output need a threshold instead. Raise it above the usual 0.5 and precision tends to rise while recall falls, so record the threshold next to the numbers.

One warning deserves attention. If the model never predicts a class, its precision divides by zero. The classification_report function handles this through a zero_division setting whose default reports 0 and raises a warning. That warning usually means a class has collapsed, so investigate before silencing it.

Whether a custom classifier is worth training at all is a separate question, covered in when to build an LLM and when to skip it. Evaluation is also the part of a system teams most often under-resource, which is the argument of the AI systems engineering problem nobody budgets for.

Before trusting any classification report, run through this list:

  1. Read support first. Scores on a handful of examples are anecdotes.
  2. Compare the macro and weighted averages. A wide gap means the small classes are struggling.
  3. Find the lowest recall in the table. That class is where the model fails quietly.
  4. Confirm the report came from a held-out test set the model never trained on.
  5. Note the decision threshold for any binary model.

Then open the confusion matrix, which shows what each class is being confused with. AI Engineering places this evaluation inside a full production pipeline.

Our cat classifier, for the record, called two dogs cats. Anyone who has lived with both will call that forgivable.

Frequently asked questions

What does support mean in a classification report?

Support is the number of true examples of each class in the test set, counted from the ground-truth labels. It does not depend on the model's predictions. It is the denominator of recall and the weight used in the weighted average, so a low support means the scores on that row rest on very few examples.

What is a good F1 score?

It depends on the task and the baseline. Compare against a trivial model that always predicts the most common class, and against the previous version of your own model. An F1 of 0.80 can be excellent on a hard, noisy task and poor on a clean benchmark.

What is the difference between micro and macro average?

Micro average pools the raw counts across all classes before dividing, so every example counts equally. Macro average calculates the score per class first and then takes a plain mean, so every class counts equally. In a single-label multi-class problem, every micro-averaged score equals accuracy.

Why does the classification report show an UndefinedMetricWarning?

It appears when a score would divide by zero, most often because the model never predicted some class, which leaves its precision undefined. Scikit-learn reports 0 for that cell by default and warns you. The zero_division setting silences it, though it is worth checking first whether a class has collapsed.

Can you calculate a classification report in PyTorch or Keras directly?

Yes. The arithmetic is short enough to write by hand from a confusion matrix, and libraries such as torchmetrics track precision and recall during training. Most teams still call scikit-learn's classification_report at the end, because its format is the one everyone recognises.

Sources