← Back to explainers
Explainer

What Is a Confusion Matrix?

The foundational building block behind most fairness metrics.

Learn how a confusion matrix breaks down accuracy into true positives, true negatives, false positives, and false negatives, and why derived metrics like FPR and FNR are essential for detecting bias.

A model can boast 90% accuracy while hiding the fact that its 10% error rate consists entirely of falsely flagging innocent people from a minority group. A confusion matrix splits a single accuracy number into the four outcomes that actually matter, revealing exactly who pays for a model's mistakes.

The One-Sentence Definition

A confusion matrix is a table that summarizes a classification model's performance by breaking its predictions down into true positives, true negatives, false positives, and false negatives - the four foundational numbers that drive almost every algorithmic fairness metric.

Why It Matters

In high-stakes systems, the type of mistake a model makes is just as important as how often it makes it. A false positive (a false alarm) and a false negative (a missed case) rarely carry the same cost.

Accuracy alone is a dangerous metric for fairness because it aggregates all correct predictions and all errors into a single percentage. A model that perfectly predicts outcomes for a 90% majority group but misclassifies everyone in a 10% minority group will still score 90% overall accuracy. By calculating a confusion matrix separately for each demographic group, you stop asking "is the model accurate?" and start asking "how does the model fail, and who absorbs the cost of those failures?"

The Core Concept

A binary classification model makes a prediction (Positive or Negative), and the real world provides the actual truth (Positive or Negative). This creates a 2x2 grid of possibilities:

Predicted Positive (1)Predicted Negative (0)
Actual Positive (1)True Positive (TP): Correctly flaggedFalse Negative (FN): Incorrectly missed
Actual Negative (0)False Positive (FP): Incorrectly flaggedTrue Negative (TN): Correctly cleared

From these four quadrants, we derive the rates that define most fairness metrics:

Concrete Example: COMPAS - Audit 01

The COMPAS recidivism algorithm perfectly illustrates why confusion matrices matter for fairness. Northpointe, the creator of COMPAS, pointed out that their model had equal Positive Predictive Value for Black and White defendants - a risk score of 7 meant roughly the same likelihood of reoffending for both groups.

However, ProPublica looked at the confusion matrix and calculated the False Positive Rate and False Negative Rate. They found that Black defendants who did not reoffend were falsely labeled as high-risk (False Positives) at nearly twice the rate of White defendants who did not reoffend. Conversely, White defendants who did reoffend were incorrectly labeled as low-risk (False Negatives) much more often than Black defendants.

Because the underlying base rates of arrest were different for the two groups, it was mathematically impossible for the model to have both equal Positive Predictive Value and equal False Positive/Negative Rates. The confusion matrix made this trade-off visible, sparking the modern debate on algorithmic fairness.

Detection Code

This code computes the four core values (TP, TN, FP, FN) and the key derived rates for each demographic group, allowing you to spot disparities in how errors are distributed.

import pandas as pd
from sklearn.metrics import confusion_matrix

def group_confusion_metrics(df, y_true_col, y_pred_col, group_col):
    """
    Computes a confusion matrix and derived rates for each group.
    
    Parameters:
        df: DataFrame with true labels, model predictions, and group membership
        y_true_col: column name of the ground-truth binary label (1 = positive)
        y_pred_col: column name of the model's binary prediction (1 = positive)
        group_col: column name of the protected attribute or group label
        
    Returns a DataFrame indexed by group with TP, TN, FP, FN, TPR, FPR, FNR, and PPV.
    """
    rows = []
    for group, sub in df.groupby(group_col):
        # Ensure labels are binary [0, 1] for reliable unraveling
        tn, fp, fn, tp = confusion_matrix(
            sub[y_true_col], sub[y_pred_col], labels=[0, 1]
        ).ravel()
        
        # Calculate rates with zero-division protection
        tpr = tp / (tp + fn) if (tp + fn) > 0 else float("nan")
        fpr = fp / (fp + tn) if (fp + tn) > 0 else float("nan")
        fnr = fn / (fn + tp) if (fn + tp) > 0 else float("nan")
        ppv = tp / (tp + fp) if (tp + fp) > 0 else float("nan")
        
        rows.append({
            "group": group,
            "TP": tp, "TN": tn, "FP": fp, "FN": fn,
            "TPR (Recall)": tpr, 
            "FPR": fpr, 
            "FNR": fnr, 
            "PPV (Precision)": ppv,
            "n": len(sub)
        })
        
    return pd.DataFrame(rows).set_index("group")

# Usage example
# metrics_df = group_confusion_metrics(compas_df, "two_year_recid", "high_risk_prediction", "race")
# print(metrics_df[["FPR", "FNR", "PPV (Precision)"]])

Limitations

1. Requires a single decision threshold

A confusion matrix evaluates a model at one specific threshold (e.g., probability > 0.5 means positive). If you change the threshold, all four numbers in the matrix change. To evaluate how a model performs across all possible thresholds, you need a ROC Curve.

2. Counts all errors equally

The standard confusion matrix treats every False Positive as identical to every other False Positive. It does not know if a model barely missed the threshold (0.49) or was wildly wrong (0.01). It also does not inherently weigh the human cost of a False Positive against a False Negative.

3. The Base Rate Problem

When two demographic groups have different base rates (prevalence) of the target variable in the dataset, it is mathematically impossible to equalize all metrics derived from the confusion matrix at the same time. You often have to choose between equalizing predictive value (calibration) and equalizing error rates (Equalized Odds).

Further Reading

Part of The Fair Code Project - exposing and fixing algorithmic bias with real data and open code.