Skip to main content

Case Study: COMPAS and the Fairness Wars — Auditing the Algorithm Yourself: Case Study: COMPAS and the Fairness Wars — Auditing the Algorithm Yourself

Case Study: COMPAS and the Fairness Wars — Auditing the Algorithm Yourself
Case Study: COMPAS and the Fairness Wars — Auditing the Algorithm Yourself
  • Show the following:

    Annotations
    Resources
  • Adjust appearance:

    Font
    Font style
    Color Scheme
    Light
    Dark
    Annotation contrast
    Low
    High
    Margins
  • Search within:
    • My Notes + Comments
    • Notifications
    • Privacy
  • Project HomeThe Making of Race
  • Projects
  • Learn more about Manifold

Notes

table of contents
  1. About This Case Study
  2. Before You Begin
  3. The Exercise
    1. Phase 1: Orientation — The Dispute (5–7 minutes)
    2. Phase 2: ProPublica's Case — Equal Error Rates (10–12 minutes)
    3. Phase 3: The Company's Case — Equal Predictive Value (10–12 minutes)
    4. Phase 4: The Impossibility and the Choice (10–12 minutes)
    5. Phase 5: From Risk Scores to Hiring (optional, 5–7 minutes)
  4. Closing Reflection
  5. A Note on Modes
  6. After the Case Study

About This Case Study

COMPAS is a proprietary software tool that scores criminal defendants by their predicted risk of reoffending; courts have used those scores to inform decisions about bail, sentencing, and parole. In 2016, the newsroom ProPublica published an investigation, "Machine Bias," arguing that COMPAS was racially biased: among defendants who did not go on to reoffend, Black defendants were far more likely than white defendants to have been labeled high-risk. The company that makes COMPAS, then Northpointe and later renamed equivant, rejected the charge — and, crucially, did so without disputing ProPublica's numbers. The company argued that by the standard that mattered, the tool was fair: a given risk score meant the same thing — the same actual rate of reoffending — whether the defendant was Black or white. Both sides were looking at the same data, and both were, by their own definition, correct. This case study puts you inside that dispute, working from the published tables themselves.

The COMPAS dispute became the most-cited example of a hard truth about algorithmic fairness: there is no single thing that "fair" means, and some reasonable definitions cannot all be satisfied at once. ProPublica's standard was equal error rates — the tool should wrongly flag innocent people at the same rate across racial groups. The company's standard was equal predictive value, also called calibration — a given score should carry the same real meaning across groups. Statisticians later proved that when two groups have different underlying base rates of the outcome being predicted, a tool generally cannot satisfy both standards at once; improving one worsens the other. The choice between them is therefore not a technical question with a correct answer but a value judgment about which kind of unfairness is worse — wrongly burdening innocent people, or making scores mean different things for different groups. The analytical move in this case study is to stop treating "remove the bias" as an engineering instruction and start treating it as a political and moral choice: to work the competing definitions on the actual numbers, see for yourself why both cannot hold, and decide which a just system should require.

Before You Begin

Have ready:

  • ProPublica, "Machine Bias" (2016) — the original investigation and its findings (the anchor).
  • Marcello Di Bello, "Algorithmic Fairness — ProPublica v. Northpointe" — a teaching handout that lays the competing tables and definitions out side by side; the most efficient way to work the numbers.
  • ProPublica, "Technical Response to Northpointe" (2016) — ProPublica's reply defending its analysis.
  • equivant Supervision, "Debunking Misconceptions About the COMPAS Core Instrument" (2024) — the company's account of why it considers the instrument fair.
  • A note-taking surface, and a calculator or spreadsheet if you want to check the rates yourself — paper, a document, or a shared doc if you are working in a group.

A note on the sources: two of these are written by the parties to the dispute — ProPublica and the company — and each makes its own case; read them as arguments, not neutral reports. The handout is a third-party teaching aid that presents both sides' figures. Note the dates, and read the company's and the newsroom's claims against each other rather than taking either at its word.

Before you begin, write down, in a sentence or two each:

  • What would it mean for a risk score to be "fair"? Try to write a one-sentence definition now, before you see the dispute — you will test it against the data.
  • If a tool is accurate overall but its mistakes fall more heavily on one group, is it biased? Does your answer depend on what kind of mistake it is?

The Exercise

Phase 1: Orientation — The Dispute (5–7 minutes)

Read ProPublica's "Machine Bias" and skim the handout to get the shape of the dispute. In your own words, state the disagreement: what did ProPublica claim about COMPAS, and on what evidence? What did the company say in response, and — this is the key point — did it dispute ProPublica's numbers or interpret them differently? Write a few sentences capturing the fact that both sides worked from the same data.

For groups: read individually, then agree on a one-paragraph statement of what, exactly, the two sides disagree about.

Phase 2: ProPublica's Case — Equal Error Rates (10–12 minutes)

Work through ProPublica's argument on the numbers, using the handout's tables. Focus on the error rates among defendants whose actual outcomes are known:

  • Among defendants who did not reoffend, what share were wrongly labeled high-risk — the false-positive rate — and how did that share differ between Black and white defendants? Pull the actual figures.
  • Among defendants who did reoffend, what share were labeled low-risk — the false-negative rate — and how did that differ by group?
  • State ProPublica's definition of fairness in your own words: what would have to be equal across groups for the tool to count as fair by this standard? Why does ProPublica treat the gap it found as evidence of racial bias?

For groups: one or two members present these figures to the others before moving on.

Phase 3: The Company's Case — Equal Predictive Value (10–12 minutes)

Now work through the company's response, using its materials and the handout. The company did not claim ProPublica's error-rate numbers were wrong; it argued they were the wrong test. Focus on a different question:

  • Take defendants who received a given risk score — say, a high score. What share of them actually went on to reoffend, and was that share roughly the same for Black and white defendants? This is calibration, or equal predictive value.
  • State the company's definition of fairness in your own words: what would have to be equal across groups for the tool to count as fair by this standard? Why does the company consider a score that means the same thing across groups to be fair, regardless of the error-rate gap?
  • Whose definition matches your own one-sentence definition from before you began? Has reading the second argument changed which one seems right?

For groups: a different one or two members present the company's figures and definition.

Phase 4: The Impossibility and the Choice (10–12 minutes)

Here is the heart of the case. The two standards — equal error rates and equal predictive value — sound like they should go together, but they generally cannot both hold when the two groups have different base rates of the outcome (here, different actual rates of reoffending, which themselves reflect a racially unequal system of arrest and prosecution). Work through why, and then decide:

  • In your own words, explain the bind: why does forcing the error rates to be equal across groups pull the predictive values apart, and the reverse, when base rates differ? You do not need the formal proof — explain the intuition from the tables in front of you.
  • This means a designer must choose which standard to meet. Frame the choice as a question of values: which is the worse injustice — a tool that wrongly burdens innocent members of one group more often, or a tool whose scores carry different real meaning for different groups? There is no neutral answer; make the case for one.
  • Now put yourself in the position of a court asked whether COMPAS may be used. Which definition of fairness should the law require, and why? What would you need to know — about the base rates, about how scores are actually used in decisions — to decide responsibly?

For groups: hold a short structured debate — assign some members to argue for each standard — then see whether the group can reach a position, and notice what the disagreement turns on.

Phase 5: From Risk Scores to Hiring (optional, 5–7 minutes)

The COMPAS dispute is about criminal risk scores, but the same questions — what counts as fair, who gets to see how a system decides, who must prove harm — arise wherever automated tools make consequential decisions. New York City's Local Law 144, the first law of its kind in the United States, requires employers using automated hiring tools to have them independently audited for bias and to disclose the results. Read the city's overview of the rule and consider: does mandatory auditing and disclosure address the problem this case raises, or only part of it? An audit can reveal a disparity, as a facial-recognition audit does — but, as the COMPAS dispute shows, it cannot by itself say which definition of fairness a tool should meet. What does a transparency requirement accomplish, and what does it leave to politics?

Closing Reflection

In two or three sentences, complete this thought:

ProPublica and the company disagreed about COMPAS even while sharing the same data because __________. A tool generally cannot satisfy both definitions of fairness at once because __________. The definition I would require, and the hardest part of choosing it, is __________.

Write the most precise version you can; the value is in stating the bind exactly and owning the value judgment the math forces.

A Note on Modes

Solo mode. Work through the phases in order, keeping your figures in brief notes. The worked error-rate and predictive-value tables from Phases 2 and 3 and your explanation of the bind in Phase 4 are the central artifacts; keep them for use with the unit's questions on fairness and risk scoring.

Group mode (3–6 people). Designate a timekeeper. Members present the two sides' figures in Phases 2 and 3; Phase 4 works as a short structured debate before a synthesis. If time is short, the closing reflection can be individual writing after the session.

After the Case Study

  • For a scholarly synthesis of the dispute, and a broader argument that opaque, proprietary risk models should be replaced with transparent ones, see Cynthia Rudin, Caroline Wang, and Beau Coker, "The Age of Secrecy and Unfairness in Recidivism Prediction," Harvard Data Science Review 2, no. 1 (2020).
  • The same competing-definitions problem recurs across automated decision-making — in lending, hiring, and benefits as well as criminal justice; the unit's readings on these systems let you ask whether the COMPAS bind is a special case or the general rule.
  • The dispute also illustrates the unit's larger argument about the appearance of objectivity: a tool can be "accurate" and still encode a contested choice about what fairness means, which connects directly to the New Jim Code's claim that automated systems can disguise value judgments as neutral computation.

Annotate

MODULE 3 CASE STUDIES
CC BY-NC-SA 4.0
Powered by Manifold Scholarship. Learn more at
Opens in new tab or windowmanifoldapp.org