About This Case Study
In 2018, the computer scientist Joy Buolamwini, then at the MIT Media Lab, published an audit — Gender Shades, with Timnit Gebru — that tested commercial facial-analysis systems and found they failed dramatically more often on darker-skinned women than on lighter-skinned men. The finding was not an impression but a measurement: a benchmark dataset built to be balanced across skin type and gender, run against products from major vendors, with the error rates reported group by group. In the years since, the same kind of technology Buolamwini audited has been deployed by police departments to identify criminal suspects from images — and it has produced false matches that led to the arrests of innocent people. This case study works in two parts and with two kinds of source: the Gender Shades project site, which presents the audit in teachable detail, and the case materials of the American Civil Liberties Union, which document the wrongful arrests that followed when systems like the ones Buolamwini tested were turned on the public.
Buolamwini gives this case its central concept: the "coded gaze." Just as a human gaze can carry the assumptions and blind spots of the person looking, an automated system carries the assumptions of the people who build it — including who counts as a normal or default user, and therefore whose faces the system is trained and tested to recognize. When the people and data behind a system underrepresent darker-skinned faces, the coded gaze "sees" those faces poorly, and the failure is built in rather than accidental. The case study asks you to do two distinct kinds of analysis with this idea. First, to evaluate Gender Shades as a method: to ask what makes an audit — a replicable test, a balanced benchmark, published numbers — a more powerful form of evidence than testimony or anecdote, and what that says about what institutions will accept as proof of discrimination. Second, to follow the coded gaze out of the lab and into the street: to trace the chain from a system's uneven error rates, through a police department's decision to deploy it anyway, to a specific person in handcuffs — and to ask, at each link, where the harm was decided.
Before You Begin
Have ready:
- Joy Buolamwini and Timnit Gebru, Gender Shades — the project site presenting the audit's dataset, error-rate findings, and the vendors' responses (the anchor for Part One).
- American Civil Liberties Union, "After Third Wrongful Arrest, ACLU Slams Detroit Police Department for Continuing to Use Faulty Facial Recognition Technology" — on the wrongful arrests produced by police use of the technology.
- American Civil Liberties Union, "Civil Rights Advocates Achieve the Nation's Strongest Police Department Policy on Facial Recognition Technology" — on the settlement and the policy changes that followed.
- A note-taking surface — paper, a document, or a shared doc if you are working in a group.
A note on currency: police use of facial recognition, the litigation over it, and the resulting policies are all changing; the settlement described here produced new rules that other jurisdictions may or may not adopt. Use the dates each source gives, and check for developments since.
Before you begin, write down, in a sentence or two each:
- If you wanted to prove that a technology discriminated, what kind of evidence would convince a skeptic — a company, a court, a police chief? A story about one person? A number? Why?
- A facial-recognition match is usually one step in a longer process, not an arrest by itself. Where, in the chain from a software result to someone in handcuffs, could the error have been caught — and who would have had to catch it?
The Exercise
Phase 1: Orientation (5–7 minutes)
Get oriented to both halves of the case. Skim the Gender Shades site to see what the audit tested and what it found, and skim the two ACLU sources to see what happened when facial recognition was used by police. Without analyzing yet, write a few sentences connecting the two: what did the audit show about how these systems perform, and what did deployment of similar systems produce? The point is to hold the lab finding and the street consequence in view together before examining either closely.
For groups: divide into those who will lead on Part One (the audit) and those who will lead on Part Two (the arrests), but make sure everyone skims both.
Phase 2: Part One — Evaluating the Audit (12–15 minutes)
Work through the Gender Shades site and evaluate the audit as a piece of evidence. Be specific:
- The benchmark. Buolamwini and Gebru built their own dataset rather than using existing ones. What was wrong with the datasets these systems were usually tested on, and how was the new benchmark constructed to be balanced across skin type and gender? Why does the choice of test set determine what an evaluation can reveal?
- The findings. What did the audit measure, and what did it find — the error rates for different groups, and the gap between the best- and worst-performing categories? Pull the actual numbers. Why is reporting results group by group, rather than as a single overall accuracy figure, essential to what the audit shows?
- The response. The site documents how the companies reacted. What did vendors do after the audit, and what does it mean that a published, replicable test could move a company in a way that complaints had not? What made this evidence hard to dismiss?
For groups: Part One leads work through these while Part Two leads observe and take notes for the comparison in Phase 4.
Phase 3: Part Two — From Error Rate to Arrest (12–15 minutes)
Now follow the technology into deployment, using the ACLU materials. The audit measured how often these systems err; this part asks what happens when they are used anyway. Work through:
- The deployment decision. Police departments adopted facial recognition to generate suspects from images. Given what the audit showed about error rates, especially for darker-skinned faces, what does it mean that departments deployed it on a general public that includes the very groups the systems identify worst? Who made that choice, and on what reasoning?
- The arrests. The ACLU materials document specific people — among them Robert Williams, arrested in Detroit in 2020, and Porcha Woodruff, arrested in 2023 — wrongly identified by facial recognition and arrested as a result. Working from the sources, reconstruct what happened: how a false match became an arrest, and what the consequences were for the people involved. Note who they are and what the technology's failure cost them.
- The remedy. The second ACLU source describes a settlement and the policy changes that came from it. What did the litigation produce — what new limits on police use of the technology — and what does it tell you that the remedy came through a lawsuit rather than through the agencies choosing to stop? What does the policy fix reach, and what does it leave in place?
For groups: Part Two leads work through these while Part One leads observe and take notes.
Phase 4: Applying the Frame — The Coded Gaze and the Burden of Proof (10–12 minutes)
Bring the two parts together through Buolamwini's idea of the coded gaze and the question of evidence. Ask:
- Trace the coded gaze across the whole chain. Where does it begin — in the data, the design, the testing — and how does a built-in failure to "see" darker-skinned faces travel from a benchmark error rate to a wrongful arrest? At which link does a technical disparity become a person's loss of liberty?
- What does this case say about the burden of proof? Buolamwini's audit succeeded partly because it produced numbers a vendor could not deny. Why might rigorous, quantified evidence be required before discrimination is believed — discrimination that the people experiencing it could already describe? Who is made to carry the burden of proving harm, and what does that cost?
- Where does the frame fit, and where does it strain? Consider that some harms of these systems are not about accuracy at all — that a facial-recognition system which worked perfectly on every face could still enable mass surveillance, and that closing the error gap would not by itself address the decision to deploy. Does the coded gaze, focused on who a system sees poorly, capture the whole problem, or only part of it? Name what it leaves out.
For groups: spend the first half on the first two questions together; spend the second on where the frame strains.
Closing Reflection
In two or three sentences, complete this thought:
The Gender Shades audit was hard to dismiss because __________. Following the technology into deployment, the chain that runs from a benchmark error rate to a wrongful arrest is __________. The hardest thing to settle, after this exercise, is __________.
Write the most specific version you can; the value is in naming exactly what made the evidence undeniable and exactly where, in the chain, the harm was decided.
A Note on Modes
Solo mode. Work through the four phases in order, keeping brief notes. The audit evaluation from Phase 2 and the deployment chain from Phase 3 are the central artifacts; keep them for use with the unit's questions on Buolamwini, audits, and facial recognition.
Group mode (3–6 people). Designate a timekeeper. Split into Part One and Part Two leads for Phases 2 and 3, with each group taking notes while the other works; come together for Phase 4. If time is short, the closing reflection can be individual writing after the session.
After the Case Study
- The peer-reviewed Gender Shades paper (Buolamwini and Gebru, Proceedings of Machine Learning Research, 2018) presents the method in full, for learners who want to see exactly how the benchmark was built and the systems tested.
- For the first-person account of the audit and what followed — the congressional testimony, the corporate pushback, the founding of the Algorithmic Justice League — see Joy Buolamwini, Unmasking AI: My Mission to Protect What Is Human in a World of Machines (Random House, 2023), in the unit's Key Scholarship.
- The accuracy problem this case centers is not the only problem with facial recognition; the unit's Further Readings on surveillance and carceral technology raise the further question of what a perfectly accurate system would still do, worth holding alongside the audit.