Skip to content
Published on

How the Stanford Prison Experiment Fell Apart — What the Archives Revealed

Share
Authors

Introduction — the Six Days That Made It into the Textbooks

The story is usually summarized like this. In the summer of 1971, a fake prison was built in the basement of the Stanford psychology building. Ordinary male college students were assigned to be guards or prisoners on a coin toss. Within days the guards had turned brutal and the prisoners had broken down. An experiment planned for two weeks was halted after six days. The conclusion always ends with the same sentence: there is no separate class of evil people, but evil situations make ordinary people act that way.

This story traveled into introductory psychology textbooks, documentaries, two feature films, and the courtroom of the Abu Ghraib prison case. And since 2018, it has become the most famous collapse in psychology.

This installment of Psychology, Straight from the Papers is a little different. In earlier installments, what collapsed was the size of an effect. Here, what collapses is the process by which the data were made.

Where the Original Data Was Published — 24 People, and a Small Journal

There is a fact to establish first. This experiment has no flagship paper of the kind we usually imagine. The most widely cited source is a 1973 report by Craig Haney, Curtis Banks and Philip Zimbardo, published in a small journal called the International Journal of Criminology and Penology. Which is to say it never passed the rigorous review of a top-tier psychology journal.

The design goes like this.

  • An ad ran in the local newspaper: "Male college students wanted for a psychological study of prison life. 15 dollars a day, 1 to 2 weeks."
  • Out of 75 respondents, 24 were selected through interviews and testing. These 24 were randomly assigned to guard and prisoner roles.
  • Zimbardo himself took the role of prison superintendent, and the graduate student David Jaffe played the warden.
  • There was no control group. With nothing to compare against, it was closer to a demonstration of what happened inside a single group.

A sample of 24, no control group, and the experimenter appearing as a character in his own study. Those three lines alone already explain why this research cannot carry the weight of its conclusion. But the real problem surfaces after that.

The Assignment Was Random, but the Applicants Were Not

In 2007, Thomas Carnahan and Sam McFarland designed a very simple test. They placed two versions of a newspaper ad. One was the original Zimbardo wording, "a psychological study of prison life," and the other simply dropped the word "prison" to read "a psychological study." Then they gave personality tests to the people who applied to each ad.

The results all pointed the same way. People who answered the prison ad scored significantly higher on aggressiveness, authoritarianism, Machiavellianism, narcissism and social dominance orientation, and lower on empathy and altruism.

The meaning of this finding has to be stated precisely. There is nothing wrong with the random assignment Zimbardo performed. The problem is that before that lottery was ever drawn, the applicant pool had already been filtered. The moment you generalize the results to "anyone who is an ordinary person," this bias eats away at the entire conclusion. The problem was not the sample but the path by which the sample was produced.

2018, the Archive Opens — Coaching and Acting

The decisive blow came from the French researcher Thibault Le Texier. He went through the original materials Stanford University had been holding — audio tapes, meeting notes, participant correspondence, unreleased film — and published his findings in a French-language book in 2018 and in an American Psychologist paper in 2019.

There are three key points.

First, the guards did not turn brutal on their own. At an orientation before the experiment began, the research team specifically instructed them to make the prisoners feel powerless and bored, subject to arbitrary control, and stripped of privacy. The tapes preserve a scene in which Jaffe corners one passive guard and coaxes him to be harder, to act "tougher." This is not demand characteristics leaking in; there were instructions.

Second, there is the testimony about the most famous scene in the experiment, the mental breakdown of prisoner number 8612. The man himself, Douglas Korpi, has said in several later interviews that he was acting. He had applied thinking he could study for graduate school exams inside the prison, then wanted out once his books were blocked, and having no convenient way to leave, he performed a breakdown. This testimony became widely known through a long 2018 article by the journalist Ben Blum.

Third, the behavior of the guards was not uniform. Only some of them acted brutally; a considerable number were lukewarm or friendly toward the prisoners. The summary that "the guards changed" erases this variance.

Same Design, Different Result — the 2006 BBC Prison Study

So what happens if you run it again? In 2001, Stephen Reicher and Alex Haslam reconstructed the prison situation together with the BBC, but gave the guards no behavioral coaching whatsoever, worked under the supervision of an independent ethics committee, and kept the researchers from stepping in as characters. The results appeared in the British Journal of Social Psychology in 2006.

ItemStanford (1971)BBC prison study (2006)
Participants2415
Duration6 days (2 weeks planned)8 days
Guard instructionsAdvance orders to induce powerlessnessNo behavioral instructions
Researcher roleZimbardo doubling as superintendentOutside observation, ethics committee oversight
ResultSome guards turned harshGuards failed to cohere as a group, prisoners cohered and the regime collapsed
Where publishedSmall criminology journalMajor social psychology journal

The guards in the BBC study were uncomfortable with their role and never built norms they agreed on among themselves. The prisoners, by contrast, formed a shared identity and challenged authority, and within days the hierarchy fell apart. Interestingly, the egalitarian self-governing regime set up afterward failed too, and after that failure some participants began proposing a more coercive order themselves.

The conclusion Reicher and Haslam drew points in a different direction from that of Zimbardo. People are not automatically absorbed into the roles they are handed. Brutality appears when there is a group you identify with and the cause of that group justifies brutality. And the powerlessness that follows the collapse of an alternative order is what invites authoritarianism. The variable is not role conformity but group identification.

The Power of the Situation Survives — Only the Address of the Evidence Moves

This is where a balance has to be struck. The collapse of the Stanford experiment does not turn into "situations do not change behavior." If anything, the opposite is closer to the truth.

There is still plenty of evidence supporting the power of the situation. The address of that evidence is simply different. What showed obedience rates moving from 0 percent up into the 90s as conditions were changed systematically was the set of 24 variations run by Milgram (covered in detail in the Milgram reexamined installment), and what showed how group identity and cause reshape behavior is the research program of Reicher and Haslam themselves. That rules, anonymity, monitoring and reward structures change behavior has been confirmed over and over in organizational research and field experiments.

What changes is the shape of the sentence. From "good people put into a bad situation automatically become monsters" to "people can turn brutal when they are persuaded by the cause of a group they feel they belong to." The second is less dramatic, but it is more useful because it tells you where you can actually intervene. There is nothing to grab hold of in something that happens automatically, whereas persuasion can be argued against.

What to Discard and What to Keep — How a Good Story Keeps Bad Evidence Alive

What to discard: the practice of citing this experiment as evidence about human nature. A sample of 24, no control group, the dual role of the experimenter, advance coaching of the guards, self-selection bias in the applicant pool, and testimony from the participant himself that the signature scene was acted. Any one of these alone would be reason to withhold judgment on the conclusion; here they are all present at once. Using this experiment as evidence in corporate training material or presentation slides is hard to recommend anymore.

What to keep 1 — random assignment does not make the applicant pool random. If you recruit participants for an internal experiment or a pilot program through an announcement, then no matter how fairly you assign them, the bias of "the people who responded to that announcement" remains untouched. The test by Carnahan and McFarland is a beautifully designed study that shows this with a one-line difference in ad copy.

What to keep 2 — a good story keeps bad evidence alive for a long time. This experiment lasted 47 years not because the data were sturdy but because the story was powerful. When Richard Griggs examined 13 introductory textbooks in 2014, essentially none of them engaged meaningfully with the criticisms. Textbooks introduced the experiment, documentaries cited the textbooks, and that fame in turn justified inclusion in the textbooks: a closed loop. Citation count is not verification.

What to keep 3 — the primary sources are still open. What Le Texier did was not a new experiment but the opening of an existing box. A great many of the recordings in the Stanford experiment archive had been available for decades, but nobody had listened to the end. The self-correction of psychology does not happen through large-scale re-verification alone. Sometimes it is enough for someone to read the original from the beginning again.

A Guide to Reading the Originals

  • Original report: Haney, C., Banks, W. C., & Zimbardo, P. G. (1973). Interpersonal dynamics in a simulated prison. International Journal of Criminology and Penology, 1, 69-97.
  • Self-selection test: Carnahan, T., & McFarland, S. (2007). Revisiting the Stanford prison experiment: Could participant self-selection have led to the cruelty? Personality and Social Psychology Bulletin, 33(5), 603-614.
  • Archival reconstruction: Le Texier, T. (2019). Debunking the Stanford prison experiment. American Psychologist, 74(7), 823-839.
  • Alternative study: Reicher, S. D., & Haslam, S. A. (2006). Rethinking the psychology of tyranny: The BBC prison study. British Journal of Social Psychology, 45(1), 1-40.
  • Textbook survey: Griggs, R. A. (2014). Coverage of the Stanford prison experiment in introductory psychology textbooks. Teaching of Psychology, 41(3), 195-203.

Reading tip: I recommend starting with Le Texier 2019. The best part of that paper is not the conclusion but the fragments of transcript quoted in the middle of the text. Read for yourself the passage where the warden talks a passive guard around, and no summary afterward will read the way it used to. Then move on to Reicher and Haslam 2006 and look at the time-series figure showing how the social identity measures of the participants moved day by day. The difference between the two prisons is contained in that single page.