- Published on
The 38 Were Never There — The Myth of the Bystander Effect and the Real Experiment
- Authors

- Name
- Youngju Kim
- @fjvbn20031
Introduction — The Number 38
In the early hours of March 13, 1964, Catherine Genovese, known to everyone as Kitty, was murdered in Kew Gardens, Queens, New York. Two weeks later an article in the New York Times reported it this way: thirty-eight neighbors watched from their windows for more than 30 minutes, and nobody called the police.
The article entered psychology textbooks, became an anecdote emblematic of big-city indifference, and served as the origin myth of the concept of the bystander effect. The number is still quoted in lectures and columns today.
The problem is that the number is not true. This installment of Psychology, Straight from the Papers is a case that has to be read in two layers. The famous anecdote did not survive fact-checking; the experiment the anecdote set in motion did survive; and the pessimistic conclusion the public drew from that experiment was in turn overturned by real data. We peel back the three layers in order.
2007 — Fact-Checking the Article
Rachel Manning, Mark Levine, and Alan Collins published a paper in American Psychologist in 2007 reexamining the record of the case. The subtitle summarizes where the authors stand: "The parable of the 38 witnesses."
Here is what they established by comparing court records with contemporaneous material.
The number in the article was less a product of reporting than of a conversation between a city editor and a police official. The headline of the article said 37 and the body said 38.
The incident was not a single scene unfolding for 30 minutes outside a window. There were two attacks in different places, and the second attack, the fatal one, happened in an interior hallway of the building and was not visible from apartment windows. The time was around 3:20 in the morning, and it was New York in March. Most of the neighbors who said they heard something heard one short scream while half asleep, and could not see the street from their beds.
Calls to the police were made. Testimony that someone called after the first attack survives in the court record. And a neighbor named Sophia Farrar went down to the hallway and held Genovese until the ambulance arrived. That is precisely the opposite of what the parable says.
The target Manning and colleagues aimed at was not only the newspaper. It was the fact that a discipline adopted an unverified newspaper article as the origin myth of one of its flagship concepts, and repeated it in textbooks for forty years because the story explained the concept so well. A good anecdote is not good evidence, and it tends to work in the direction of delaying verification.
For balance, this does not mean nothing was wrong that night. There certainly were people who heard something and did nothing. What is not true is the sentence "38 people watched."
And Yet the Experiment Was Real
Two young researchers who read the article, John Darley and Bibb Latane, set aside indifference as an explanation and formed a different hypothesis: the number of people might itself be the cause. And they turned it into an experiment.
The design of the 1968 paper still looks clever today. Students at New York University each sat alone in a separate booth. With the explanation that it was to preserve anonymity, they conversed only over an intercom, and only one person could speak at a time. The topic was personal difficulties in college life. During the conversation one participant mentions that he suffers from seizures, and a little later actually has one. He asks for help, saying he cannot breathe, and then the sound cuts off. That participant is in fact a recording.
Exactly one variable was manipulated: how many other people the participant believed were on the intercom.
| People on the intercom (including the victim) | n | Reported during the seizure | Mean response time |
|---|---|---|---|
| 2 | 13 | 85 percent | 52 seconds |
| 3 | 26 | 62 percent | 93 seconds |
| 6 | 13 | 31 percent | 166 seconds |
When they believed nobody else was there, 85 percent moved. When they believed four other people were present, it fell to 31 percent. The person who needed help, the disposition of the participant, and the physical conditions were all identical. The only thing that changed was the headcount in their heads.
The most striking passage in the paper is not the table but the observational notes. The participants who did not report were not calm. They trembled, sweated, and were visibly agitated, and afterward asked the experimenter urgently whether that person was all right. They were not indifferent; they were trapped. This distinction is the part of the whole study that is forgotten most often.
The sample across the three conditions totals 52 people. It is a small study by present standards, and it rests on the unusual structure of participants isolated in booths where they cannot see one another. Yet the result survived. The meta-analysis by Fischer and colleagues, published in Psychological Bulletin in 2011, pooled 105 studies and more than 7,700 participants and confirmed that the group size effect is real. It also identified an important condition. When the situation is unambiguously dangerous, when the other people are physically present, and when they are capable of helping, the effect shrinks and sometimes reverses direction.
Three Mechanisms
The bystander effect is not one phenomenon but a bundle of three different failures. Latane and Darley laid out the path to intervention as five steps: noticing, interpreting it as an emergency, taking it on as your own responsibility, knowing how, and acting. The three mechanisms jam at different steps.
Diffusion of responsibility is a failure at the third step. If six people could help, the guilt of not helping is divided six ways too. The sentence "somebody will do it" runs in six heads at once. The six-person condition in the 1968 experiment measured exactly this.
Pluralistic ignorance is a failure at the second step. The smoke experiment Latane and Darley published the same year isolates it on its own. Smoke flows through a vent into a room where a participant is filling out a questionnaire. Of participants who were alone, 75 percent reported it. But among participants sitting with two confederates instructed to show no reaction, only 10 percent reported it. Diffusion of responsibility does not apply here, because the smoke is a matter of your own safety. What broke down was the interpretation. Since the others are unruffled, it must be nothing. The trouble is that those others were reading it the same way.
Evaluation apprehension is a failure at the last step. What if I have judged this wrong, what if I step up clumsily and look ridiculous. The more spectators there are, the higher this cost climbs.
Different diagnoses call for different prescriptions. This distinction gets used in the last section.
Understanding It in Code — The Probability for One Person Is Not the Probability for the Group
The popular summary renders the experimental result like this: "In a crowd, nobody helps." But what the 1968 experiment measured is the probability that one individual helps. What matters to the victim is a different quantity: the probability that at least one person helps. The two do not move in the same direction.
# Each bystander's own willingness decays as the crowd grows
# (diffusion of responsibility), here as p1 * n ** -DECAY.
P1 = 0.60 # a lone bystander helps 60% of the time
DECAY = 0.5 # how fast individual responsibility dilutes
print(" n p(this person helps) p(at least one helps)")
for n in (1, 2, 3, 5, 10, 20):
p_individual = P1 * n**-DECAY
p_group = 1 - (1 - p_individual) ** n
print(f"{n:2d} {p_individual:.3f} {p_group:.3f}")
# output:
# n p(this person helps) p(at least one helps)
# 1 0.600 0.600
# 2 0.424 0.669
# 3 0.346 0.721
# 5 0.268 0.790
# 10 0.190 0.878
# 20 0.134 0.944
The left column is the bystander effect exactly as advertised. The probability of an individual intervening collapses from 0.60 to 0.13. The right column moves the other way. The probability that somebody helps rises from 0.60 to 0.94. Both columns come out of the same model, and both are true.
Of course this depends on the assumptions. Raise the decay exponent to 1.0 so that individual probability falls in inverse proportion to headcount, and the right column comes down with it, from 0.60 to around 0.46. In other words, the answer to "does anyone end up helping in a big crowd" is settled not by theory but by the rate of decay. And the rate of decay cannot be known by calculation. It has to be measured.
2020 — What 219 Pieces of CCTV Footage Answered
What Richard Philpot and his colleagues did is exactly that measurement. They obtained 219 real public conflicts recorded by surveillance cameras in the city centers of Amsterdam, Lancaster, and Cape Town, and coded them frame by frame. Not laboratory scripts but real fights. The three cities were chosen deliberately because their rates of violent crime differ sharply.
The result was this. In 90.9 percent of the 219 incidents, at least one bystander intervened. The average number of people who intervened was 3.76. And the group size effect ran the other way: the more people around, the higher the chance that somebody intervened. Differences in intervention rates among the three cities were not pronounced either.
You have to be precise about what this study overturns and what it does not. It did not refute the 1968 experiment. The laboratory measured the probability for an individual, and the cameras measured the outcome for the victim. It is the same structure as the left column and the right column of the table in the previous section being true at once. What collapsed is not the experiment but the conclusion the public drew from it — the sentence that when there are many people, nobody helps.
The limits deserve a fair statement too. The cameras show whether an intervention occurred but tell you nothing about whether it helped. The sample consists of visible conflicts in public places where cameras are installed, so it cannot be generalized to domestic violence or to situations in enclosed spaces.
What to Discard and What to Keep
Discard 1. The number 38. It did not pass verification. The practice of citing the anecdote in talks and articles has to go with it. Fortunately there is a far better story left in this case: the story of one neighbor who went down to the hallway at three in the morning.
Discard 2. The line that says in a crowd nobody helps. In more than 90 percent of real public conflicts, somebody intervenes.
Discard 3. Using this research as cynicism about human nature. The participants in the 1968 experiment who did not report were trembling. The bystander effect is not evidence of malice but a problem of situational design. Which is why design can fix it.
Keep 1. Why pointing works. "You there in the blue shirt, call for an ambulance and come back and tell me what they said." There is a reason first-aid training drills this sentence form over and over. This one sentence repairs all three failures at once. It specifies a target, making the denominator of responsibility one; it declares out loud that this is an emergency, breaking pluralistic ignorance; and it specifies what to do, removing evaluation apprehension. Learn the three mechanisms separately and you can see why this sentence has the shape it has.
Keep 2. The size of a single request. Moriarty ran an experiment on a beach in 1975. If you asked the person next to you to keep an eye on your radio for a moment and then left, 95 percent stopped a staged thief who came to take it. In the condition where no favor was asked and the person was only asked for the time, it was 20 percent. People who have explicitly been handed responsibility behave differently.
Keep 3. Rules for when you are the crowd. Two will do. One: act on the assumption that if you do not, nobody will. Two: do not read the calm of other people as information. They are reading your calm as information right now.
Keep 4. The same structure outside the emergency room. An email with several people in the recipient list, an incident alert channel with hundreds of members, a code review request everybody can see. All of them are places where the same failure happens. The prescription is the same too: name the owner. What an on-call rotation and a code owners file ultimately do is make the denominator of responsibility one.
Reading Guide
- Fact-checking the article: Manning, R., Levine, M., & Collins, A. (2007). The Kitty Genovese murder and the social psychology of helping: The parable of the 38 witnesses. American Psychologist, 62(6), 555-562.
- The original experiment: Darley, J. M., & Latane, B. (1968). Bystander intervention in emergencies: Diffusion of responsibility. Journal of Personality and Social Psychology, 8(4), 377-383.
- Pluralistic ignorance: Latane, B., & Darley, J. M. (1968). Group inhibition of bystander intervention in emergencies. Journal of Personality and Social Psychology, 10(3), 215-221.
- The effect of delegating responsibility: Moriarty, T. (1975). Crime, commitment, and the responsive bystander: Two field experiments. Journal of Personality and Social Psychology, 31(2), 370-376.
- Meta-analysis: Fischer, P., Krueger, J. I., Greitemeyer, T., Vogrincic, C., Kastenmuller, A., Frey, D., Heene, M., Wicher, M., & Kainbacher, M. (2011). The bystander-effect: A meta-analytic review on bystander intervention in dangerous and non-dangerous emergencies. Psychological Bulletin, 137(4), 517-537.
- Real-world field data: Philpot, R., Liebst, L. S., Levine, M., Bernasco, W., & Lindegaard, M. R. (2020). Would I be helped? Cross-national CCTV footage shows that intervention is the norm in public conflicts. American Psychologist, 75(1), 66-75.
Reading tip: in the 1968 paper, one table is the whole paper. Look first at the three-line table that sets reporting rates by group size beside mean response times, then find and read the paragraph in the discussion section describing the state of the non-intervening participants. That paragraph is the human core of this study. In the 2020 paper by Philpot and colleagues, the thing to look at is the graph of intervention probability by number of bystanders. That line, rising in exactly the opposite direction from the 1968 table, is the evidence that the two studies were measuring different things. The next installment is a reexamination of the Milgram obedience experiments, which asked the same question in the same period from the opposite side.