Sheep–Goat Effect
The sheep-goat effect is the finding, first reported by Gertrude Schmeidler in the 1940s, that people who believe ESP is possible tend to score above chance on ESP tests, while people who reject that possibility tend to score below chance. Both groups pull away from the baseline in opposite directions. That gap between them is one of the most replicated patterns in experimental parapsychology. It is also one of the most debated, with critics arguing that psychology, not psi, explains it entirely.
Key findings
- Schmeidler’s original studies at City College of New York showed that believers in ESP score above chance while disbelievers score below it. Both groups depart from the baseline in opposite directions.
- A 1993 meta-analysis by Tony Lawrence combined results across decades of studies. It confirmed the sheep-goat effect as statistically solid, with an overall effect that is small but points the same way across many labs.
- The effect has appeared in both card-guessing tasks and free-response tasks. This suggests it is not tied to one particular test format.
- Reactance theory research found that when disbelievers were put under psychological pressure, their below-chance scoring got stronger. This suggests that a person’s mental state shapes the effect.1
- A meta-analysis of sheep-goat ESP studies from 1947 to 1993 produced a result that would be extraordinarily unlikely if there were no real effect. This supports the idea that belief tracks experimental success.
- Psi belief is one of the more reliable individual predictors of psi performance across many kinds of studies.5
Overview
In the early 1940s, psychologist Gertrude Schmeidler asked a simple question before running ESP card tests: do you think you could score above chance here? Participants who said yes she called sheep. Those who rejected the possibility she called goats. What she found was striking. Sheep tended to score above the level expected by pure chance. Goats tended to score below it. Neither group showed a dramatic effect on its own. But the gap between them was consistent enough to attract serious attention.
Deeper dive: Schmeidler’s original classification
Schmeidler’s sheep-goat classification began in 1943 at City College of New York. She defined sheep as participants who accepted the possibility of ESP operating in that specific test. Goats were those who rejected that possibility (Schmeidler, 1943, 1952, 1959; Schmeidler and McConnell, 1958).910 The distinction was not about general paranormal enthusiasm. It was specifically about whether the participant thought ESP could show up in that particular experiment. This matters because it ties the classification to the test context, not just to a stable personality trait. Later researchers replaced the binary sheep-or-goat split with continuous belief scales. This allowed more sensitive statistical analysis.1
The phenomenon matters for a specific reason. In most ESP experiments, the overall hit rate stays close to chance. That makes it hard to claim any effect at all. The sheep-goat effect offers a different angle. Even when the average score across all participants looks unremarkable, splitting participants by belief can reveal that the two groups are pulling in opposite directions. That bidirectional pattern is harder to dismiss as random noise.4
Deeper dive: Why bidirectional scoring matters statistically
When sheep score above chance and goats score below chance, the two effects can cancel each other out in the overall data. The total sample then looks like a null result. This is why researchers focus on the sheep-goat difference score rather than the absolute hit rate. Palmer (1977) noted the sheep-goat effect as one of the more successfully replicated relationships in experimental ESP research.11 The overall effect size reported by Lawrence (1993) was very small (0.03 on a standard scale).12 A small effect size means the average score difference between sheep and goats is modest in absolute terms. But the consistency of direction across many independent studies is what makes the pattern notable.
How the Experiments Work
The basic design is straightforward. Before any ESP task begins, participants fill out a questionnaire or answer a direct question about whether they believe ESP is possible. This labels them as sheep or goats. They then attempt the ESP task. Usually this means guessing which of several symbols or images was randomly chosen as the target. After the session, researchers compare the hit rates of the two groups against the level expected by chance alone.
Deeper dive: Forced-choice and free-response formats
Forced-choice tasks ask participants to pick from a fixed set of options, such as five Zener card symbols or numbered balls. The chance baseline is mathematically precise, which makes the statistics clean. Free-response tasks give participants a broader target, such as a photograph or video clip. Participants describe or rank what they perceive. Both formats have been used to study the sheep-goat effect, and the pattern has appeared in both. One forced-choice tool used in recent research is Ertel’s Ball Selection Test. Participants predict which numbered ball will be drawn from a set. This test has been used to examine how belief and mood shape psi performance.1
A key design rule is that the belief measure must be collected before the ESP task, not after. If researchers asked about beliefs only after seeing scores, the classification could be contaminated by how well participants happened to perform. Collecting belief data first keeps the two measures independent.
Deeper dive: Blinding and randomization in sheep-goat studies
Blinding means the person scoring the ESP responses does not know which answer the participant gave, or which target was selected, until after the judgment is recorded. This prevents unconscious bias from shaping the outcome. Randomization means the target for each trial is chosen by a process that cannot be predicted or influenced by the participant, such as a random number generator or a shuffled deck. Both controls are standard in well-designed sheep-goat studies. When these controls are absent, results are more open to the criticism that sensory cues or experimenter bias produced the apparent effect, rather than any real link between belief and psi performance.
What the Evidence Shows
Any single experiment produces a small, ambiguous result. To see whether a pattern holds across many studies, researchers use a method called meta-analysis. The idea is to combine data from many separate experiments into one large analysis. When you pool that much data, even a small consistent effect becomes detectable. A result is called statistically significant when it is unlikely enough to have occurred by chance alone. The standard threshold is a probability below 5 in 100, written as p less than 0.05. Results well below that are described as highly significant.
Tony Lawrence’s 1993 meta-analysis combined sheep-goat ESP studies spanning decades and multiple labs. The overall effect was small in absolute terms but pointed the same way across studies. The combined result was statistically significant. A separate analysis covering studies from 1947 to 1993 produced a result that was extraordinarily unlikely to be chance. This supports the conclusion that belief in psi tracks experimental success across a wide range of settings.
“The sheep-goat effect is one of the more successfully replicated relationships in experimental ESP research” (Palmer, 1977), even though the overall effect size is very small (Lawrence, 1993).
Deeper dive: Lawrence’s meta-analysis and effect size
Lawrence (1993) is the main quantitative summary of the sheep-goat literature. The overall effect size was 0.03. By standard conventions, a Cohen’s d below 0.2 is considered small. However, the consistency of direction across independent studies is what gives the meta-analysis its force. The combined z-value from the 1947-1993 analysis reported in one theoretical review was 8.17. This corresponds to a probability of approximately 1.33 times 10 to the power of negative 16, meaning the result is extraordinarily unlikely to be a chance fluctuation. A later meta-analytic update combined forced-choice sheep-goat studies from 1947 to 1993 and found the effect statistically robust and stable across that span: neither study quality nor effect size declined over the 46 years. This runs counter to the decline effects sometimes claimed in parapsychology.12
The effect has also appeared in studies not primarily designed to test it. Researchers working on ganzfeld ESP, dream telepathy, and telephone telepathy have all noted that belief in psi predicts better performance. This fits the sheep-goat pattern.37 This consistency across different test types is one reason defenders argue the effect reflects something real about expectation and psi performance.
Deeper dive: Sheep-goat effect across paradigms
In ganzfeld research, belief in psi predicts above-chance performance. The sheep-goat link is described as fairly consistent throughout the parapsychological literature.3 In telephone telepathy studies, participants who accepted the possibility of psi showed higher hit rates than those who did not. Researchers described this explicitly as an example of the sheep-goat effect.7 In dream precognition research, psi belief was one of the only individual measures to correlate significantly with psi performance.8 The pattern also appears in student research projects at the Koestler Parapsychology Unit, where a small number of process-oriented projects found evidence supporting the sheep-goat effect.
Psychological Explanations
Even researchers who accept the sheep-goat effect as a real statistical pattern disagree about what it means. One prominent line of work asks whether the effect reflects psi at all, or whether ordinary psychological processes produce the same scoring pattern.
One psychological framework is reactance theory. Reactance is a motivational state that arises when a person feels their freedom is being threatened. When someone is pushed to do something, they often push back. Applied to psi testing, the idea is that goats may respond to the implicit demand of the experiment by actively resisting. That resistance could push their scores below chance. Not because of any psi process, but because of ordinary psychological noncompliance.1
Deeper dive: Reactance experiments and the sheep-goat effect
Storm (2013) tested reactance theory directly using the Ball Selection Test with 82 participants. They were randomly assigned to a control condition or a reactance condition. The reactance condition used a message designed to threaten participants’ sense of freedom. The overall hit rate across all 12,016 trials was 21.06% against a chance baseline of 20%. This was statistically significant (p = .002, meaning: if there were no real effect, a result this strong would occur by chance only about 2 times in 1,000). The reactance group scored lower (20.26%) than the control group (21.74%). The difference reached conventional significance (F(1,77) = 2.75, p = .05). Reactant goats scored significantly lower than control sheep. A continuous belief measure (the RASGS scale) correlated positively with psi hit rates (r = 0.20, p = .036). This supports the sheep-goat relationship even when belief is treated as a sliding scale rather than a binary split. Pre-test tension and confusion also predicted psi outcomes, suggesting mood states play a role.1
A follow-up study tested both reactance and a positive psychological technique called imagery cultivation. This technique uses guided visualization to encourage a receptive mental state. The results were more complex than expected. Goats showed significantly greater disagreement with the reactance message than sheep did. Participants who were not bothered by the threatening message tended to score better than those who were. This suggests that how a person responds to the experimental situation matters more than their belief label alone.2
Deeper dive: Imagery cultivation and reactance in a 2×2 design
Storm (2019) used a 2×2 factorial design with 240 psychology students. Participants were assigned to receive or not receive an imagery cultivation protocol and a reactance treatment. The imagery cultivation condition produced a non-significant but higher hit rate (21.0%) compared to control (20.7%). Contrary to the hypothesis, the reactance treatment produced higher hit rates (22.3%) than the no-reactance condition (19.3%). Neither difference reached conventional significance (F(1,223) = 1.84, p = .176). Goats showed significantly greater disagreement with the reactance message than sheep (t(238) = 8.40, p less than .001). A marginally significant discrepancy effect emerged: non-discrepants scored 24.8% versus discrepants at 16.8% (F(1,223) = 2.27, p = .067). The sheep-goat effect itself was marginally significant, with sheep at 22.0% versus goats at 19.7%.2
Another angle comes from cognitive dissonance research. The argument is that sheep score above chance because hitting the target fits their belief. Goats score below chance because hitting would contradict their worldview and create psychological discomfort. On this view, the scoring pattern reflects motivated thinking rather than any information moving between minds without a normal explanation.
Skeptical critiques and debates
The sheep-goat difference is produced by differential motivation rather than psi: sheep try harder and engage more seriously with the task, so better effort produces better scores through ordinary means.
Skeptic source: Critique drawn from the introductory-parapsychology literature on belief and scoring.
Rebuttal: Defenders note that the below-chance scoring of goats is hard to explain by motivation alone. If goats simply tried less hard, their scores should drift toward chance, not below it. The systematic below-chance pattern suggests active suppression of some kind. Reactance theory tries to explain this psychologically, but it still needs to account for why scores go in a specific direction rather than simply becoming random.1
Analysis. The motivation account and the psi account both predict that sheep outscore goats, so the headline direction does not separate them. The open question is the below-chance scoring of goats: a pure effort account predicts drift toward chance, while the observed systematic deviation below baseline is what reactance-theory work attempts to model. Whether ordinary noncompliance fully accounts for that directional shift, or whether something additional is required, remains unresolved in the cited record.
Sensory cues and experimenter effects explain the pattern: sheep may pick up subtle cues from experimenters who share their belief, producing the scoring difference through ordinary social psychology rather than any psi mechanism.
Skeptic source: Methodological critique on cueing and experimenter influence in ESP testing.
Rebuttal: Some studies have found sheep-goat effects even when the experimenter’s belief was controlled, or when automated testing removed direct experimenter contact. The Ball Selection Test research addressed sensory leakage by running trials with blindfolds and gloves, and the psi effect remained significant under those conditions. Experimenter effects nonetheless remain a genuine concern across parapsychology, and no single study has fully ruled them out.6
Analysis. The cited automated and shielded protocols address direct sensory leakage and report an effect persisting under those controls, which narrows the cueing explanation for those specific studies. Experimenter-effect work with believer and disbeliever experimenters keeps the broader concern live, since experimenter belief can itself operate as a variable. The debate centers on how far automation closes the gap and whether experimenter effects are a confound or themselves part of the phenomenon under study.
Optional stopping, flexible analysis, and the file-drawer problem inflate the effect: researchers may have stopped data collection when results looked favorable, or tested multiple participant splits, and null studies may never have been published.
Skeptic source: Critique on analytic flexibility and publication bias in the sheep-goat literature.
Rebuttal: The meta-analytic record tried to account for the file-drawer problem by estimating how many unpublished null results would be needed to reduce the combined effect to non-significance; the number required was large, which argues against publication bias as a sole explanation. More recent work using preregistered designs and continuous belief measures has also found the pattern, which reduces the force of the optional-stopping critique for those studies.1
Analysis. Fail-safe-style estimates and preregistered replications speak to different parts of the critique: the former addresses how much unpublished null evidence would be needed to overturn the pooled estimate, the latter constrains analytic flexibility within individual studies. Neither fully eliminates the concern for the older literature, where reporting practices were less standardized. The state of the debate is that the more recent, design-controlled subset is harder to attribute to these artifacts than the historical corpus.
The effect size is too small to be scientifically meaningful: even accepting the meta-analytic result, the overall effect is very small and could reflect minor systematic flaws rather than a genuine belief-performance link.
Skeptic source: Critique citing the small overall meta-analytic effect size.
Rebuttal: Defenders argue that small effect sizes are common across psychology and do not by themselves mean an effect is an artifact. They note that the sheep-goat effect size is comparable to effect sizes accepted as meaningful in other areas of behavioral research, and that the consistency of direction across independent labs is more informative than the absolute size of the effect.5
Analysis. Both sides agree the absolute effect (about 0.03 in the cited meta-analysis) is small; they disagree on what a small-but-consistent effect implies. One reading treats a small effect as within the range plausibly generated by uncontrolled minor flaws; the other treats directional consistency across many independent labs as the load-bearing observation. Resolving this turns on whether residual systematic flaws can plausibly produce a uniform direction across decades and laboratories, which the cited sources do not settle.
Where the Field Stands
The sheep-goat effect holds an unusual position in parapsychology. It is among the most replicated findings the field has produced. The pattern keeps appearing across labs, decades, and test formats. At the same time, the effect is small enough that its meaning stays genuinely open. Replication means the pattern keeps showing up. It does not settle whether the pattern reflects psi, psychology, or some mix of both.
Deeper dive: Replication and what it does and does not establish
Replication in science means that independent researchers, using different participants and sometimes different methods, get results pointing the same way as the original finding. A replication failure means a study using similar methods does not reproduce the original result. The sheep-goat effect has a strong replication record in terms of directional consistency: sheep tend to outscore goats across many independent studies. However, replicating a statistical pattern does not by itself establish its cause. A confound is a variable that was not controlled and that could explain the result through ordinary means. The central debate about the sheep-goat effect is whether known confounds, such as differential motivation, experimenter effects, or response biases, fully account for the replicated pattern, or whether something additional is needed.
One theoretical proposal holds that the sheep-goat effect is not incidental to psi research but central to it. On this view, skeptical attitudes do not merely reflect a personality trait. They actively suppress whatever process produces psi effects. This would explain why replication attempts by skeptical researchers sometimes fail: the experimenter’s own disbelief functions as a goat variable at the lab level. This idea is provocative precisely because it makes the effect hard to test under standard scientific conditions, where skeptical scrutiny is considered a virtue rather than a confound.
Deeper dive: The Model of Pragmatic Information and belief as a systemic variable
Walter von Lucadou’s Model of Pragmatic Information proposes that psi phenomena have a built-in tendency to evade definitive proof. The model predicts that high expectations and systematic replication attempts reduce or eliminate effects. Within this framework, the sheep-goat effect is not just a participant-level variable. Skeptical attitudes at any level of the experimental system, including the experimenter, the institution, or the broader scientific community, should suppress psi effects. A meta-analysis of sheep-goat ESP studies from 1947 to 1993 produced a combined z-value of 8.17 (p = 1.33 times 10 to the power of negative 16). That means: if there were no real effect, a result this extreme would occur by chance far less than once in a trillion tries. The model interprets this as evidence that belief in psi’s possibility tracks experimental success at a level far beyond chance. Critics note that this framing makes the hypothesis unfalsifiable: any failure to replicate can be attributed to the skepticism of those attempting the replication.
For researchers who accept the effect as genuine, the next question is what it tells us about how psi works. If believing you can do something helps you do it, that is not unique to parapsychology. Expectation shapes performance in many domains. What would be unusual is if the below-chance scoring of goats reflects genuine psi operating in reverse, guided by the expectation of failure. That possibility keeps the sheep-goat effect at the center of theoretical debates about the nature of psi, not just its existence.
Deeper dive: First Sight theory and the sheep-goat effect
James Carpenter’s First Sight theory proposes that psi is a continuous, preconscious process. The mind uses it to orient toward future and distant events before conscious awareness kicks in. On this model, the sheep-goat effect arises because sheep are open to using psi information in their responses, while goats actively avoid it. The below-chance scoring of goats is not a failure of psi. It is psi operating in the service of the goat’s intention to avoid confirming the hypothesis. This framework treats the sheep-goat effect as evidence for psi rather than against it. It predicts both above- and below-chance scoring as functions of the participant’s orientation toward the task.
References
- Storm, L. (2013). The Sheep-Goat Effect as a Matter of Compliance vs. Noncompliance: The Effect of Reactance in a Forced-Choice Ball Selection Test. Journal of Scientific Exploration, 27(3), 21. ↩︎
- Storm, L. (2019). Imagination and Reactance in a Psi Task using the Imagery Cultivation Model and a Fuzzy Set Encoded Target Pool. Journal of Scientific Exploration, 33(2), 16. ↩︎
- Dalton, K. S. (1997). Is there a formula to success in the ganzfeld? Observations on predictors of psi-ganzfeld performance. European Journal of Parapsychology, 13, 10. ↩︎
- Morris, R. L. (1999). Experimental systems in mind-matter research. Journal of Scientific Exploration, 13, 17. ↩︎
- Williams, B. J. (2019). Reassessing the “Impossible”: A Critical Commentary on Reber and Alcock’s “Why Parapsychological Claims Cannot Be True”. Journal of Scientific Exploration, 33(4), 18. ↩︎
- Watt, C. (2003). Experimenter effects with a remote facilitation of attention focusing task: A study with multiple believer and disbeliever experimenters. Journal of Parapsychology, 10. ↩︎
- Sheldrake, R., Stedall, T., & Tressoldi, P. (2025). Telecommunication Telepathy: A Meta-Analysis. Journal of Anomalous Experience and Cognition, 5(1), 23. ↩︎
- Luke, D. P., & Zychowicz, K. (2014). Working the graveyard shift at the witching hour: Further exploration of dreams, psi and circadian rhythms. International Journal of Dream Research, 8. ↩︎
- Schmeidler, G. R. (1943). Predicting good and bad scores in a clairvoyance experiment: A preliminary report. Journal of the American Society for Psychical Research, 37, 103-110. ↩︎
- Schmeidler, G. R., & McConnell, R. A. (1958). ESP and Personality Patterns. Yale University Press. ↩︎
- Palmer, J. (1977). Attitudes and personality traits in experimental ESP research. In B. B. Wolman (Ed.), Handbook of Parapsychology (pp. 175-201). Van Nostrand Reinhold. ↩︎
- Lawrence, T. (1993). Gathering in the sheep and goats: A meta-analysis of forced-choice sheep/goat ESP studies, 1947-1993. Proceedings of Presented Papers: The Parapsychological Association 36th Annual Convention, 75-86. ↩︎
Related topics: