Experimenter Effects
The experimenter effect is the influence the person running a study has on its outcome, through expectancy, manner with participants, and choices in procedure and analysis, or, on a stronger reading sometimes proposed in parapsychology, through a putative psi-mediated effect of the experimenter’s own intentions. In psi research it is a double-edged methodological concern: it is a candidate ordinary explanation for why some laboratories consistently obtain effects while others do not, and it is itself a phenomenon some researchers have tried to study directly. The most-cited example is the Wiseman-Schlitz collaboration, in which a skeptic and a proponent ran the same remote-staring protocol and reported different results.
Overview
An experimenter effect is the influence the person running a study has on its results. That influence can come through expectancy, the way the experimenter treats participants, and the small choices made in procedure and analysis. In parapsychology the idea carries a second, stronger reading: some researchers propose that the experimenter’s own intentions might shape outcomes through a psi-mediated process of their own. The two readings are very different, and keeping them apart is the central discipline of this topic.
The concern is double-edged. On one hand, the ordinary experimenter effect is a candidate normal explanation for why some laboratories keep getting positive results while others, running what looks like the same study, do not. On the other hand, the effect is itself a phenomenon that a few researchers have tried to measure directly. The best-known test is the collaboration between a skeptic and a proponent who ran the same remote-staring protocol together and reported different results.[1]
Two cautions frame everything below. First, the ordinary experimenter effect is a control issue, not a paranormal claim. Second, an experimenter effect is not fraud: expectancy and procedural influence can be unconscious and unintended, whereas fraud is deliberate.
Protocol
There is no single experiment that defines the experimenter effect. Instead there is a family of designs meant to isolate the experimenter as a variable. The cleanest approach holds everything else fixed, the participant pool, the equipment, the written procedure, the analysis plan, and varies only the person in charge. If outcomes still differ, the difference is attributed to the experimenter rather than to the participants or apparatus.
The most-cited template used identical protocol, equipment, and the same participant pool, with a skeptic and a proponent each running sessions of a remote-staring task. In that task a participant is watched or not watched from another room while electrodermal activity is recorded, and the question is whether the body registers being stared at.[1] Later variants moved the focus to a remote facilitation of attention task, where one person tries to help a distant participant concentrate.[2]
A complementary design recruits many experimenters at once, some believers, some disbelievers, so that the experimenter’s stance can be treated as a measurable factor rather than a single comparison.[2] A statistical version of this idea treats the experimenter as a random factor in a mixed model, which in principle lets an analyst separate experimenter-linked variance from any underlying psi signal.[3]
History
The experimenter effect entered psychology as a general methodological worry. Robert Rosenthal’s work in the 1960s on experimenter expectancy showed that an experimenter’s beliefs could quietly steer results, which is why blind procedures became standard practice across the behavioral sciences. That background was imported into psi research, where reviews in the 1970s catalogued how strongly outcomes seemed to track who was running the study. Later experimenter-effect reports trace that catalog to reviews by Kennedy and Taddonio (1976) and by Palmer, which enumerated the mechanisms proposed for the effect, from unintended differences in procedure to the experimenter’s own putative psi.[13]
The issue came to a head with the 1997 collaboration between a proponent and a skeptic on remote staring. Because the two ran identical procedures with the same participants, ordinary differences in equipment or sampling could be ruled out, leaving the experimenter as the prime suspect for the divergence.[2] Follow-up studies pursued the same question through the attention-focusing task, and one of them turned up a procedural artifact that had to be corrected before the experimenter comparison could be trusted.[4]
A third joint study later added two further investigators and repeated the comparison under a shared cross-over design, run at a single site, the Institute of Noetic Sciences’ psychophysiology laboratory.[15] A structured interview project then tried to pin down the tacit knowledge involved, the qualitative differences in how each researcher prepared, interacted with participants, and held intention during a session.[5]
Key findings
The headline result is that experimenters who differ in belief sometimes obtain different outcomes from the same protocol, but the pattern is not stable enough to call settled. In the original staring collaboration the proponent’s sessions and the skeptic’s sessions diverged even with shared participants and equipment.[1] Attempts to reproduce that contrast in the attention-focusing paradigm gave mixed answers, and at least one promising signal dissolved once an artifact was found.[4]
Studies that recruited multiple believer and disbeliever experimenters did not cleanly confirm a belief-driven split, which weakens the simplest version of the story. Interviews with the principals suggested the difference, where it appears, may lie less in stated belief than in subtle habits of preparation and rapport that resist standardization.[5]
When the same remote-staring protocol, equipment, and participant pool were shared between a skeptic and a proponent, the outcomes still diverged, pointing the question at the experimenter rather than the participant.[1]
The joint studies’ own numbers carry both the pattern and its instability. In the 1997 study, receivers run by Wiseman did not differ from chance expectation (Wilcoxon z = −0.44, p = .64, two-tailed), while receivers run by Schlitz showed a significant effect (z = −2.02, p = .04), with electrodermal activity higher during stare than non-stare trials; the direct comparison between the two experimenters’ participants was not significant (t = 1.39, df = 30, p = .17).[1] In the 1999 replication, run with 35 participants per experimenter, Wiseman’s participants were again at chance (z = −0.39, p = .69, effect size −0.07) and Schlitz’s again reached the threshold (z = −1.93, p = .05, effect size −0.33), but this time her participants were significantly less activated during the stare than non-stare periods, the opposite direction from 1997, and the between-experimenter comparison was again not significant (t = −0.77, df = 68, p = .44).[14]
The third collaboration added Caroline Watt and Dean Radin and moved into a double steel-walled, electromagnetically and acoustically shielded chamber, with a 2 × 2 cross-over design separating who greeted the participant from who carried out the staring. Neither experimenter’s condition differed from chance: with Schlitz as both greeter and sender, z = −0.17, p = .87 (effect size −0.03, N = 25); with Wiseman in both roles, z = −0.35, p = .72 (effect size −0.07, N = 26). The main effects of greeter (p = .50) and sender (p = .64) and their interaction (p = .85) were all non-significant, and neither coded greeter-participant rapport (r = −.028, p = .86) nor Schlitz’s self-reported focus (p = .69) or expectation of success (p = .40) correlated with session outcome. The authors present the series as open to two competing readings, a genuine effect disrupted in the third study, or chance findings or undetected subtle artifacts in the first two, and do not adjudicate between them.[15]
Crucially, none of these results decides between the two readings. A divergence between experimenters is equally consistent with ordinary interpersonal and procedural differences and with a psi-mediated experimenter effect, and the collaboration’s authors themselves disagreed on which explanation fits.[4] Reviews of the broader claim, that an experimenter’s own psi could shape a study, have treated it as a live but unproven hypothesis rather than an established mechanism.
Individual studies in the reference library
The entries below list the experimenter-effect studies and quantitative syntheses currently held in the ESP-Nexus reference library on this question, including the joint staring series discussed above. For the electrodermal literature those staring studies sit in, Schmidt, Schneider, Utts and Walach (2004) enumerate 36 direct-mental-interaction studies and 15 remote-staring studies; the library holds the three joint Wiseman-Schlitz studies and the two syntheses listed here rather than that full set, so the list reports what the library holds, not what has been published. The entries also report different kinds of number, Wilcoxon z scores, effect sizes on different conventions, correlations, and hit rates, and these cannot be added together into one bottom-line figure.
Methodological critiques
The strongest critique is that experimenter effects are hard to distinguish from the very biases they are meant to reveal. In micro-psychokinesis research, one re-analysis argued that experimenter expectancy, conformity, and publication bias together could account for the small deviations from chance, leaving no residue for a paranormal explanation.[6] A pre-registered replication of a correlational micro-PK effect failed, and its post-hoc checks found that false experimenter expectations had little visible impact on the data, while noting that the initial-success-then-decline shape still echoed claims of unconscious experimenter psi.[7]
Skeptics also warn against a moving target. When a failed replication is explained away as a psi-mediated experimenter effect, the hypothesis can become unfalsifiable. Commentators have linked this slipperiness to a “trickster” pattern in the field, cautioning that long-term experimenter-psi and decline effects often fail to survive rigorous meta-analysis.[8] A blind-analysis matrix experiment that found no main effect but a marginal secondary signal was read by its author as exactly this kind of ambiguous case, where experimenter psi is invoked precisely where the data are weakest.[9]
Replication efforts on related paradigms have been blunt about the problem. Two pre-registered failures to reproduce a matrix experiment concluded the paradigm is not reliably repeatable.[10] Expectancy-as-moderator hypotheses fared no better in two large pre-registered replications of a time-reversed priming task, where the main psi prediction did not hold.[11]
Current practice
Contemporary psi work treats the experimenter effect mainly as something to control. Standard defenses include blinding the experimenter to condition, automating stimulus presentation and scoring so a human cannot nudge outcomes, and pre-registering the analysis plan to remove discretionary choices after the data arrive. These measures target the ordinary reading directly.
Where researchers want to study the experimenter as a variable rather than suppress it, the favored move is to run multi-experimenter projects and model the experimenter as a random factor, which can in principle separate experimenter variance from any intrinsic effect.[3] Some teams have built remote internet platforms to gather many sessions across many handlers, partly to make this kind of partition possible at scale.
Broad reviews of the experimental literature acknowledge that small effect sizes and patchy replication keep the field’s claims open, and they place the experimenter question among the reasons results travel unevenly between laboratories. Umbrella syntheses report that some protocols yield steadier signals than others, which is itself relevant to whether a given lab’s success reflects method, participant selection, or the experimenter.[12] The stronger psi-mediated reading remains unsettled: it is taken seriously enough to design studies around, but it has not been established, and much of the field still treats the ordinary experimenter effect as the more defensible explanation for divergent results.
References
- Wiseman, R., & Schlitz, M. (1997). Experimenter effects and the remote detection of staring. Journal of Parapsychology, 61(3), 197–208. http://www.richardwiseman.com/resources/staring1.pdf R001 [Wiseman 1997] ↩︎
- Watt, C., & Ramakers, P. (2003). Experimenter effects with a remote facilitation of attention focusing task: A study with multiple believer and disbeliever experimenters. Journal of Parapsychology, 67, 99–116. https://koestlerunit.wordpress.com/wp-content/uploads/2015/06/watt-ramakers-2003.pdf R002 [Watt 2003] ↩︎
- Bierman, D., & Jolij, J. (2020). Dealing with the Experimenter Effect. Journal of Scientific Exploration, 34(4), 703–709. https://journalofscientificexploration.org/index.php/jse/article/view/1871 R003 [Bierman 2020] ↩︎
- Watt, C. (2002). Experimenter effects and the remote facilitation of attention focusing: Two studies and the discovery of an artifact. The Journal of Parapsychology, 66, 49–71. https://koestlerunit.wordpress.com/wp-content/uploads/2015/06/watt-brady-2002.pdf R004 [Watt 2002] ↩︎
- Watt, C., Wiseman, R., & Schlitz, M. (2005). Tacit knowledge in remote staring research: An interview with Marilyn Schlitz and Richard Wiseman. Zeitschrift für Anomalistik / Journal of Anomalistics, 5, 244–256. https://www.anomalistik.de/images/pdf/zfa/zfa2005_23_244_watt.pdf R005 [Watt 2005] ↩︎
- Pallikari, F. (2023). Understanding the Nature of Psychokinesis. Zeitschrift für Anomalistik / Journal of Anomalistics, 23(1), 103–131. https://doi.org/10.23793/zfa.2023.103 R006 [Pallikari 2023] ↩︎
- Maier, M., & Dechamps, M. (2022). A Pre-Registered Test of a Correlational Micro-PK Effect: Efforts to Learn from a Failure to Replicate. Journal of Scientific Exploration, 36(2), 251–263. https://journalofscientificexploration.org/index.php/jse/article/view/2235 R007 [Maier 2022] ↩︎
- Storm, L. (2023). The Dark Spirit of the Trickster Archetype in Parapsychology. Journal of Scientific Exploration, 37(4), 665–682. https://journalofscientificexploration.org/index.php/jse/article/view/2715 R008 [Storm 2023] ↩︎
- Grote, H. (2021). Mind-Matter Entanglement Correlations: Blind Analysis of a new Correlation Matrix Experiment. Journal of Scientific Exploration, 35(2), 287–310. https://journalofscientificexploration.org/index.php/jse/article/view/1931 R009 [Grote 2021] ↩︎
- Walach, H., Kirmse, K., Sedlmeier, P., Vogt, H., Hinterberger, T., & von Lucadou, W. (2021). Nailing Jelly: The Replication Problem Seems to Be Unsurmountable. Two Failed Replications of the Matrix Experiment. Journal of Scientific Exploration, 35(4), 788–828. https://journalofscientificexploration.org/index.php/jse/article/view/2031 R010 [Walach 2021] ↩︎
- Schlitz, M., Bem, D., Marcusson-Clavertz, D., Cardeña, E., Lyke, J., Grover, R., Blackmore, S., Tressoldi, P., Roney-Dougal, S., Bierman, D., Jolij, J., Lobach, E., Hartelius, G., Rabeyron, T., Bengston, W., Nelson, S., Moddel, G., & Delorme, A. (2021). Two Replication Studies of a Time-Reversed (Psi) Priming Task and the Role of Expectancy in Reaction Times. Journal of Scientific Exploration, 35(1), 65–90. https://journalofscientificexploration.org/index.php/jse/article/view/1903 R011 [Schlitz 2021] ↩︎
- Tressoldi, P., & Storm, L. (2021). Anomalous Cognition: An Umbrella Review of the Meta-Analytic Evidence. Journal of Anomalous Experience and Cognition, 1(1-2), 55–72. https://journals.lub.lu.se/jaex/article/view/23206 R012 [Tressoldi 2021] ↩︎
- Watt, C., & Wiseman, R. (2002). Experimenter differences in cognitive correlates of paranormal belief and in psi. Journal of Parapsychology, 66, 371–385. https://koestlerunit.wordpress.com/wp-content/uploads/2015/06/watt-wiseman-2002.pdf [PDF] R013 [Watt & Wiseman 2002] ↩︎
- Wiseman, R., & Schlitz, M. (1999). Experimenter effects and the remote detection of staring: An attempted replication. Proceedings of Presented Papers: The Parapsychological Association 42nd Annual Convention. [citation pending verification] R014 [Wiseman 1999] ↩︎
- Schlitz, M., Wiseman, R., Watt, C., & Radin, D. (2006). Of two minds: Sceptic–proponent collaboration within parapsychology. British Journal of Psychology, 97, 313–322. https://doi.org/10.1348/000712605X80704 [DOI] R015 [Schlitz 2006] ↩︎
- Schlitz, M., & Braud, W. (1997). Distant intentionality and healing: Assessing the evidence. Alternative Therapies in Health and Medicine, 3(6), 62–73. https://pubmed.ncbi.nlm.nih.gov/9375431/ R016 [Schlitz & Braud 1997] ↩︎
- Schmidt, S., Schneider, R., Utts, J., & Walach, H. (2004). Distant intentionality and the feeling of being stared at: Two meta-analyses. British Journal of Psychology, 95(2), 235–247. https://doi.org/10.1348/000712604773952449 [DOI] R017 [Schmidt 2004] ↩︎
- Sherwood, S. J., Roe, C. A., Holt, N. J., & Wilson, S. (2005). Interpersonal psi — Exploring the role of the experimenter and the experimental climate in a ganzfeld telepathy task. European Journal of Parapsychology, 20(2), 150–172. https://ejp.wyrdwise.com/EJP%20v20-2.pdf [PDF] R018 [Sherwood 2005] ↩︎
- Roe, C. A., Davey, R., & Stevens, P. (2006). Experimenter effects in laboratory tests of ESP and PK using a common protocol. Journal of Scientific Exploration, 20(2), 239–253. https://koestlerunit.wordpress.com/wp-content/uploads/2015/06/roe-2006.pdf [PDF] R019 [Roe 2006] ↩︎
- Roe, C. A., Sherwood, S. J., Farrell, L., Savva, L., & Baker, I. S. (2007). Assessing the roles of sender and experimenter in dream ESP research. European Journal of Parapsychology, 22(2), 175–192. https://koestlerunit.wordpress.com/wp-content/uploads/2015/06/roe-2007.pdf [PDF] R020 [Roe 2007] ↩︎
- Smith, M. D., & Savva, L. (2008). Experimenter effects in the ganzfeld. Proceedings of Presented Papers: The Parapsychological Association 51st Annual Convention, 238–249. https://fbial.yggycloud.com/getmedia.aspx?guid=cfcdceb45caa20df39cc0b0f07a3ae45#pp238-249 [PDF] R021 [Smith 2008] ↩︎