Marilyn J. Schlitz, PhD Sources:

Experimenter Effects in Parapsychology Research

Experimenter effects, the phenomenon whereby the identity or beliefs of the researcher systematically influence experimental outcomes, represent one of the most methodologically challenging puzzles in parapsychology. Marilyn Schlitz has been at the center of the field’s most rigorous investigation of this problem, conducting a decade-long collaborative program with skeptic Richard Wiseman in which identical protocols produced divergent results depending on who ran the session.

Key findings

  • In two initial joint studies using identical protocols, Schlitz’s participants showed statistically significant electrodermal responses during remote staring periods, while Wiseman’s participants did not.1
  • A third collaborative study failed to replicate the earlier experimenter-effect pattern, leaving the interpretation of the first two studies unresolved.2
  • The Schlitz–Wiseman program represents one of the few systematic, multi-study skeptic–proponent collaborations in parapsychology’s history, using simultaneous data collection at the same location with the same equipment and participant pool.
  • A large international multi-experimenter replication study of a time-reversed priming task found that the preregistered confirmatory hypothesis was not supported, though exploratory analyses identified language and expectancy as potential moderators.3
  • Schlitz has described experimenter effects as one of the central methodological challenges facing parapsychology, alongside funding constraints and replication difficulties.4

Overview

Experimenter effects in parapsychology refer to the documented pattern in which the same experimental protocol yields different outcomes depending on the researcher conducting the study. This is distinct from ordinary experimenter bias, it is not simply a matter of recording errors or demand characteristics, but a systematic divergence in results that correlates with the experimenter’s beliefs, expectations, or identity. The phenomenon poses a Type-II vulnerability for the field: if genuine psi effects are state-dependent and sensitive to the psychological context created by the experimenter, then null results obtained by skeptical investigators cannot straightforwardly falsify the hypothesis, and positive results obtained by proponents cannot straightforwardly confirm it. Schlitz has argued that this ambiguity is not a reason to abandon the research program but rather a reason to design studies that treat experimenter identity as an explicit variable.4

Experimenter Effects as a Type-II Vulnerability

A Type-II vulnerability in parapsychology is a methodological feature that makes genuine effects harder to detect, thereby inflating false-negative rates. Experimenter effects are particularly acute because they operate at the level of the experimental session itself: if a skeptical experimenter’s presence or demeanor suppresses psi performance, then well-powered studies run by skeptics will systematically underestimate effect sizes. Schlitz has noted that this creates a structural problem for replication: a skeptic who cannot replicate a proponent’s result has not necessarily disconfirmed the effect, they may have introduced a moderating variable. She has described this as one of the field’s core challenges alongside funding and the small researcher pool.4 The Schlitz–Wiseman collaboration was designed explicitly to test whether experimenter identity was such a moderating variable by holding all other conditions constant.

The Schlitz–Wiseman Joint Studies

The core of Schlitz’s contribution to experimenter-effects research is a systematic collaborative program conducted with Richard Wiseman, a prominent skeptic of parapsychological claims. The first two joint studies were designed to isolate experimenter identity as the sole varying factor: both researchers ran sessions simultaneously, at the same location, using the same equipment, drawing from the same participant pool, and following identical procedures. The measure was electrodermal activity (EDA), a physiological index of autonomic arousal, recorded while a remote experimenter either stared at a live video feed of the participant or looked away. Schlitz’s participants showed significantly elevated EDA during staring periods; Wiseman’s did not.1

Protocol Design of the First Two Joint Studies

In the joint studies, participants were connected to EDA recording equipment. A video camera captured a live image of the participant and transmitted it to a monitor in a separate room where the experimenter sat. Each session comprised 32 thirty-second periods, half randomly allocated to a staring condition and half to a non-staring condition. During staring periods the experimenter looked directly at the monitor image; during non-staring periods the experimenter looked away and directed attention elsewhere. The two experimenters ran their respective participant sets at the same time and in the same facility, eliminating site-level confounds. Schlitz’s data showed statistically significant EDA differences between stare and non-stare conditions; Wiseman’s data did not reach significance. The design was intended to rule out equipment artifacts (same equipment), participant sampling bias (same pool), and procedural variation (same protocol), leaving experimenter identity as the primary candidate explanation for the divergence.1

The Third Collaboration and Non-Replication

A third joint study, conducted with Caroline Watt and Dean Radin as additional collaborators, attempted to replicate the experimenter-effect pattern from the first two studies and to explore potential explanations for the earlier divergence. This study did not replicate the previous findings: neither Schlitz’s nor Wiseman’s data showed the pattern observed before.2 The non-replication left the interpretation of the original two studies genuinely open.

What the Third Study Found and What It Left Unresolved

The third study, reported in Schlitz, Wiseman, Watt, and Radin (2006), used the same basic EDA-staring paradigm as the earlier collaborations. The paper explicitly frames two competing interpretations of the overall three-study pattern: (1) the first two studies captured a genuine psi effect that was disrupted in the third study by some unidentified aspect of the new design or context; or (2) the first two studies represented chance findings or undetected subtle artifacts, and the third study’s null result more accurately reflects the absence of a remote staring effect. The paper does not adjudicate between these interpretations. The authors note that the collaborative design itself, a skeptic and proponent working jointly, offers a model for resolving disagreements in other controversial areas of psychology, regardless of the ultimate verdict on the psi question.2

Competing Interpretations

The Schlitz–Wiseman dataset admits at least three distinct interpretations, none of which is definitively ruled out by the available evidence. The first is a genuine psi interpretation: Schlitz’s belief in and openness to psi effects created a psychological or physical context that facilitated anomalous information transfer, while Wiseman’s skepticism suppressed it. The second is a subtle-artifact interpretation: undetected differences in how the two experimenters interacted with participants, tone, body language, instructions, or rapport, produced differential demand characteristics or autonomic priming that drove the EDA divergence through conventional channels. The third is a chance interpretation: the first two positive results were statistical fluctuations, and the third study’s null result is the more reliable estimate.

Demand Characteristics and Subtle Artifact as Non-Psi Explanations

The subtle-artifact interpretation centers on demand characteristics: participants assigned to Schlitz may have received implicit cues, through pre-session interaction, tone of voice, or framing of the task, that primed them toward greater autonomic reactivity, independently of any remote staring. This artifact was partially addressed by the joint-study design (same location, same equipment, same pool), which eliminated many obvious confounds. However, the design did not fully address experimenter–participant interaction during the consent and setup phases, which occurred separately for each experimenter’s participants. Schlitz and Wiseman acknowledge in the 2006 paper that subtle interaction differences during setup remain an unresolved potential confound, mitigated but not eliminated by the simultaneous same-site design. The psi interpretation, by contrast, proposes that the experimenter’s beliefs or intentions constitute a causal variable in the production of the effect, a claim that, if true, would have significant implications for how replication is understood across the field. The researcher’s preferred interpretation, as stated in the 2006 paper, is that the data are genuinely ambiguous and that further collaborative research is needed.2

Expectancy Effects in a Multi-Experimenter Replication Study

A separate line of evidence bearing on experimenter and participant expectancy comes from a large international multi-experimenter replication of Daryl Bem’s time-reversed priming task, coordinated by Schlitz and involving researchers from more than a dozen institutions. The preregistered confirmatory hypothesis, that response times to incongruent stimuli would be longer than to congruent stimuli even before the prime appeared, was not supported in Experiment 1 (N not specified in the extract; international multi-site sample, fixed-N stopping rule, preregistered).3 Exploratory analyses in Experiment 1 found that participants completing the English-language version showed a significant effect, while those using translated versions did not, a language-moderator pattern that was present in the original Bem study.3 Experiment 2 tested whether reading a pro-psi versus anti-psi statement at the outset would modulate performance. The primary psi hypothesis was not supported. However, exploratory analyses found that participants who received the pro-psi statement showed a larger psi score than those who received the anti-psi statement, a pattern consistent with expectancy moderating performance, though not reaching the threshold for the confirmatory hypothesis.3 Neither experimenters’ nor participants’ beliefs were consistently associated with the dependent measure across the full dataset. The personality variable Sensation Seeking, a component of extraversion, emerged as a correlate of psi performance in exploratory analyses.3

Broader Context: Expectancy and Belief in Psi Research

Schlitz has situated the experimenter-effects question within a broader argument about the sociology of science and the role of belief in shaping empirical outcomes. Drawing on her career-long observation that proponent and skeptic researchers tend to cluster with like-minded colleagues, she has argued that the Schlitz–Wiseman collaboration model, in which researchers with opposing priors run identical protocols, offers a structural solution to the credibility problem in parapsychology.4 The 2006 paper explicitly frames this as a methodological contribution independent of the psi verdict: skeptic–proponent collaborations can help resolve disagreements in any controversial area of psychology by making the experimenter a measured variable rather than an uncontrolled one.

The Collaborative Model as a Methodological Contribution

The Schlitz–Wiseman program is notable not only for its findings but for its design logic. By having a proponent and a skeptic run simultaneous sessions with the same equipment, same participant pool, and same procedures, the collaboration operationalized experimenter identity as an independent variable. This design addresses the multiple-comparisons artifact (both experimenters’ data are reported together, preventing selective reporting of only the significant arm), partially addresses demand characteristics (same setup procedures, though pre-session interaction was not fully standardized), and eliminates equipment and site confounds. The design does not address optional stopping (session counts were not preregistered in the early studies), and the third study’s failure to replicate introduces the possibility that the earlier positive results reflected chance fluctuations, an interpretation the authors explicitly acknowledge. Schlitz has described the collaborative model as potentially applicable to other contested areas of psychology where researcher belief may function as an uncontrolled moderator.4

Schlitz’s Autobiographical Account of Experimenter Effects

In her autobiographical account of coming of age in parapsychology, Schlitz describes experimenter effects as one of the field’s defining methodological puzzles, one she encountered early in her career and that has shaped her research agenda across institutions from the Rhine Research Center to the Institute of Noetic Sciences. She frames the Schlitz–Wiseman collaboration as a deliberate attempt to make the experimenter’s role explicit and measurable rather than treating it as noise to be averaged away. She also notes that the field’s small size and funding constraints make it difficult to accumulate the large, multi-lab datasets that would be needed to fully characterize experimenter effects as a moderating variable.4

Modern Context

Mainstream replication research offers useful framing for the Schlitz-Wiseman divergence. A direct experimental test of experimenter expectancy found that a classic priming effect reappeared only when experimenters anticipated the predicted outcome, demonstrating that expectancy alone can mechanistically drive results under otherwise identical conditions.[5] Separately, a large multi-lab replication project found that cross-site variation was modest compared with effect-level differences, suggesting that who runs a study can matter more than where it is run.[6] Both findings situate the experimenter-effect question the Schlitz-Wiseman collaboration was designed to probe within active, unresolved concerns in mainstream methodology.

Skeptical Critiques and Discussion

Critique 1: The experimenter-effect pattern may reflect undetected subtle artifacts in experimenter–participant interaction rather than a genuine psi phenomenon

Skeptic source: This interpretation is explicitly raised within the collaborative research program itself. Wiseman, as the skeptic co-author, argued that undetected differences in how Schlitz and Wiseman interacted with participants during setup and consent, differences in tone, framing, or rapport, could have produced differential autonomic priming through entirely conventional channels, mimicking a psi-mediated experimenter effect.

Response: The joint-study design partially addressed this artifact by holding location, equipment, participant pool, and procedures constant, eliminating the most obvious conventional confounds. However, the design did not fully standardize pre-session experimenter–participant interaction, leaving the demand-characteristics pathway unresolved. The 2006 paper acknowledges this explicitly: the subtle-artifact interpretation cannot be ruled out, and the third study’s non-replication is consistent with either the artifact account or the chance account.21

Analysis. The joint-study design mitigated but did not eliminate the subtle-artifact explanation. Across the three studies, two favored the experimenter-effect pattern and one did not, and the design as implemented could not distinguish psi-mediated experimenter effects from conventional demand-characteristics effects with the data collected.

Critique 2: The failure to replicate in the third study suggests the earlier positive results were chance findings, and the overall dataset does not support a reliable experimenter effect

Skeptic source: The third collaborative study, which used the same basic paradigm as the first two, failed to reproduce the pattern in which Schlitz’s participants showed significantly elevated EDA during staring periods. The skeptic interpretation of the three-study arc is that the first two positive results were statistical fluctuations, a plausible account given that neither study was preregistered and stopping rules were not disclosed, and the third study’s null result is the more accurate estimate of the true effect size.

Response: Schlitz and co-authors acknowledge this interpretation as one of two equally viable readings of the three-study dataset. The alternative reading, that some aspect of the third study’s design disrupted the production of the effect, cannot be ruled out either. The 2006 paper does not adjudicate between these interpretations, and the authors note that the collaborative design itself is a methodological contribution regardless of the psi verdict. The multi-experimenter replication study of Bem’s precognition task, which Schlitz coordinated, similarly found that the preregistered confirmatory hypothesis was not supported, though exploratory analyses identified potential moderators including language and expectancy framing.3

Analysis. The three-study Schlitz–Wiseman dataset is genuinely ambiguous: two positive results followed by a null result, with no preregistration and unresolved stopping rules across the series. The null result in the third study carries evidential weight but does not definitively resolve the question, because the design change between studies introduces the possibility of a disrupted effect rather than an absent one.

Study design ledger

Study design ledger — experimenter-effects evidence
StudyYearN (sessions/trials)PreregistrationPrimary outcomeResult-class
Wiseman & Schlitz, remote-staring (1st)199732 participantsnot-disclosedEDA differentiation; experimenter-effect pattern observedmixed (MS positive; RW null; between-experimenter not-significant)
Schlitz, Wiseman, Watt, Radin — Of two minds2006not-disclosed in pool entrynot-disclosedThird-collaboration test of experimenter-effect patternnull (failed to replicate prior pattern)
Schlitz, Bem, Cardeña et al., Bem precognition replication2021N₁ = 512; N₂ = 586 (two studies)preregistered (confirmatory)Time-reversed priming task; psi-priming hypothesisnull on confirmatory hypotheses; exploratory language + expectancy moderators
Schlitz, “Boundless mind”2001n/a (career retrospective)not-applicableCareer-context essay on parapsychologyessay
Doyen et al., behavioral priming replication2012not-disclosed in pool entrynot-disclosedReplication failure of social-priming-with-experimenter-blinding studynull (informs experimenter-expectancy literature)
Klein et al., Many Labs 1201436 independent samplespreregistered (Many Labs protocol)Cross-lab replication variability for 13 classic effectsmixed (variability quantified as evidence-class observation)

Cells marked “not-disclosed” reflect genuine absence from the source paper (typical of pre-2015 parapsychology corpus where preregistration + stopping rules were not standard methodological practice). Per the ESP-Nexus honesty mandate, the gap is surfaced rather than filled with inference.

References
  1. Wiseman, R., & Schlitz, M. (1997). Experimenter effects and the remote detection of staring. Journal of Parapsychology, 61(3), 197–208. http://www.richardwiseman.com/resources/staring1.pdf [Wiseman & Schlitz 1997] R001 ↩︎
  2. Schlitz, M., Wiseman, R., Watt, C., & Radin, D. I. (2006). Of two minds: Skeptic‐proponent collaboration within parapsychology. British Journal of Psychology, 97(3), 313–322. https://doi.org/10.1348/000712605X80704 [Schlitz 2006 Of Two Minds] R002 ↩︎
  3. Schlitz, M., Bem, D., Cardeña, E., et al. (2021). Two Replication Studies of a Time-Reversed (Psi) Priming Task and the Role of Expectancy in Reaction Times. Journal of Scientific Exploration, 35(1), 65–90. https://journalofscientificexploration.org/index.php/jse/article/view/1903 [Schlitz 2021] R003 ↩︎
  4. Schlitz, M. (2001). Boundless mind: Coming of age in parapsychology. Journal of Parapsychology, 65, 335-350. marilynschlitz.com [Schlitz 2001] R004 ↩︎
  5. [Doyen 2012] Doyen, S., Klein, O., Pichon, C.-L., & Cleeremans, A. (2012). Behavioral priming: It’s all in the mind, but whose mind? PLOS ONE, 7(1), e29081. https://doi.org/10.1371/journal.pone.0029081 R005 ↩︎
  6. [Klein 2014] Klein, R. A., et al. (2014). Investigating variation in replicability: A “many labs” replication project. Social Psychology, 45(3), 142–152. https://doi.org/10.1027/1864-9335/a000178 R006 ↩︎

Last updated: 2026-07-03 18:26:23

Copyright © 2026 Innovative Software Design. All rights reserved.

ESP-Nexus reports published parapsychology research as written by its own authors. Source-fidelity to cited documents is audited on every page; the site takes no position on the epistemic status of parapsychological claims themselves.