Jessica M. Utts, PhD Sources:
Assessing Student Retention of Essential Statistical Ideas: Perspectives, Priorities, and Possibilities
This 2007 paper by Hayden, Utts, Mendoza, Roiter, and Garcia, presented in the proceedings of the Joint Statistical Meetings of the American Statistical Association, asks a basic but rarely-studied question of statistics education: what do introductory-statistics students still understand about essential statistical concepts months or years after the course ends? The paper situates the retention question within Utts’s broader literacy framework and identifies practical priorities for curriculum and assessment design.
Deeper dives — Utts:
Key findings
- The paper identifies long-term retention of statistical concepts, rather than immediate post-course performance, as the appropriate target for evaluating whether introductory statistics education has succeeded.1
- The authors argue that standard course-end exams measure short-term recall and do not establish whether students retain conceptual understanding beyond the course’s immediate horizon.1
- The paper proposes a set of priorities for which essential statistical ideas should be the focus of retention assessment, drawing on the conceptual targets Utts had articulated in her 2003 citizen-literacy paper.2
- The paper sits within an ongoing methodological discussion in statistics education about how to design assessments that measure conceptual understanding rather than procedural fluency, a discussion that subsequently informed assessment-instrument development including the ARTIST (Assessment Resource Tools for Improving Statistical Thinking) and CAOS (Comprehensive Assessment of Outcomes in a first Statistics course) projects.3
- The retention question reframes the goal of introductory statistics education: not “what do students know at the final exam” but “what do students still understand a year later when they encounter a statistical claim in news coverage or clinical practice.”1
Overview
The paper’s central observation is that the standard evaluation cycle in introductory statistics, course-end exams measuring immediate post-instruction performance, systematically misses the question that matters most for citizen literacy: do students still possess functional statistical reasoning when they encounter statistical claims months or years after the course ends? The authors argue that retention is a different cognitive target from acquisition, that the two can diverge sharply, and that an introductory course optimized for course-end performance may produce graduates who cannot apply the same concepts a year later.1
Connection to Utts’s Broader Literacy Agenda
The retention question is the natural follow-up to the literacy question Utts had posed in her 2003 American Statistician paper: if educated citizens should understand seven specific topics (causation inference, statistical-vs-practical significance, power and absence of evidence, survey bias, the prevalence of coincidence in large samples, the non-equivalence of conditional probability and its inverse, and the difference between “normal” and “average”), then the relevant question for curriculum evaluation is whether students retain understanding of those topics post-course.2 The 2007 retention paper formalizes this connection and proposes assessment priorities accordingly.
The Retention Question
The paper distinguishes three temporal horizons for evaluating statistical understanding: immediate post-instruction performance (the standard final-exam measurement), short-term retention (typically measured weeks to months after course completion), and long-term retention (measured a year or more after the course ends). The authors argue that the standard evaluation cycle conflates these three horizons, treating immediate performance as a proxy for the longer-term targets that actually matter for citizen literacy.1
Why the Three Horizons Can Diverge
The literature on long-term retention of academic content, drawn from broader cognitive-psychology research on learning and forgetting, indicates that immediate post-instruction performance is a poor predictor of long-term retention for most academic material; the two correlate but the relationship is far from one-to-one, and content that is well-retained immediately can be substantially forgotten over months and years.4 Statistics education has not been an exception to this pattern, and the 2007 paper draws on broader retention research to argue that introductory statistics curricula should be evaluated against retention measurements rather than against immediate post-course performance.
Priorities for Assessment Design
Rather than attempting to measure retention of every concept covered in an introductory course, the paper argues for prioritization: assessments should focus on the concepts whose retention most directly determines whether a graduate functions as a statistically literate citizen. The priorities follow from Utts’s 2003 framework but are adapted for the retention question: which conceptual targets, if retained, would most reduce the rate of consequential statistical error in adult life, and which, if forgotten, would most leave the graduate vulnerable to the standard failure modes documented in media coverage and clinical practice?1
Conceptual Understanding vs. Procedural Fluency
A recurring methodological tension in the assessment literature is the distinction between conceptual understanding (the ability to recognize and reason about a statistical claim in context) and procedural fluency (the ability to compute a statistic or apply a formula). The authors of the 2007 paper argue that retention assessments should prioritize the conceptual side: procedural fluency is the easier of the two to measure but the less consequential of the two to retain. A graduate who cannot recall how to compute a confidence interval from raw data is less impaired in citizen life than one who cannot recognize that a “50% increase in risk” can describe radically different absolute risks depending on the baseline.1 This priority connects directly to the seven-topic framework: each of the seven concepts is fundamentally conceptual rather than computational.
Modern Context
The statistics-education retention question Utts examined connects to a broader research program in statistics-education research that has matured since 2000. Garfield and Ben-Zvi’s Developing Students’ Statistical Reasoning: Connecting Research and Teaching Practice (Springer, 2008; 10.1007/978-1-4020-8383-9) provides the canonical synthesis of research on how introductory statistics students develop (and lose) statistical reasoning capabilities over time. Cobb’s influential 2007 essay “The Introductory Statistics Course: A Ptolemaic Curriculum?” (Technology Innovations in Statistics Education, 1(1); 10.5070/T511000028) argued that retention failures reflect a curriculum built around computational procedures rather than conceptual reasoning — a framing directly relevant to Utts’s empirical findings on what graduates retain.
Discussion: Retention as Curriculum Diagnostic
The paper’s broader argument is that retention measurement is not only a tool for evaluating individual students; it is a tool for evaluating curricula. A curriculum that produces high course-end performance but poor long-term retention is failing at the citizen-literacy goal even if it succeeds at the immediate-assessment goal. The retention measurement reveals the divergence and motivates curriculum reform: instructional methods that produce more durable understanding (worked examples connected to authentic contexts, spaced practice, conceptual elaboration) should be preferred over methods that maximize short-term performance at the expense of durability.1
The Paper’s Place in a Series
The 2007 paper is one of several Utts-co-authored or Utts-cited contributions to a broader discussion in statistics education about the goals and metrics of introductory instruction. It connects forward to the 2015 Baldi & Utts paper on biostatistics for future doctors (which extends the citizen-literacy framework to a specific professional audience), to the 2015 Utts historical review of 175 years of statistics-education themes, and to the 2022 Raman, Utts, Cohen, & Hayat paper on integrating ethics into the GAISE guidelines.5 The retention-assessment focus the 2007 paper advocates was subsequently incorporated, in modified form, into the ARTIST and CAOS assessment-instrument projects that became reference tools for evaluating introductory statistics curricula.3
Skeptical Critiques and Discussion
Critique 1: Long-term retention assessment is logistically impractical at the scale needed to inform curriculum policy
Skeptic source: A practical critique of retention-based curriculum evaluation is that tracking students for a year or more after course completion is logistically expensive, attrition-prone, and resistant to the rapid feedback cycles curriculum committees typically operate under. The argument is that short-term retention proxies, while imperfect, may be the most that curriculum-evaluation systems can realistically use.
Response: The 2007 paper acknowledges the logistical challenge but argues that the alternative, optimizing curricula against immediate-performance proxies that diverge from long-term retention, produces graduates who do not function as statistically literate citizens. The cost of retention measurement is high, but the cost of not measuring it, in the form of curricula that succeed on the wrong target, is higher. The authors propose that periodic retention studies, conducted on representative student samples, can inform curriculum policy without requiring tracking of every student.1
Analysis. The logistical concern is real; the authors’ proposed mitigation (representative-sample periodic studies rather than universal tracking) is methodologically defensible but has not been adopted as the dominant evaluation paradigm in the years since. The empirical question of whether short-term proxies adequately predict long-term retention remains partially open.
Critique 2: Focusing assessment on a small set of “essential” concepts risks narrowing the curriculum to the assessment list
Skeptic source: A standard concern in any assessment-driven reform proposal is Campbell’s Law: when an assessment becomes the target of curriculum optimization, instructional time and attention concentrate on the assessed concepts at the expense of unassessed but valuable material. The 2007 paper’s prioritization of essential concepts for retention assessment risks producing curricula that teach only those concepts.
Response: The authors are explicit that the prioritization is for assessment, not for instruction: instructors should teach a richer curriculum than the assessment measures, with the assessed concepts forming a literacy floor rather than a curriculum ceiling.1 Whether this distinction is preserved in practice, as the assessment instruments come into use, is an empirical question. Subsequent ARTIST and CAOS development has attempted to address the narrowing concern by including conceptual targets across a broader range than any single curriculum would emphasize, on the theory that students who pass the assessment will have encountered varied conceptual contexts.3
Analysis. The narrowing concern is methodologically real; whether assessment-driven reform in statistics education has narrowed curricula in practice is an empirical question on which the existing literature is mixed.
References
- Hayden, R. W., Utts, J., Mendoza, F. M., Roiter, K., & Garcia, F. (2007). Assessing Student Retention of Essential Statistical Ideas: Perspectives, Priorities, and Possibilities. Proceedings of the Joint Statistical Meetings, American Statistical Association. R001 [Hayden 2007] ↩︎
- Utts, J. (2003). What Educated Citizens Should Know About Statistics and Probability. The American Statistician, 57(2), 74–79. https://doi.org/10.1198/0003130031630 R002 [Utts 2003] ↩︎
- delMas, R., Garfield, J., Ooms, A., & Chance, B. (2007). Assessing students’ conceptual understanding after a first course in statistics. Statistics Education Research Journal, 6(2), 28–58. R003 [delMas 2007] ↩︎
- Roediger, H. L., & Karpicke, J. D. (2006). Test-enhanced learning: Taking memory tests improves long-term retention. Psychological Science, 17(3), 249–255. R004 [Roediger 2006] ↩︎
- Baldi, B., & Utts, J. (2015). What Your Future Doctor Should Know About Statistics: Must-Include Topics for Introductory Undergraduate Biostatistics. The American Statistician, 69(3), 231–240. https://doi.org/10.1080/00031305.2015.1048903 R005 [Baldi 2015] ↩︎
Deeper dives — Utts:
See hub COI disclosure for subject-coauthorship transparency.