Science classrooms and laboratories around the world are being reshaped by an idea that has quietly accumulated four decades of evidence: students learn science and engineering best not by memorizing facts, but by building, testing, and revising models of the systems they study. A sweeping new systematic review published in the International Journal of STEM Education confirms that this approach, known as model-based reasoning, is now one of the most robust frameworks for understanding how scientific thinking actually works—and how it can be taught.
The review, conducted by Abasiafak N. Udosen and Alejandra J. Magana of Purdue University, synthesized 146 peer-reviewed studies published between 1980 and 2025. Following the PRISMA 2020 reporting guidelines, the researchers started from an enormous initial pool of over 9.4 million records across nine bibliographic databases, including Web of Science, ACM Digital Library, Google Scholar, ProQuest, SpringerLink, Wiley Online Library, Taylor & Francis, ERIC, and APA PsycInfo. After multiple stages of identification, screening, and eligibility assessment, the final corpus comprised 99 peer-reviewed journal articles, 24 book chapters, 14 books, 8 conference papers, and one thesis. The sheer scale of the filtering process underscores both the richness of the field and the difficulty of pinning down what model-based reasoning actually means.
At its core, model-based reasoning is the iterative process of constructing, retrieving, using, evaluating, and refining scientific models—whether computational, mathematical, diagrammatic, physical, or mechanistic—to make predictions or explain observed outcomes of real-world systems. The theoretical foundation traces back to the mental models framework developed by cognitive scientist Philip Johnson-Laird, which holds that humans reason by constructing internal, situation-specific simulations that represent the structure and behavior of external systems. Rather than relying purely on formal deductive logic, which often proves too rigid for the complexity and open-endedness of real scientific problems, model-based reasoning integrates abductive, inductive, deductive, causal-mechanistic, analogical, and computational forms of reasoning into a single flexible architecture. As philosopher of science Ronald Giere famously put it, scientific reasoning is “models almost all the way up and models almost all the way down.”
One of the review’s most striking findings is the identification of a consistent temporal structure in how different types of reasoning dominate different stages of the modeling cycle. During early problem analysis and problem formulation, learners rely heavily on abductive reasoning—generating hypotheses that might explain puzzling observations—supported by analogical reasoning, visual reasoning, and causal-mechanistic thinking. As they move into model construction and execution, deductive, quantitative, algorithmic, and reductive reasoning take over, driving the formulation of equations, the writing of code, and the running of simulations. Finally, during verification, validation, and debugging, diagnostic, inductive, probabilistic, and quantitative reasoning become dominant as modelers compare predictions against evidence, identify mismatches, and refine their work. This stage-based pattern, the authors argue, is not a rigid prescription but a robust epistemic organization visible across classrooms and professional laboratories alike.
The review also highlights a deep theoretical tension within the field: is model-based reasoning fundamentally an individual cognitive process rooted in mental models, or is it a socially and materially distributed practice? The evidence increasingly supports the latter view. Nancy Nersessian’s landmark five-year cognitive-historical ethnography of two university biomedical engineering laboratories—documented through roughly 800 hours of field notes and complete transcripts of 148 interviews—showed that graduate students solve problems by coordinating an “inter-locking models” ecosystem of computational flow models, benchtop prototypes, tissue-engineered constructs, and differential equations. Mastery emerged not from any single model but from the distributed coordination of models across people, tools, and time. This finding reframes modeling as a fundamentally collective and material enterprise rather than a purely internal mental exercise.
The review identifies broad consensus across the literature on several points. Virtually all accounts agree that model-based reasoning is an iterative process of constructing, testing, and revising models that stand in for real-world systems. There is also widespread agreement that external representations—diagrams, equations, prototypes, code, and simulations—do more than display ideas; they actively mediate reasoning by offloading cognitive load, coordinating collaborative talk, and preserving revision history. Scaffolding in the form of structured tasks, code prompts, project milestones, and software tools consistently amplifies the quality of model-based reasoning, helping learners progress from interpreting existing models to building and defending their own.
Yet disagreements persist on several fronts. Scholars remain divided over whether model-based reasoning is best grounded in mental-model theory, abductive cognition, distributed cognition, or socially regulated frameworks. There is also disagreement about whether different reasoning modes should be treated as analytically separable—drawing some support from neuroimaging evidence showing that inductive and deductive reasoning activate distinct brain regions—or whether they are best understood as hybrid, multimodal blends that resist clean partitioning. A third fault line concerns domain specificity: mental-model theorists often present their accounts as broadly cognitive and cross-disciplinary, while discipline-specific researchers argue that each field sets its own “rules of the game” for what counts as a good model, whether mechanism-rich explanation in biology, quantitative prediction in physics, or design-oriented intervention in engineering.
How researchers measure model-based reasoning turns out to shape what they can claim about it, and the review identifies three distinct levels of analysis. At the micro level, think-aloud protocols and time-stamped coding capture moment-to-moment reasoning moves. One illustrative study by Ríos and colleagues had ten upper-division physics students troubleshoot an inverting-amplifier circuit while verbalizing every thought, with synchronized audio-video capture segmented into 30-second intervals coded for five modeling subtasks: construct, measure, compare, propose cause, and revise. Students spent most of their time in rapid-fire loops of measuring, comparing, and revising—often cycling through all three moves in under a minute. At the meso level, computational notebooks, simulation logs, and rubric-scored artifacts reveal workflow structure and representational competence. Magana and colleagues analyzed scaffolded Jupyter notebooks by sorting each cell into one of four modeling phases and applying validated rubrics for code accuracy, graphical interpretation, and explanatory coherence, producing numeric indices of how well students reasoned with their models. At the macro level, Model-Evidence Link diagrams and portfolios capture longer-term development over weeks or semesters, tracking how students’ coordination of evidence and explanation grows in sophistication over instructional time.
The pedagogical implications of the review are concrete and actionable. Courses should be organized around visible iteration—build-run-compare-revise loops—with explicit handoffs between diagrams, equations, code, and graphs, and with routine opportunities for students to reconcile mismatches between prediction and observation. Reasoning-mode scaffolds should be matched to modeling stage: analogies during problem analysis, unit checks and small parameter changes during solution construction, and targeted verification and validation near the end. Assessment should credit the quality of assumptions, traceable revisions, explicit validation criteria, and model-evidence coordination rather than rewarding only a final correct answer. The authors also emphasize that evidence-based reasoning and model-based reasoning should not be scored as separate activities but treated as intertwined components of a single sensemaking practice, since models provide the conceptual and material space in which diverse forms of reasoning interact and cross-check one another.
The review acknowledges several limitations. The final corpus is weighted more heavily toward science, engineering, and computing contexts than toward technology education or mathematics education. English-language and peer-review filters excluded potentially relevant work published in other languages or non-indexed formats. The mapping of studies to reasoning modes and modeling stages involved subjective interpretation, and the lack of inter-rater reliability on conceptual classifications may affect reproducibility. Many findings are also tied to specific disciplines, tools, and instructional environments, making transfer across STEM domains uneven. Finally, the field lacks standardized assessment instruments, complicating direct comparison across studies.
Despite these caveats, the synthesis offers a clear, evidence-based account of how learners use models to construct, test, evaluate, and refine explanations and predictions across STEM contexts. Model-based reasoning, the authors conclude, is not a niche technique but a common epistemic engine that can be tuned to biology, physics, engineering, and computing without abandoning its core architecture of iterative refinement. When instruction makes the modeling cycle visible, when students are supported to move across representations, and when assessment attends to process as well as product, learners develop the representational competence and metacognitive habits—planning, monitoring, evaluating—needed for authentic scientific inquiry. In an era where computational modeling and simulation are central to scientific practice, model-based reasoning offers a shared language through which diverse disciplines can cultivate the habits of mind that define genuine scientific work.
Cite Scienmag News
Courtney Benton. (September 6, 2026). Model-based reasoning in STEM education: systematic review of literature. Scienmag. https://scienmag.com/model-based-reasoning-in-stem-education-systematic-review-of-literature/
Courtney Benton. "Model-based reasoning in STEM education: systematic review of literature." Scienmag, 6 September 2026, https://scienmag.com/model-based-reasoning-in-stem-education-systematic-review-of-literature/. Accessed 6 September 2026.
Courtney Benton. "Model-based reasoning in STEM education: systematic review of literature." Scienmag. September 6, 2026. https://scienmag.com/model-based-reasoning-in-stem-education-systematic-review-of-literature/

