Why Misconceptions Survive Traditional Exams
August 11, 2026
5 min read
A student can select the correct answer on an exam question and still be wrong about the concept behind it. Researchers studying the Force Concept Inventory, one of the most widely used conceptual tests in physics education, call this a false positive: a correct response reached through flawed reasoning. In one analysis of student explanations for a question on atmospheric force, students who had never studied air pressure guessed correctly more often than students who had studied it but misunderstood it, and their post-instruction accuracy actually declined (Physical Review Physics Education Research).
That result cuts against how most exams are graded. A score assumes that a right answer reflects right understanding. When it doesn't, the misconception behind the wrong reasoning stays intact, unflagged, and ready to resurface in the next course that depends on it.
The exam format hides the reasoning it should be testing
Multiple-choice and short-answer exams are built around a single output: the final answer. They are efficient to grade and easy to scale, but they were never designed to expose how a student arrived at that answer. A student can memorize the correct label for a process, pattern-match it to a familiar question format, or narrow four options down through elimination, all without holding a coherent model of the underlying concept.
Misconceptions are not random errors. They are internally consistent, often useful in narrow contexts, and resistant to correction because they let students solve a subset of problems successfully. A physics student who believes force is required to sustain motion, rather than only to change it, will still get many textbook problems right. The exam rewards the output and never asks the student to defend the model that produced it.
Repetition without correction reinforces the wrong idea
The problem compounds with exposure. Research on misconception remediation in science education describes a documented risk: presenting a misconception again, even to correct it, increases its familiarity, and familiarity is one of the strongest predictors of perceived truth. If a course repeatedly surfaces a misconception on practice sets and exams without ever requiring the student to articulate why it's wrong, the misconception can become more entrenched, not less.
This is part of why simply adding more practice questions doesn't reliably fix conceptual gaps. Research from the National Research Council notes that even strong students frequently give correct answers using only memorized terms, and when questioned further, reveal that they never developed the underlying concept in the first place (National Academies Press). The gap isn't visible until someone asks the student to explain, not just answer.
Explaining a concept exposes what selecting an answer conceals
Written self-generated questions offer one alternative signal. In a study of medical students, researchers analyzed the questions students wrote about course material and found that the presence of a detectable misconception in a student's own question was negatively associated with their formal exam score, meaning the misconception was catchable before the exam, not just visible in hindsight (NCBI). Generating an explanation, rather than selecting one, forces the gap into the open.
This is the mechanism Axiom Flow is built around. Rather than asking a student to pick an answer, Axiom Flow has the student teach the concept to an AI student, Sam, who starts every session holding a set of misconceptions generated by Atlas, Axiom Flow's assessment designer. Sam has no independent way to check what's true. He accepts what he's taught, and his understanding shifts only when the student corrects him directly. A student who has memorized the right label without the right model will struggle to correct Sam's misconceptions convincingly, because Sam will keep reflecting the gap back.
This teaching phase is unscored. Once it ends, Atlas evaluates Sam's answers to exam questions mapped one to one to the original misconceptions, and produces a scored result along with a report on which misconceptions were resolved and which remain. Axiom Flow is not purely a formative assessment platform in the way that phrase usually gets used to describe recall checking tools. It combines a formative, no score teaching phase with a summative, scored exam, and the two phases are connected: what a student fails to correct in Sam during teaching becomes visible as a gap in the exam that follows.
Teaching is the assessment, not preparation for it
This is closer to what assessment for learning describes in practice: the assessment activity itself builds and reveals understanding, rather than simply measuring it after the fact. Assessment for learning, applied properly, requires the act of assessment to do more than record a score. Having a student teach a concept and watching what they fail to correct is one of the few assessment structures that does this by design, rather than as an add on to a testing sequence.
Unlike a standard formative assessment platform that monitors whether a student can recall a fact under a time limit, this kind of conceptual mastery assessment asks something harder: can the student rebuild the concept well enough that someone else, or something else, comes away with an accurate model of it. A misconception that survives an exam usually survives because nothing in the exam ever asked the student to teach it back.
Enjoyed reading this? Share this article with your network.


