Back to blog
Product

What Makes a Good Diagnostic Question

Teacher writing a question on a whiteboard

A diagnostic question is not the same as a test question. A test question measures performance on a skill that has already been taught. A diagnostic question measures whether a conceptual prerequisite is reliably in place. These are different goals, and they require different design choices.

Most questions teachers write or select from textbooks are optimized for the test use case: they cover the content of the current unit, they are graduated in difficulty, and they reward students who have learned what was taught. This makes them poor diagnostic instruments for prerequisite assessment. A question designed to assess algebra mastery will not tell you whether the student has a fraction gap, even though the fraction gap might be exactly what is causing the algebra difficulty.

The core design requirement: concept isolation

The most important property of a good diagnostic question is concept isolation: the question should be solvable only by applying the specific concept being probed, without requiring other skills that could confound the result. If a question designed to probe fraction division also requires multi-step arithmetic or reading comprehension of a word problem, then a student who gets it wrong might have failed on the fraction division or might have failed on the arithmetic or the reading. You cannot tell which, and the diagnostic signal is lost.

This means that good diagnostic questions are typically shorter and simpler in their surface presentation than typical test questions. A diagnostic question targeting fraction division might be as simple as: "What is 3/4 divided by 1/2?" But this simple form should not be confused with an easy question. It is specifically designed to be impossible to answer correctly through any means other than the target skill -- there is no context to guess from, no pattern to recognize, no related procedure that accidentally produces the right answer.

Concept isolation also means avoiding questions that can be solved by a procedure without requiring any understanding. "What is 3/4 divided by 1/2?" can be answered by a student who has memorized "invert and multiply" without understanding why the rule works. For a preliminary screen, this is acceptable -- even procedural competence indicates some learning. But for deeper prerequisite assessment, you need question variants that probe whether the understanding is flexible enough to transfer. A student who can compute 3/4 / 1/2 = 3/2 but cannot tell you whether 3/5 divided by something equals 6/5 without setting up the full procedure is showing a different level of mastery than a student who can reason directly about the relationship.

Using distractors to carry diagnostic information

In multiple-choice diagnostic questions, the wrong answers -- distractors -- carry as much diagnostic information as the right answer does, if they are designed carefully. A good diagnostic distractor corresponds to a specific, predictable misconception that reveals something about the nature of the gap.

For fraction division, common distractors include the answer you get by dividing straight through instead of inverting (students who have not fully internalized the invert-and-multiply rule), the answer you get by inverting both fractions (a plausible but incorrect extension of the rule), and the original fraction unchanged (students who are guessing or who confuse division by a fraction with multiplication). If five students get the wrong answer and four of them give the "divided straight through" response, that tells you something different than if those four students each gave different wrong answers. The pattern in distractor choices reveals whether the gap is a specific misconception or just general unfamiliarity.

This is also why the distractor pool for a diagnostic question should be small -- three or four options at most, where each wrong option is specifically chosen to correspond to a real misconception rather than to be a plausible-sounding distractor. The goal is not to make the question harder by adding confusing options; it is to make the wrong choices informative.

Question sequences versus single questions

A single question is almost never sufficient to diagnose a prerequisite gap reliably. A student who answers one question correctly might have guessed. A student who answers one question incorrectly might have made a careless arithmetic error. A reliable diagnostic signal requires a sequence of three to five questions targeting the same concept from different angles.

The sequence should include direct computation, conceptual interpretation, and at least one question that requires applying the concept in a slightly unfamiliar form. A student who gets all three versions of the same underlying concept right is almost certainly not guessing. A student who gets the direct computation right but fails the conceptual interpretation is showing procedural competence without conceptual understanding -- which is a specific and important kind of gap. A student who fails the direct computation is showing a more fundamental problem.

The sequence design also has to manage student time carefully. A five-question diagnostic that takes fifteen minutes is not practical to embed in regular classroom activity. The target is three to five questions that a student can answer in under five minutes -- which means the questions have to be genuinely short. No word problems, no multi-step computations, no diagrams that require interpretation before the target skill can be applied. The simpler the question surface, the faster the student can answer, and the cleaner the diagnostic signal.

What happens when the question design is wrong

Poor diagnostic question design generates misleading signals in two directions. False negatives occur when a student who has a genuine prerequisite gap manages to answer the question correctly through a means other than the target skill -- guessing, applying a memorized procedure to a form they recognize, or using a related skill that happens to produce the right answer in this case. This student will appear gap-free and will not receive the intervention they need.

False positives occur when a student who has the target concept fails the question for a different reason -- they misread a number, they were distracted, they encountered a question format they had not seen before, or the question was ambiguous in a way that made their correct reasoning produce the wrong answer. This student will appear to have a gap and may be unnecessarily pulled aside for remediation.

Both errors are costly. False negatives allow gaps to propagate unaddressed. False positives waste teacher time and, depending on how remediation is delivered, can be demoralizing for students who are actually fine. The goal of careful diagnostic question design is to minimize both error types -- which requires iterating on question items based on data from real students and tracking which questions produce suspicious results (too many right answers on hard concepts, too many wrong answers on concepts students should have solidly).

This is ongoing work, not a one-time design task. A question that worked well as a diagnostic for one cohort of students may show unexpected results with a different cohort if the way the concept was taught differs. Good diagnostic question design is a practice, not a product.

See gap detection in action

Join the early-access program and run a pilot with your classroom or program.