Question quality lab

How to write better distractors for AI-generated multiple-choice questions

A failure taxonomy, worked example and repeatable review test for plausible wrong answers that measure understanding.

Published
24 August 2026
Reading time
11 min
Evidence
Distractor failure taxonomy and QA test
Author
The AI Quiz Lab

A good distractor is a plausible wrong answer tied to a real misconception. It should attract someone who has not mastered the idea, while remaining clearly incorrect under the stated source and question wording.

Key findings, 24 August 2026

  • Fluent wrong answers are not necessarily diagnostic wrong answers.
  • Grammar clues, unequal length and joke options make questions easier without measuring knowledge.
  • Each distractor should have a documented reason why a learner might choose it and a clear reason why it is wrong.

What is a distractor?

A distractor is an incorrect option in a selected-response question. Its job is to distinguish partial understanding from mastery, not merely to fill space. The University of Waterloo recommends plausible alternatives, one clearly best answer, parallel grammar and avoiding clues such as absolute wording. See its institutional guide to designing multiple-choice questions.

Recent evidence reinforces the need for review. A 2026 peer-reviewed comparison reported a higher share of nonfunctioning AI-generated distractors than human-written distractors in its sample. A separate 2026 study with 80 students found human distractors were rated highest and AI-generated alternatives ranked second. These studies are context-specific, so they support testing rather than a universal failure rate. See the SAGE study and the AAAI study.

What are the six common distractor failures?

FailureWhy it failsRepair
Unrelated optionNo informed respondent would select itUse a nearby misconception
Grammar mismatchOnly the key completes the stemMake every option grammatically parallel
Length clueThe detailed option looks correctKeep comparable detail and structure
OverlapTwo options can both be trueMake options mutually exclusive
Absolute language“Always” or “never” signals a trapUse source-appropriate qualifiers
Invented misconceptionThe wrong answer has no diagnostic valueName the likely reasoning error

How do you prompt for better distractors?

Reusable distractor specification

For each multiple-choice item, write one best answer and three plausible distractors. Base each distractor on a distinct misconception, reversed relationship, boundary error or misapplied rule supported by the source. Keep all options parallel in grammar, length and specificity. After each distractor, privately label the misconception and cite the evidence that makes it wrong.

Worked example

Concept: Correlation does not by itself establish causation.

Question: A study finds that two variables increase together. What can be concluded from that result alone?

  • Best answer: The variables are associated, but causation is not established.
  • Distractor 1: The first variable caused the second. Misconception: temporal or causal leap.
  • Distractor 2: The second variable caused the first. Misconception: reversed direction.
  • Distractor 3: No relationship exists because other factors may be involved. Misconception: confounding erases observed association.

A weak fourth option such as “The study is about weather” would be unrelated and non-diagnostic.

What is the repeatable QA test?

  1. Hide the answer key and ask a reviewer to identify the best answer.
  2. Require the reviewer to explain the misconception behind every wrong option.
  3. Check options for grammar, length and specificity clues.
  4. Verify that the source makes each distractor wrong.
  5. After real use, inspect option selections. A never-selected distractor may be too implausible, though a small sample is inconclusive.

When should a multiple-choice item be replaced?

Replace it when the concept has no plausible alternatives, two answers remain defensible after editing, or the only way to make distractors harder is trick wording. A short-answer or application task may measure the objective more honestly.

Use this alongside the question quality checklist, the six-part prompt specification and the AI Quiz Lab methodology.

Sources and verification

This guide was verified on 24 August 2026. Facts are linked to primary or institutional sources. The worked models and checklists are original AI Quiz Lab evidence assets, designed to be repeated with any generator.

If a source changes or you find an error, use our correction path. The publication-neutral next step is to test the framework on a small, low-stakes quiz before using it at scale.