AI for Work 6 min read

Why AI-Written Exam Questions Are So Obvious — and How to Fix It

The problem isn't that AI can't write exam questions. It can. The problem is that it's answering a question you didn't ask — and the mismatch shows up in every single paper it produces.

SP
Founder, NovaKit
📝
NovaKit Skill
Exam Paper Generator — curriculum-calibrated exam papers with full mark schemes, built for educators inside Claude.
Quick answer: The problem isn't that AI can't write exam questions. It can. The problem is that it's answering a question you didn't ask — and the mismatch shows up in every single paper it produces.
In this guide

Exam Paper Generator is a Claude AI skill — curriculum-calibrated exam papers with full mark schemes, built for educators inside Claude.

  1. The Three Ways Generic AI Misses the Mark
  2. What Asking First Actually Changes
  3. The Questions That Reveal the Difference
  4. What the Quality Gate Catches Before You See the Output
  5. What a Good AI Exam Paper Actually Requires

Ask any teacher who's tried to use AI to write an exam and they'll describe the same experience. The output looked like an exam. The formatting was fine. And something about every single question felt slightly off — too broad, too easy, pitched at the wrong level, testing things students covered three months ago rather than what was taught last week. They spent more time editing it than if they'd written it themselves. And they couldn't quite explain why.

The reason is specific and fixable. When you type "write a geography exam on coastal processes for Year 9," Claude does exactly what you asked. It writes a geography exam on coastal processes. But "coastal processes for Year 9" means something completely different in a school where students have spent five weeks on erosion, transportation, and deposition than in one where they've just been introduced to the topic. Claude doesn't know which school you're in. It doesn't know which aspects you've covered, which you've skipped, or what your students were tested on last month. So it writes for the average — and the average is never quite right.

That subtle wrongness isn't a quality problem. It's a context problem. The distinction matters, because one of these is fixable.

The Three Ways Generic AI Misses the Mark

Generic AI exam output tends to fail in the same three ways, regardless of subject. Identifying them makes the fix obvious.

The topic scope problem. "Waves and coastal erosion" is a topic. But within that topic, a teacher might have focused on discordant and concordant coastlines, or headlands and bays, or hydraulic action versus abrasion versus solution. Claude doesn't know which. So it covers everything at surface level rather than going deep on what students actually learned. The result is a paper that tests breadth when your class needs a test of depth.

The cognitive level problem. An exam question that asks students to "describe" a process is testing recall. One that asks them to "explain" is testing comprehension. One that asks them to "evaluate" or "assess" is testing higher-order thinking. A well-designed paper distributes questions across this range intentionally — you might want 40% recall, 40% application, and 20% evaluation, because that's the balance your syllabus requires. Generic AI almost always clusters in the middle, producing a paper dominated by "explain" questions because that's the safest territory. The distribution never matches what you actually need.

The command word problem. Different exam boards use command words differently. "Outline" means one thing in AQA Geography and something slightly different in OCR. "Discuss" carries specific mark scheme expectations that vary by specification. Generic AI uses command words based on general conventions, not your exam board's specific requirements. Teachers notice this immediately. Students who've been trained to respond to their board's command words will read the paper differently than you intended.

💡
The core problem

Every AI exam paper failure traces back to the same root cause: the tool was never told what the teacher actually needs. It filled in the blanks with assumptions — and assumptions produce generic output.

None of these failure modes require a better AI to fix. They require the AI to ask three questions before it starts writing.

That gap is exactly what the Exam Paper Generator skill for Claude was built to close.

What Asking First Actually Changes

The Exam Paper Generator skill opens with calibration, not generation. It asks about the specific topic and subtopics taught, the year group and qualification tier, and the cognitive demand distribution required. These aren't optional enrichment inputs — they're the information the skill needs before it can write a single question that's actually aligned to your class.

The difference this makes is structural. When Claude knows you're writing a Higher GCSE paper for a class that has covered the full coastal processes unit including discordant coastlines, cliff profiles, and coastal management strategies, it's no longer guessing at scope. When it knows you need 30% recall, 40% application, and 30% evaluation questions, it builds the paper to that distribution rather than defaulting to comfortable middle ground. When it knows your exam board, the command words it uses match your mark scheme conventions rather than a generic interpretation.

The gap between a paper you can use and one you have to rewrite is almost always a missing conversation before the generation started.

This also changes what the mark scheme looks like. A mark scheme built from calibrated questions uses the same cognitive level language as the questions themselves. A six-mark evaluation question gets a level-descriptor mark scheme, not a list of bullet points. A two-mark recall question gets a simple credit-point scheme. The mark scheme structure matches the question structure because both were designed from the same starting parameters.

The Questions That Reveal the Difference

The clearest way to see what calibration changes is to look at what it produces for the same subject at two different levels. Both papers below are on respiration for a science class. Only one is on the right respiration for the right class.

Generic AI — Year 10 Biology, Respiration
Q1. What is respiration? [2 marks]

Q2. Write the word equation for aerobic respiration. [2 marks]

Q3. Explain the difference between aerobic and anaerobic respiration. [4 marks]

Q4. Describe what happens during anaerobic respiration in muscle cells. [3 marks]

All questions cluster at recall and basic comprehension. No higher-order questions. No reference to whether this class has covered ATP synthesis, oxygen debt, or fermentation.
✓ Calibrated — Year 10 Higher, post-ATP and fermentation unit
Q1. State the products of anaerobic respiration in yeast cells. [2 marks]

Q2. A student exercises at high intensity for 90 seconds. Explain why their muscles switch to anaerobic respiration and describe how oxygen debt is resolved after exercise. [6 marks]

Q3. Evaluate the efficiency of aerobic versus anaerobic respiration with reference to ATP yield per glucose molecule. [5 marks]

Mark scheme: Q3 — Level 3 (4–5 marks): clear comparative analysis of ATP yield (approximately 38 ATP aerobic vs 2 ATP anaerobic); addresses efficiency in terms of incomplete oxidation in anaerobic pathway; considers contexts where anaerobic respiration is necessary despite lower yield.

The left paper would frustrate a Year 10 Higher class that's spent three weeks on ATP synthesis and fermentation. Every question tests content from the first two lessons of the unit. The right paper builds directly on what this class has actually covered — it tests application and evaluation of the concepts they've spent the most time on, with a mark scheme that reflects the analytical depth required. The difference isn't AI quality. It's context.


What the Quality Gate Catches Before You See the Output

Beyond the calibration questions, the Exam Paper Generator runs an automatic quality check on the output before it reaches you. This catches the failure modes that context alone doesn't fix: questions where the command word and the mark allocation are mismatched, mark scheme entries that don't map cleanly to the question asked, or cognitive level distributions that drifted from what you specified. You don't see the draft that had a six-mark "state" question or a mark scheme that awarded points for content outside the specified topic — those get caught and corrected before delivery.

This matters because even well-calibrated questions can have structural problems that only show up at the mark scheme stage. A question that asks students to "analyse" coastal erosion processes but then provides a two-mark mark scheme based on naming processes isn't coherent — the cognitive demand of the question and the mark scheme logic are at different levels. The quality check aligns these before you see the output, so you're not discovering the mismatch when you hand the paper to a colleague for review.

NovaKit Skill
Exam Paper Generator — asks before it writes
Works inside Claude. No technical setup. Full exam paper with complete mark scheme, calibrated to your topic, level, and cognitive demand requirements before a single question is written.
See the skill from $9 · instant download

What a Good AI Exam Paper Actually Requires

The comparison below isn't about hard versus easy questions. It's about what makes a paper professionally usable versus what makes it a starting point that still needs significant work.

Generic AI output Calibrated skill output
Questions cover the subject broadlyQuestions cover the specific subtopics taught
Cognitive levels cluster mid-rangeCognitive demand distributed to your specification
Command words follow general conventionsCommand words match your exam board's usage
Mark scheme uses bullet points throughoutMark scheme format matches question type (level descriptors or credit points)
Requires editing before useRequires review before use — not rebuilding

The practical difference between "requires editing" and "requires review" is about 45 minutes of a teacher's time. Editing means rewriting questions, adjusting the mark scheme, removing things that don't fit. Review means reading through what's there, adding any topic-specific nuance the skill didn't have access to, and formatting it for your school's template. Both involve professional judgment — but only one of them produces a paper the teacher could hand to a colleague as-is.

Teachers who write exams understand this distinction immediately. The ones who've tried generic AI and dismissed it are usually describing the editing experience, not the review experience. Those are different products, and the difference is in the calibration step — not the underlying capability.

The next piece most people tackle from here is a personal statement that reads specific, not generic. If you're working across the full Educator workflow, the Educator bundle covers everything in one place.

Ready to try it?
Exam Paper Generator for Claude
Complete, curriculum-calibrated exam paper with full mark scheme. Asks the right questions before writing a single one. Works with your free Claude account.
Get the skill $9 · instant download · 7-day refund

Put this to work: the Exam Paper Generator skill for Claude turns everything above into one guided workflow you run in a normal Claude chat. Not ready to buy? Start with a free Claude skill and see how it works first.

Related reading: The Exam Paper That Matches Your Curriculum — Not Claude's Assumptions

Tags Exam Paper Generator Claude AI AI for Educators Assessment Writing Claude Skills
Free skill
Try NovaKit before
you spend a dollar.

Get the LinkedIn Post Engine free — the same skill that runs live trend research before every post. Drop your email and it lands in your inbox in seconds.

💼
LinkedIn Post Engine
Social · normally $9 · free today
Live trend research before every post
Hook variants calibrated to what's converting this week
Works on a free Claude account

No spam. No account. Unsubscribe any time.