Exam Paper Generator is a Claude AI skill — curriculum-calibrated exam papers with full mark schemes, built for educators inside Claude.
Ask any teacher who's tried to use AI to write an exam and they'll describe the same experience. The output looked like an exam. The formatting was fine. And something about every single question felt slightly off — too broad, too easy, pitched at the wrong level, testing things students covered three months ago rather than what was taught last week. They spent more time editing it than if they'd written it themselves. And they couldn't quite explain why.
The reason is specific and fixable. When you type "write a geography exam on coastal processes for Year 9," Claude does exactly what you asked. It writes a geography exam on coastal processes. But "coastal processes for Year 9" means something completely different in a school where students have spent five weeks on erosion, transportation, and deposition than in one where they've just been introduced to the topic. Claude doesn't know which school you're in. It doesn't know which aspects you've covered, which you've skipped, or what your students were tested on last month. So it writes for the average — and the average is never quite right.
That subtle wrongness isn't a quality problem. It's a context problem. The distinction matters, because one of these is fixable.
The Three Ways Generic AI Misses the Mark
Generic AI exam output tends to fail in the same three ways, regardless of subject. Identifying them makes the fix obvious.
The topic scope problem. "Waves and coastal erosion" is a topic. But within that topic, a teacher might have focused on discordant and concordant coastlines, or headlands and bays, or hydraulic action versus abrasion versus solution. Claude doesn't know which. So it covers everything at surface level rather than going deep on what students actually learned. The result is a paper that tests breadth when your class needs a test of depth.
The cognitive level problem. An exam question that asks students to "describe" a process is testing recall. One that asks them to "explain" is testing comprehension. One that asks them to "evaluate" or "assess" is testing higher-order thinking. A well-designed paper distributes questions across this range intentionally — you might want 40% recall, 40% application, and 20% evaluation, because that's the balance your syllabus requires. Generic AI almost always clusters in the middle, producing a paper dominated by "explain" questions because that's the safest territory. The distribution never matches what you actually need.
The command word problem. Different exam boards use command words differently. "Outline" means one thing in AQA Geography and something slightly different in OCR. "Discuss" carries specific mark scheme expectations that vary by specification. Generic AI uses command words based on general conventions, not your exam board's specific requirements. Teachers notice this immediately. Students who've been trained to respond to their board's command words will read the paper differently than you intended.
Every AI exam paper failure traces back to the same root cause: the tool was never told what the teacher actually needs. It filled in the blanks with assumptions — and assumptions produce generic output.
None of these failure modes require a better AI to fix. They require the AI to ask three questions before it starts writing.
That gap is exactly what the Exam Paper Generator skill for Claude was built to close.
What Asking First Actually Changes
The Exam Paper Generator skill opens with calibration, not generation. It asks about the specific topic and subtopics taught, the year group and qualification tier, and the cognitive demand distribution required. These aren't optional enrichment inputs — they're the information the skill needs before it can write a single question that's actually aligned to your class.
The difference this makes is structural. When Claude knows you're writing a Higher GCSE paper for a class that has covered the full coastal processes unit including discordant coastlines, cliff profiles, and coastal management strategies, it's no longer guessing at scope. When it knows you need 30% recall, 40% application, and 30% evaluation questions, it builds the paper to that distribution rather than defaulting to comfortable middle ground. When it knows your exam board, the command words it uses match your mark scheme conventions rather than a generic interpretation.
The gap between a paper you can use and one you have to rewrite is almost always a missing conversation before the generation started.
This also changes what the mark scheme looks like. A mark scheme built from calibrated questions uses the same cognitive level language as the questions themselves. A six-mark evaluation question gets a level-descriptor mark scheme, not a list of bullet points. A two-mark recall question gets a simple credit-point scheme. The mark scheme structure matches the question structure because both were designed from the same starting parameters.
The Questions That Reveal the Difference
The clearest way to see what calibration changes is to look at what it produces for the same subject at two different levels. Both papers below are on respiration for a science class. Only one is on the right respiration for the right class.
Q2. Write the word equation for aerobic respiration. [2 marks]
Q3. Explain the difference between aerobic and anaerobic respiration. [4 marks]
Q4. Describe what happens during anaerobic respiration in muscle cells. [3 marks]
All questions cluster at recall and basic comprehension. No higher-order questions. No reference to whether this class has covered ATP synthesis, oxygen debt, or fermentation.
Q2. A student exercises at high intensity for 90 seconds. Explain why their muscles switch to anaerobic respiration and describe how oxygen debt is resolved after exercise. [6 marks]
Q3. Evaluate the efficiency of aerobic versus anaerobic respiration with reference to ATP yield per glucose molecule. [5 marks]
Mark scheme: Q3 — Level 3 (4–5 marks): clear comparative analysis of ATP yield (approximately 38 ATP aerobic vs 2 ATP anaerobic); addresses efficiency in terms of incomplete oxidation in anaerobic pathway; considers contexts where anaerobic respiration is necessary despite lower yield.
The left paper would frustrate a Year 10 Higher class that's spent three weeks on ATP synthesis and fermentation. Every question tests content from the first two lessons of the unit. The right paper builds directly on what this class has actually covered — it tests application and evaluation of the concepts they've spent the most time on, with a mark scheme that reflects the analytical depth required. The difference isn't AI quality. It's context.
What the Quality Gate Catches Before You See the Output
Beyond the calibration questions, the Exam Paper Generator runs an automatic quality check on the output before it reaches you. This catches the failure modes that context alone doesn't fix: questions where the command word and the mark allocation are mismatched, mark scheme entries that don't map cleanly to the question asked, or cognitive level distributions that drifted from what you specified. You don't see the draft that had a six-mark "state" question or a mark scheme that awarded points for content outside the specified topic — those get caught and corrected before delivery.
This matters because even well-calibrated questions can have structural problems that only show up at the mark scheme stage. A question that asks students to "analyse" coastal erosion processes but then provides a two-mark mark scheme based on naming processes isn't coherent — the cognitive demand of the question and the mark scheme logic are at different levels. The quality check aligns these before you see the output, so you're not discovering the mismatch when you hand the paper to a colleague for review.
What a Good AI Exam Paper Actually Requires
The comparison below isn't about hard versus easy questions. It's about what makes a paper professionally usable versus what makes it a starting point that still needs significant work.
| Generic AI output | Calibrated skill output |
|---|---|
| Questions cover the subject broadly | Questions cover the specific subtopics taught |
| Cognitive levels cluster mid-range | Cognitive demand distributed to your specification |
| Command words follow general conventions | Command words match your exam board's usage |
| Mark scheme uses bullet points throughout | Mark scheme format matches question type (level descriptors or credit points) |
| Requires editing before use | Requires review before use — not rebuilding |
The practical difference between "requires editing" and "requires review" is about 45 minutes of a teacher's time. Editing means rewriting questions, adjusting the mark scheme, removing things that don't fit. Review means reading through what's there, adding any topic-specific nuance the skill didn't have access to, and formatting it for your school's template. Both involve professional judgment — but only one of them produces a paper the teacher could hand to a colleague as-is.
Teachers who write exams understand this distinction immediately. The ones who've tried generic AI and dismissed it are usually describing the editing experience, not the review experience. Those are different products, and the difference is in the calibration step — not the underlying capability.
The next piece most people tackle from here is a personal statement that reads specific, not generic. If you're working across the full Educator workflow, the Educator bundle covers everything in one place.
Put this to work: the Exam Paper Generator skill for Claude turns everything above into one guided workflow you run in a normal Claude chat. Not ready to buy? Start with a free Claude skill and see how it works first.
Related reading: The Exam Paper That Matches Your Curriculum — Not Claude's Assumptions