Exam Paper Generator is a Claude AI skill — curriculum-calibrated exam papers with full mark schemes, built for educators inside Claude.
You typed "Generate a Year 10 biology exam on cell division" into Claude. You got ten questions. They were coherent, correctly spelled, and completely useless — asking students to "describe the stages of mitosis" and "explain the importance of cell division" when your class has spent three weeks on meiosis, crossing-over, and genetic variation. The questions tested the right subject in the wrong direction entirely.
This is the most common frustration educators report when using generic AI for assessment writing. The output isn't wrong in an obvious way — it's wrong in the way that only someone who teaches the class would notice. Claude doesn't know which chapters you covered. It doesn't know that your Year 10s are sitting a higher-tier paper, or that you've already assessed recall-level questions this term and need application questions now. So it defaults to the middle: broad, safe, textbook-adjacent questions that could have come from any teacher at any school.
What happens when the tool asks those questions before generating the paper — and calibrates every question to the answers?
Why Generic Exam Questions Feel Wrong to Every Teacher
The gap isn't quality — it's alignment. A generic AI-written question about photosynthesis isn't badly written. It's just written for a hypothetical student rather than your actual class. It doesn't know whether your students have done limiting factors yet, whether you're assessing at GCSE or A-Level, or whether you want two-mark recall questions or six-mark extended answers. So it produces questions that are fine in the abstract and misaligned in practice.
Teachers who try to fix this spend longer editing than they would have spent writing. They delete four questions, rewrite three, adjust the mark scheme, add the command words their exam board requires, and wonder why they bothered asking AI at all. The output became a starting point — a rough first draft that needed as much thought as the original task.
Generic AI generates exam questions for a theoretical version of your subject — not for the specific topic, level, and cognitive demands of the paper you actually need to set.
The deeper issue is that exam design isn't just about correct content — it's about cognitive demand. Bloom's taxonomy exists because "describe," "explain," and "evaluate" are not interchangeable. A question that asks students to describe a process tests entirely different knowledge than one asking them to evaluate its limitations. Generic AI doesn't make this distinction unless you push it hard, and even then it tends to cluster questions at the same cognitive level rather than distributing them across the range your assessment needs.
That gap is exactly what the Exam Paper Generator skill for Claude was built to close.
What Context Calibration Changes in Assessment Writing
The Exam Paper Generator skill doesn't open with a blank generation. Before writing a single question, it asks you three things: the subject and specific topic, the year group and tier, and the cognitive demand distribution you need — recall, application, analysis, evaluation. Those three answers are the difference between a paper that fits your class and one that could have been downloaded from any revision website.
This mirrors how experienced teachers actually think about assessment design. You don't sit down and generate questions randomly — you map the topic first, decide how many marks to weight toward higher-order thinking, then write to those constraints. The skill replicates that process. It understands that a GCSE Chemistry paper on rates of reaction needs different questions than an A-Level one on the same topic — not because the chemistry is different, but because the expected depth of understanding, the required command words, and the mark allocation logic are completely different.
An exam paper only works when every question is written for your students' specific point in the curriculum — not for a generalised version of the subject.
What that calibration produces concretely: questions that use the right command words for your exam board's expectations, a mark scheme with credit points that match the cognitive level of each question, and a paper structure that distributes marks across the topic rather than clustering everything in the obvious areas. You get something you can hand to a colleague for review, not something you need to rebuild from scratch.
What the Exam Paper Generator Skill Actually Does
The skill runs a defined sequence — not a single large request, but a structured process that earns the right to write by understanding your context first.
Generic Questions vs Calibrated Questions
The difference between these two outputs is the difference between a paper you have to edit and one you can use. Both are on the same topic. Only one is on your topic.
Q2. Explain why cell division is important for living organisms. [3 marks]
Q3. What is the difference between mitosis and meiosis? [4 marks]
Mark scheme: Award 1 mark per correct stage named. Award marks for correct explanation of growth and repair. Award marks for describing differences in chromosome number.
Q2. A student claims that crossing over during meiosis always increases genetic variation. Evaluate this statement with reference to the process of recombination. [6 marks]
Mark scheme: Q1 — Award 1 mark each for: formation of bivalents (synapsis); crossing over / chiasmata formation. Q2 — Level 3 (5–6 marks): clear and detailed explanation that crossing over shuffles allele combinations within chromosomes and that variation depends on where chiasmata form; addresses the "always" claim with reference to identical crossover positions as a limiting case...
The left column produced a passable set of questions that any student who's opened a revision guide could answer. The right column is built for a class that has already covered the stages, understands the basic differences, and is now being assessed at the analytical level the topic demands. The mark scheme on the right uses level-descriptor language and credit-point logic — it's usable by any teacher marking the paper, not just the one who wrote it. That difference comes entirely from the calibration step.
Who Gets the Most from the Exam Paper Generator
Secondary and post-secondary teachers who write their own assessments. Course creators and online educators building module tests. Private tutors who need custom papers for individual students rather than generic past-paper questions.
The biggest gains come for teachers writing assessments across multiple year groups or classes — the calibration step means one skill run can produce a GCSE Foundation paper and a separate Higher paper on the same topic, properly differentiated rather than just shortened. For tutors, the value is even more specific: a student preparing for a resit needs questions that target the exact gaps in their understanding, not a generic paper that re-covers everything equally. The skill produces what you spec, not what it assumes you need.
Online course creators benefit in a slightly different way. Their assessments need to test module-specific content without leaning on institutional exam board conventions — they can specify the cognitive level distribution that matches their course structure, and the skill builds to that rather than defaulting to a traditional exam format their students didn't sign up for.
The Output You Walk Away With
A complete exam paper: questions numbered and formatted, mark allocation stated per question, command words consistent with your level and specification, and a full mark scheme with credit points specified per question. For higher-order questions, the mark scheme includes level-descriptor guidance so any teacher can apply it consistently. The paper is structured to your stated cognitive demand distribution — if you asked for 30% recall, 40% application, 30% evaluation, that's what the question mix reflects.
The path from download to usable paper is short. Run the skill in your Claude session, answer the three calibration questions, and the output arrives formatted and complete. Copy it into your preferred document tool, adjust any formatting specifics for your school's template, and it's ready. The editing stage isn't "rewrite half the questions" — it's "check this against my lesson notes and add any topic-specific context the skill didn't have access to."
Most teachers expect to spend the first five minutes after an AI-generated paper crossing out what doesn't fit. The calibration step is specifically designed to eliminate that. The skill doesn't know your students — but it does know your topic, your level, and the cognitive demand you asked for. That's usually enough to make the output usable rather than merely salvageable.
The next piece most people tackle from here is a personal statement that reads specific, not generic. If you're working across the full Educator workflow, the Educator bundle covers everything in one place.
Put this to work: the Exam Paper Generator skill for Claude turns everything above into one guided workflow you run in a normal Claude chat. Not ready to buy? Start with a free Claude skill and see how it works first.
Related reading: Why AI-Written Exam Questions Are So Obvious — and How to Fix It