Claude is better for assessment writing. ChatGPT is better for quick conversational drafts. Those are two different workflows, and the distinction matters more than the question of which model is "smarter." Most teachers want one thing from AI: material they can use in class this week without spending as long fixing it as writing it from scratch. Which model gets you there depends on the task.
For exam papers, mark schemes, and structured assessment: Claude. For rapid first-draft lesson plans that you'll refine through back-and-forth dialogue: ChatGPT is marginally faster to iterate with. For anything that needs to hold tight curriculum alignment across a 90-minute paper: Claude by a significant margin.
Exam Papers: Where the Gap Is Clearest
The difference between a usable exam question and a wasted one is precision in three areas: the command term (describe, explain, analyse, evaluate — each has a specific scope), the mark allocation (a 6-mark question needs a different level of response than a 2-mark question), and the curriculum alignment (which assessment objective the question tests, and at which cognitive level).
ChatGPT produces exam questions that look correct but frequently fail the precision test. Command terms get used loosely — "discuss" where the mark scheme expects "analyse." Mark allocation doesn't match the expected depth of response. Questions test surface recall when the curriculum objective specifies application or synthesis. These errors are invisible to a glance and only surface when you compare the output against the actual assessment criteria.
Claude, given a structured brief that specifies the board, subject, year group, and command term requirements, produces questions that stay aligned with those constraints across a full paper. It is not perfect — subject-specific factual errors still occur, particularly in sciences — but the structural alignment is more reliable.
| Assessment task | ChatGPT | Claude |
|---|---|---|
| Command term precision | Uses command terms loosely; "discuss" and "analyse" treated interchangeably | Distinguishes command terms correctly when brief specifies the board's taxonomy |
| Mark allocation calibration | Question depth rarely matches the mark allocation without explicit prompting | Calibrates expected response depth to mark allocation more reliably |
| Multi-part question structure | Parts (a)(b)(c) often repeat similar cognitive demand rather than scaffolding | Structures multi-part questions with increasing cognitive demand when instructed |
| Mark scheme generation | Produces mark scheme with broad acceptable answers; misses specific point-scoring criteria | Produces more specific mark points when the assessment objective is named |
| Worked solutions | Correct in most cases for standard question types; errors appear in multi-step problems | More reliable on multi-step problems; shows working more consistently |
Lesson Plans: Where ChatGPT Catches Up
Lesson plan writing is a different task from assessment writing. The quality bar is lower in one specific way: a lesson plan is a teacher's working document, not a student-facing one. A lesson plan with slightly vague phrasing gets refined in the classroom. An exam question with vague phrasing creates a marking dispute.
ChatGPT's conversational iteration makes it useful for lesson planning. A teacher can describe a lesson concept, get a draft structure, then ask for the introduction activity to be shorter and the group task to involve more peer explanation — and get a revised plan in 30 seconds. The back-and-forth refining is natural.
Claude does the same, and often produces better initial learning objectives and clearer success criteria. But if a teacher's workflow is "quick draft, then talk it into shape," ChatGPT's conversational feel has a slight edge in comfort.
The Brief Problem — Neither Model Knows Your Curriculum
"Both models produce generic exam questions when given generic inputs. 'Write ten questions on photosynthesis for Year 10' is not a brief — it's a topic. The board, the cognitive level, and the mark allocation are the brief."
The most common mistake teachers make with AI assessment tools is treating a topic as a brief. "Write ten questions on photosynthesis for Year 10" tells the model almost nothing useful: which board (AQA, OCR, Edexcel, Cambridge), which specification point, which command terms are required, what the difficulty distribution should be (recall, application, analysis), and what mark weight each question carries.
Without this information, both Claude and ChatGPT produce generic questions that could appear in a revision worksheet but not in a board-aligned assessment. With it, Claude in particular produces questions that require minimal editing before use.
Board and specification: AQA GCSE Biology, OCR A-Level Chemistry, Cambridge IGCSE Maths — each has different command term conventions and mark allocation norms. Year group and tier: Foundation vs Higher changes both the question style and the acceptable response level. Topic weighting: which sub-topics to include and in what proportion. Difficulty distribution: percentage of recall, application, and analysis questions. Total marks: so the model can calibrate individual question weight correctly.
A skill that runs this intake systematically changes the output from "generic questions on the topic" to "questions aligned with this paper's structure." For teachers producing assessment across multiple classes and year groups, the brief consistency is more valuable than the model capability gap between Claude and ChatGPT.
What AI Does Well for Teachers (and What It Doesn't)
AI is genuinely useful for three teaching tasks: generating question variants at different difficulty levels from a master question, drafting lesson structure so teachers spend their time on subject knowledge rather than format, and producing worked solutions for assessment tasks faster than writing them from scratch.
AI is unreliable for two tasks that look similar but aren't: verifying that subject content is factually correct (especially in sciences and mathematics where errors in AI output are invisible until a student or colleague catches them), and calibrating difficulty to a specific cohort without knowing how that cohort has been taught.
The practical rule: use AI to produce the structure and the question wording, then verify the subject content yourself. The time saving is in the structural work — that's where AI is reliable. Subject fact-checking is still the teacher's job.
For lesson planning at scale, the Lesson Plan Builder skill runs the same structured intake — learning objective, year group, prior knowledge, time allocation — before producing a full plan with starter, main activities, and exit assessment built in.