Why ChatGPT struggles to level the same text consistently
Paste a passage, ask for an easier version, and you get one. Do it for four levels across a term and the cracks show — usually in the direction of everything drifting toward the middle.
8 min readA general chatbot has no fixed reference for what a reading level is, so "simpler" is judged relative to the text in front of it rather than against a standard. Ask separately for three levels and you get three texts that are each simpler than the last, but not reliably at the levels you named — and different on Tuesday than on Monday.
What actually goes wrong
It is not that the output is bad. It usually reads well, which is exactly why the problems are easy to miss.
- Levels drift toward the middle. Ask for A2 and C1 versions and both tend to land nearer B1 than requested. The gap between your top and bottom students is the thing that gets compressed.
- Nothing is anchored. "Simpler" is relative to the input, not to a standard. Two different source texts asked for the same level produce two different difficulties.
- It varies between attempts. The same prompt on two days gives noticeably different results, so a text prepared in September is not comparable with one prepared in March.
- Content quietly changes. Simplifying prose often drops a clause, and with it a fact your comprehension questions depend on.
That last one is the dangerous one, because it fails silently. The lower-level version reads fine and simply no longer contains the answer to question four.
Prompting your way to something usable
If you are doing this by hand — and plenty of teachers are, effectively — a few things measurably improve the output.
- 1Describe the level concretely. "Average sentence length under twelve words, no subordinate clauses, everyday vocabulary" beats "make it A2", because the model has something checkable to aim at.
- 2Generate all versions in one request. Asking for every level together gives it the spread to work against. Separate requests drift independently.
- 3Pin the facts. List what must survive in every version, then check they did.
- 4Write the questions from the lowest version. If the easiest text can answer them, all of them can.
When the manual approach stops paying
For one text, occasionally, hand-prompting is completely reasonable and you should not buy software for it. The arithmetic changes with repetition.
A careful pass — prompting, checking the facts survived, writing matched questions, verifying the lowest version answers them — runs to perhaps forty minutes. Weekly, across a year, that is a working week and a half spent on something that also has to be redone whenever the scheme of work changes.
What changes with a fixed scale
The difference is not that the writing is better. It is that the levels mean something stable.
In Dily a topic is written across ten levels against the same fixed scale every time, with comprehension questions, vocabulary support and a writing prompt matched to each version. A level-four text in September and a level-four text in March are the same difficulty, so the work is comparable across a year — which is what makes progress reporting mean anything.
Bring Dily to your classrooms
Tell us about your school and we'll set up a demo tailored to how your teachers work. Pilot programs available.