Practical AI

Why ChatGPT struggles to level the same text consistently

Paste a passage, ask for an easier version, and you get one. Do it for four levels across a term and the cracks show — usually in the direction of everything drifting toward the middle.

8 min read
Short answer

A general chatbot has no fixed reference for what a reading level is, so "simpler" is judged relative to the text in front of it rather than against a standard. Ask separately for three levels and you get three texts that are each simpler than the last, but not reliably at the levels you named — and different on Tuesday than on Monday.

What actually goes wrong

It is not that the output is bad. It usually reads well, which is exactly why the problems are easy to miss.

  • Levels drift toward the middle. Ask for A2 and C1 versions and both tend to land nearer B1 than requested. The gap between your top and bottom students is the thing that gets compressed.
  • Nothing is anchored. "Simpler" is relative to the input, not to a standard. Two different source texts asked for the same level produce two different difficulties.
  • It varies between attempts. The same prompt on two days gives noticeably different results, so a text prepared in September is not comparable with one prepared in March.
  • Content quietly changes. Simplifying prose often drops a clause, and with it a fact your comprehension questions depend on.

That last one is the dangerous one, because it fails silently. The lower-level version reads fine and simply no longer contains the answer to question four.

Prompting your way to something usable

If you are doing this by hand — and plenty of teachers are, effectively — a few things measurably improve the output.

  1. 1Describe the level concretely. "Average sentence length under twelve words, no subordinate clauses, everyday vocabulary" beats "make it A2", because the model has something checkable to aim at.
  2. 2Generate all versions in one request. Asking for every level together gives it the spread to work against. Separate requests drift independently.
  3. 3Pin the facts. List what must survive in every version, then check they did.
  4. 4Write the questions from the lowest version. If the easiest text can answer them, all of them can.
Better"Write this at four levels, from simplest to hardest, in one response."
UnreliableFour separate chats asking for "a bit simpler" each time.
BetterNaming the sentence length and vocabulary you want.
UnreliableNaming a level and assuming a shared definition.
BetterChecking every version answers every question.
UnreliableAssuming simplification preserved the content.
A fixed scale, every time
A1A2B1B2C1C2
Levels produced together against a fixed scale, rather than each one rewritten relative to the last.

When the manual approach stops paying

For one text, occasionally, hand-prompting is completely reasonable and you should not buy software for it. The arithmetic changes with repetition.

A careful pass — prompting, checking the facts survived, writing matched questions, verifying the lowest version answers them — runs to perhaps forty minutes. Weekly, across a year, that is a working week and a half spent on something that also has to be redone whenever the scheme of work changes.

~40 minA careful hand-levelled text with checked questions
WeeklyHow often a scheme of work needs a new one
Every timeHow often the checking has to be repeated

What changes with a fixed scale

The difference is not that the writing is better. It is that the levels mean something stable.

In Dily a topic is written across ten levels against the same fixed scale every time, with comprehension questions, vocabulary support and a writing prompt matched to each version. A level-four text in September and a level-four text in March are the same difficulty, so the work is comparable across a year — which is what makes progress reporting mean anything.

Get started

Bring Dily to your classrooms

Tell us about your school and we'll set up a demo tailored to how your teachers work. Pilot programs available.

We usually respond within 1–2 school days.