Writing assessment

An ESL writing rubric that works

Most writing rubrics were built for native speakers. Here is a four-criteria version made for English language learners, why four is the right number, and how to keep it consistent when six teachers are marking the same year group.

8 min read
Short answer

Score EFL writing on four criteria: Task Achievement, Grammar & Accuracy, Vocabulary, and Cohesion & Coherence. Four is enough to separate what a student meant from how they said it, and few enough that teachers apply it the same way twice. Grade each criterion 1–5, then report the overall score alongside the four parts rather than instead of them.

Why a general writing rubric misfires on EFL work

A rubric written for first-language writers assumes the language is available and asks what the student did with it. For an English learner that assumption breaks. A student can have a sharp argument and sequence it well, and still produce sentences with tense errors throughout. Score that on a single holistic scale and the grammar swamps everything else — the student reads the mark as "my thinking was bad," which is not what happened.

The opposite failure is just as common. A student writes four flawless simple sentences, avoids every structure they are unsure of, and scores well because nothing is wrong. Nothing is wrong because nothing was attempted.

A rubric for language learners has to separate what the student meant from how they managed to say it — and then reward the attempt as well as the accuracy.

The four criteria

These map onto the descriptor families teachers already recognise from CEFR and IELTS, which matters: if your department has ever looked at a band descriptor, this vocabulary will already be familiar.

CriterionWhat it asksWhat a 5 looks like
Task AchievementDid the student do what the prompt asked, at the length asked for?Every part of the prompt is addressed, with enough development that the reader is not left guessing.
Grammar & AccuracyIs the language controlled — and controlled across a range, not just in safe structures?Errors are rare and do not obscure meaning. Complex structures are attempted, not avoided.
VocabularyIs word choice precise, varied, and appropriate to the register?Words are chosen, not reached for. Repetition is deliberate rather than a symptom of a small range.
Cohesion & CoherenceDo the ideas connect, and does the whole piece hold together?Paragraphs follow an order the reader can predict, and linking is varied rather than a chain of "and then".
Writing
4/ 5
Task4
Grammar3
Vocab4
Cohesion5
The same four criteria reported alongside the overall score, so a student can see which part moved.

Turning criteria into bands

A 1–5 scale per criterion is enough. What makes it usable is deciding, as a department, what the middle band means — because that is where most writing lands and where markers drift apart fastest.

  1. 15 — Consistent. The criterion is met throughout. A reader would not stop.
  2. 24 — Mostly consistent. Occasional slips that a reader notices but moves past.
  3. 33 — Uneven. The criterion is met in places and missed in others. This is the anchor band; agree on it first.
  4. 42 — Emerging. Attempted, but the reader has to work to follow.
  5. 51 — Not yet. The criterion is not visible in this piece.

Write one real student sample against each band and keep those samples where markers can see them. A shared example does more for consistency than another paragraph of descriptor wording.

Keeping it consistent across a department

The rubric is not the hard part. Six teachers applying it the same way in March as they did in October is the hard part. Two things move the needle more than anything else:

  • Mark one piece together before the term starts. Everyone scores the same anonymised sample, then compares. Disagreement in the room is cheap; disagreement in the reports is not.
  • Fix the criteria, not just the rubric. If one teacher scores Vocabulary and another folds it into Grammar, the numbers are not comparable no matter how good the descriptors are.

This is also where automated scoring earns its place — not because a model marks better than a good teacher, but because it applies the same four criteria in exactly the same way to every piece, in every class, all year. Teacher judgement then has something stable to argue with.

How this works in Dily

Dily scores every submitted piece of writing on these four criteria, gives an overall 1–5, and writes a short feedback paragraph in plain language. It also highlights specific strong moments in the student's own text, which matters more than it sounds — students who only ever see corrections stop attempting anything difficult.

Alongside the score it produces before-and-after rewrites: the student's original sentence, and a better version of the same idea. The research on written corrective feedback is still unsettled on most questions, but direct correction paired with an explanation consistently outperforms both circling errors and rewriting silently.

Get started

Bring Dily to your classrooms

Tell us about your school and we'll set up a demo tailored to how your teachers work. Pilot programs available.

We usually respond within 1–2 school days.