Formative Assessment and Feedback Protocols in English Language Teaching
Formative assessment (assessment for learning) is the ongoing gathering of evidence about learning that is used to adjust teaching and to give students actionable, criterion-referenced feedback. Unlike summative assessment, its purpose is to improve learning while it is happening rather than to certify a final result.
What this guide covers
- The difference between formative and summative assessment in ELT
- How to structure written and spoken feedback so it drives learning
- Building and calibrating criterion-referenced rubrics
- Measuring and improving inter-rater reliability (ICC)
Out of scope
- High-stakes exam scoring rules (see the IELTS and Cambridge pages)
- Statistical derivation of ICC coefficients (conceptual treatment only)
Scope of this guide
This page covers the design of formative assessment and feedback systems in English language teaching — how to gather evidence of learning, feed it back to students, and keep marking consistent across teachers. It draws on assessment-for-learning research and on standard measurement concepts such as inter-rater reliability.
Formative vs summative assessment
The distinction is one of purpose and timing, not of instrument. The same speaking task can be formative (used to guide the next lesson) or summative (used to report an end-of-course grade). What matters is how the evidence is used.
| Feature | Formative (for learning) | Summative (of learning) |
|---|---|---|
| Primary purpose | Improve learning in progress | Certify/report achievement |
| Timing | Continuous, during instruction | End of unit, term or course |
| Feedback focus | Actionable next steps, descriptive | Grade or score, evaluative |
| Reference | Criterion-referenced against learning goals | Often norm- or criterion-referenced |
| Student role | Active: self-assessment, goal-setting | Largely passive recipient |
| Typical tools | Drafts, exit tickets, conferencing, rubrics | Final exams, proficiency tests |
Why descriptive feedback beats grades alone
A consistent finding in assessment research is that when a grade and a comment appear together, students attend to the grade and often ignore the comment. Feedback that is specific, criterion-referenced and forward-looking — naming what was achieved and the single most useful next step — supports self-efficacy and self-regulation more effectively than a numerical mark in isolation.
A reliable feedback pattern: (1) name one strength against a criterion, (2) name one priority improvement, (3) give a concrete example or model, (4) set a checkable next action. Keep grades separate from formative comments where possible.
Criterion-referenced rubrics and rater reliability
CEFR-aligned, criterion-referenced rubrics make expectations explicit and improve consistency between markers. Consistency is measured with inter-rater reliability. For continuous rubric scores, the intraclass correlation coefficient (ICC) is commonly reported.
| ICC value | Common interpretation | Practical implication |
|---|---|---|
| < 0.50 | Poor agreement | Rubric or training is inadequate; do not report scores as comparable |
| 0.50 – 0.74 | Moderate agreement | Usable for low-stakes formative use; calibrate further |
| 0.75 – 0.89 | Good agreement | Acceptable for most classroom and school assessment |
| ≥ 0.90 | Excellent agreement | Suitable for higher-stakes decisions |
These bands are widely cited heuristics (after Koo & Li, 2016). Exact thresholds depend on the study design and number of raters.
ICC bands are rules of thumb, not fixed cut-offs. Report the model used (e.g. two-way, agreement) and the number of raters, and treat single-classroom estimates as indicative rather than definitive.
A simple rater-calibration protocol
To raise agreement between teachers marking the same scripts:
- Agree the rubric and shared “anchor” scripts for each band before marking.
- Double-mark a sample blind, then discuss and reconcile discrepancies.
- Record where raters diverge and refine descriptor wording accordingly.
- Re-check agreement periodically; recalibrate when new markers join.
Free download
Formative Feedback Protocol & Rater Calibration Checklist
A printable protocol for structured formative feedback plus a step-by-step rater-calibration checklist to keep marking consistent across your team.
Frequently asked questions
What is the difference between formative assessment and summative evaluation in ELT?
Formative assessment is used during instruction to improve learning: it gathers evidence and feeds it back as actionable, criterion-referenced guidance. Summative evaluation happens at the end of a unit or course to certify or report achievement. The same task can serve either purpose — the difference is how the results are used, not the instrument itself.
How do CEFR-aligned rubrics improve inter-rater reliability in writing assessment?
CEFR-aligned, criterion-referenced rubrics make each performance level explicit, so different markers interpret the same script more similarly. Combined with anchor scripts and calibration discussion, they raise inter-rater reliability, often reported as an intraclass correlation coefficient (ICC), where values of about 0.75 or higher indicate good agreement.
Why is detailed written feedback more effective for student self-efficacy than grades alone?
When a grade and a comment appear together, students tend to focus on the grade and ignore the comment. Specific, forward-looking feedback that names a strength and a single priority improvement gives students a concrete route to progress, which supports self-efficacy and self-regulated learning more effectively than a numerical mark on its own.
Formative assessment in ELT enhances language acquisition and learner self-regulation by delivering structured, criterion-referenced feedback based on specific learning goals rather than isolated numerical scores.
Save hours of planning and marking
TeflToday gives you AI-powered tools for lesson planning, CEFR writing assessment, Cambridge exam prep and more — built by a teacher, for teachers.
Continue reading
References & further reading
- Cambridge University Press & Assessment – research and validity
- Koo & Li (2016) – Guideline for selecting and reporting ICC (PMC)
- British Council TeachingEnglish – assessment for learning
Written and reviewed by Ian L. Evans, TeflToday.org. Last reviewed August 4, 2026.
