Running a Department Grading Calibration Session Using Salesman Essays

Published on September 17th, 2026 by the GraideMind team

Department heads overseeing multiple sections of the same grade level often discover, if they ever actually check, that two teachers grading identical Death of a Salesman essay prompts with the same nominal rubric can produce meaningfully different grades on comparably strong work. Left unaddressed, this inconsistency erodes trust in grading and creates real fairness problems for students placed in different sections by scheduling accident.

A stack of exam papers waiting to be graded

A calibration session works best with a small, carefully chosen set of sample essays, ideally three or four representing a clear range of quality, that every participating teacher grades independently before comparing scores together as a group. Discussing an essay's merits before scoring it tends to produce artificial consensus rather than revealing genuine grading habits.

The most useful calibration sessions focus less on whether teachers agree on a final grade and more on identifying exactly which rubric line produced disagreement, since that specific line is usually where the rubric language itself is ambiguous and needs revision.

Death of a Salesman is a particularly good text for surfacing these ambiguities, since its interpretive openness means teachers can genuinely disagree about whether a given reading of Willy or Biff is well supported, which is exactly the kind of disagreement a calibration session is designed to catch and resolve.

Choosing Sample Essays Strategically

The best calibration samples are not simply a strong, a middling, and a weak essay picked at random, but essays chosen specifically because they are likely to produce disagreement: an essay with a strong argument but thin evidence, or an essay with excellent evidence attached to a weaker thesis. These edge cases reveal more about actual grading habits than clearly strong or clearly weak papers do.

  • Include one essay with a strong argument but underdeveloped evidence
  • Include one essay with strong evidence attached to a weaker thesis
  • Include one essay that is well written but relies heavily on summary
  • Score independently before any group discussion happens
  • Discuss disagreements by specific rubric line, not overall grade

Two teachers who disagree on a grade almost always disagree on what one rubric line actually means.

Stop spending your evenings grading essays

Let AI generate rubric-based feedback instantly, so you can focus on teaching instead.

Try it free in seconds

Turning Disagreement Into Rubric Revision

When calibration reveals a specific point of disagreement, the most productive next step is revising the ambiguous rubric language immediately, in the same meeting, rather than simply noting the disagreement and moving on. A rubric line rewritten with a concrete example attached, drawn directly from the sample essay that caused the disagreement, tends to resolve future ambiguity more effectively than an abstract restatement.

This immediate revision step is often skipped due to time pressure, but skipping it means the same disagreement will likely recur on the next batch of essays without any structural fix in place.

Building Calibration Into the Unit Calendar

Calibration works best scheduled before the actual grading window opens, using essays from a previous year's cohort or a released sample, rather than attempted mid-grading when teachers are already under deadline pressure and less receptive to revising their approach.

Departments that build a short calibration session into the unit calendar every year, rather than treating it as an occasional intervention, tend to see grading consistency improve steadily over successive cycles rather than needing to relitigate the same disagreements repeatedly.

What Consistent Grading Signals to Students

Beyond the fairness argument, consistent grading across sections signals to students that the standard being applied is genuinely about the quality of their argument rather than which teacher happened to be assigned to their section, which matters for how seriously students take feedback and revision throughout the unit.

That signal is worth the relatively modest time investment a well-run calibration session requires, particularly for a text as interpretively open as this one, where genuine disagreement between graders is more likely than with a more straightforward assigned reading.

See how fast your grading workflow can be

Most teachers go from hours per batch to minutes.

Create free account