Department Calibration: Scoring Salvage the Bones Essays Consistently Across Teachers
Published on October 9th, 2026 by the GraideMind team
In many schools, several English teachers assign the same Salvage the Bones essay and grade their own sections. Without coordination, an essay that earns a high score in one classroom might receive a much lower one in another. Students and families notice these differences, and they erode trust in the grading process. Calibration is the practice that closes the gap.

Calibration begins with a shared understanding of the criteria. Teachers might all agree that the rubric rewards interpretation, yet disagree about what interpretation looks like in a ninth-grade essay on Esch. Discussing concrete examples brings these differences to the surface. Surfacing disagreement is the whole point of the exercise.
A calibration session need not be long. Thirty to forty-five minutes is enough to read three or four sample essays, score them independently, and compare results. The conversation that follows, in which teachers explain their reasoning, is where alignment actually happens. Documenting the outcomes preserves the work for later use.
Running an Effective Calibration Session
Choose anonymous samples that represent a range of quality, including at least one borderline case. Ask each teacher to score without discussion, then reveal scores and discuss differences. The borderline essay produces the most useful debate because it forces teachers to articulate their thresholds. By the end, the group should agree on annotated anchors for each level.
- Select three to four anonymous samples across the quality range
- Have each teacher score independently before any discussion
- Compare scores and ask teachers to explain their reasoning with evidence from the essay
- Agree on anchor papers and write short annotations for each performance level
- Revise unclear rubric language based on points of disagreement
Consistency is not about identical opinions but about shared standards that everyone can explain.
Stop spending your evenings grading essays
Let AI generate rubric-based feedback instantly, so you can focus on teaching instead.
Try it free in secondsCommon Sources of Disagreement
Teachers often diverge on how much weight to give conventions versus ideas. One may penalize frequent errors heavily, while another focuses on the strength of the argument. Another source of difference is how generously to treat ambitious but flawed analysis. Naming these tendencies in the open helps the department decide on common practice.
Experience also matters, since veteran teachers sometimes hold implicit standards that newer colleagues have not learned. Calibration provides a structured way to share this knowledge. It benefits new teachers by clarifying expectations and benefits experienced teachers by prompting reflection. The result is a more cohesive department.
Maintaining Consistency Over Time
Calibration fades if it is a one-time event. Schedule brief check-ins midway through grading, where teachers exchange a few essays and compare scores. Keep anchor papers in a shared folder and update them annually. These small practices keep standards stable as staff and student populations change.
Consider tracking score distributions across sections. If one class shows a markedly different average, it may signal a calibration issue or a difference in instruction worth investigating. Data should prompt conversation rather than judgment. Used constructively, it supports continuous improvement.
How AI Can Support Departmental Consistency
Rubric-based AI grading tools apply the same criteria in the same way to every essay, which can act as a steady reference point. Teachers can compare their own scores to the tool's first-pass evaluation and investigate meaningful differences. This is not a replacement for professional judgment but a check on drift. Departments can also use it to examine whether rubric language is clear enough to be applied consistently.
Share findings with the department and revise the rubric where the tool and teachers consistently disagree. Those disagreements often reveal ambiguous wording. Clarifying it benefits human and automated grading alike. The outcome is a rubric that truly communicates standards.
See how fast your grading workflow can be
Most teachers go from hours per batch to minutes.
Create free account


