Running a Department Common Assessment on Monster With Calibrated Scoring

Published on September 25th, 2026 by the GraideMind team

Many English departments choose "Monster" for a common assessment because the book is accessible, widely available, and rich enough to support strong writing. The trouble begins when four or five teachers grade the same prompt and give very different scores to similar essays. Without calibration, a common assessment can produce data that says more about the grader than the students.

A stack of exam papers waiting to be graded

Calibration begins with a shared understanding of the rubric. Teachers should read each row together and discuss what each performance level looks like in practice, using real student writing when possible. These conversations often reveal that colleagues interpret terms like "insightful analysis" quite differently.

A typical calibration session takes about an hour and follows a simple format. Everyone scores the same three or four essays independently, then the group compares results and discusses any differences. The goal is not perfect agreement but a common standard that keeps scores meaningfully comparable.

Choosing Anchor Papers

Anchor papers are sample essays that represent each performance level, and they are the backbone of calibrated scoring. Choose examples from a previous year or from a pilot set, and remove identifying information. Include at least one essay that sits on the borderline between two levels, since those are the cases that cause the most disagreement.

  • One anchor for each performance level in the rubric
  • At least one borderline paper for discussion
  • Anonymous examples with student names removed
  • Short annotations explaining why each essay earned its score
  • A shared folder where all teachers can access the anchors

A common assessment is only common if every teacher reads the rubric the same way.

Stop spending your evenings grading essays

Let AI generate rubric-based feedback instantly, so you can focus on teaching instead.

Try it free in seconds

Checking Agreement After Grading

Calibration does not end when grading begins. A helpful practice is to have teachers swap a small sample of essays and score each other's papers, then compare results. Large gaps between scores highlight areas where the rubric needs clarification or where one teacher may be applying standards differently.

Tracking agreement rates over time also helps departments improve. If scores on the evidence row consistently diverge, that row probably needs clearer descriptors. Small adjustments made each year lead to steadily more reliable results.

Where AI Grading Helps a Department

An AI grading tool can act as an additional, consistent scorer. When every teacher uses the same rubric and the tool applies it uniformly, it creates a shared reference point that makes disagreements easier to spot. Teachers can compare their own scores against the tool's and discuss meaningful differences.

The tool also saves time on the initial grading pass, which leaves more room for collaborative conversations. Departments can spend meetings on interpreting results and planning instruction rather than on the mechanics of scoring. That shift makes common assessments far more useful.

Using the Data to Improve Instruction

The real value of a common assessment lies in what the department learns. If most students struggle to explain their evidence, that points to a need for shared instruction on that skill. Sharing successful lessons across classrooms turns the data into action.

Document what the department decides so that next year's team can build on it. A short summary of strengths, weaknesses, and planned changes keeps improvements from disappearing. Over several years, this process can transform how a department teaches writing.

See how fast your grading workflow can be

Most teachers go from hours per batch to minutes.

Create free account