How to Test an AI Essay Grader Using Our Town Sample Essays
Published on September 18th, 2026 by the GraideMind team
A product demo will always look good. The essays are clean, the feedback is polished, and nothing goes wrong. What matters for a school or department is how the tool performs on your own students' writing, with your rubric, on a text you actually teach.

Our Town is a practical choice for a trial because so many teachers know it well. You can judge the feedback against your own understanding of the play, and you likely already have a range of student essays from past years. That gives you a ready-made test set.
Begin by assembling eight to ten essays that cover the full range of quality. Include at least one strong paper, a few in the middle, and a couple of weak ones. Remove student names and, if your school requires it, follow your privacy and data policies before anything is uploaded.
Grade the set yourself first, using your rubric, and record both scores and the comments you would give. Do this before looking at any tool output so your judgment stays independent. This becomes your benchmark.
What to Check in the Results
Once you have run the essays through the tool, compare its output against your benchmark on several dimensions. Scores matter, but they are only one piece. The quality and accuracy of the feedback often tell you more about whether the tool will hold up in practice.
Stop spending your evenings grading essays
Let AI generate rubric-based feedback instantly, so you can focus on teaching instead.
Try it free in seconds- Score alignment: how closely the tool's rubric scores match yours, and whether the differences are random or systematic
- Feedback specificity: whether comments refer to what the student actually wrote instead of offering generic advice
- Accuracy about the text: whether the tool correctly handles details from Our Town without inventing or misattributing lines
- Edge cases: how it handles very short essays, off-topic responses, and unconventional but well-argued interpretations
- Teacher control: how easily you can edit scores and comments before anything reaches a student
A tool that only looks good on average essays has not really been tested.
Testing the Difficult Cases
The middle of the pack is easy. The hard cases are the essays that break patterns: a brilliant but messy paper, a polished essay with no real argument, or a response that takes an unusual view of the third act. Include one or two of these deliberately, because they reveal how much nuance the tool can handle.
Also test consistency by running the same essay more than once, if the tool allows it. Large differences between runs would be a warning sign. A platform like GraideMind, which grades against a rubric you supply, should give stable and explainable results that you can trace back to your own criteria.
Making a Decision With Evidence
Set your standards before you look at results. For example, you might decide that scores must fall within a certain range of yours for most essays and that feedback must be usable with only light editing. Having a threshold in advance keeps you from being swayed by a good first impression.
Share your findings with colleagues, especially those who teach the same course. A department that has run the same trial together will have a common understanding of what the tool does well and where it needs oversight. That shared view makes any later rollout smoother.
See how fast your grading workflow can be
Most teachers go from hours per batch to minutes.
Create free account