There is no assessment strategy that gives you time, consistency, and depth all at once. Every choice trades one for another — the trick is choosing on purpose.

Somewhere around the fifteenth persuasive writing piece on a Sunday night, most of us have had the same thought: could I just use a checklist instead? The rich rubric felt right in the planning meeting. At 9pm, with thirteen scripts still in the pile, it feels like a decision someone else made on your behalf.
That moment is worth sitting with, because it’s not really about tiredness. It’s about a trade-off nobody named out loud when the assessment was designed. In Part 2 of this series, we looked at matching assessment strategies to record-keeping systems. This final part asks the harder question: what does each strategy quietly cost you, and is that cost one you’d choose if someone made you say it out loud?
Three costs show up again and again in the research: time, consistency, and depth of insight. No strategy escapes all three. The honest move isn’t finding the mythical assessment that avoids every trade-off — it’s knowing which one you’re carrying, for which task, and why.
Trade-Off One: Time
Let’s start with the cost every teacher already feels in their body. Australian Education Union data puts full-time teachers at an average of 49.3 hours a week, with marking and planning dominating the hours beyond the timetable. A 2022 AEU South Australian survey found teachers were doing 30 hours a week beyond face-to-face teaching alone. Richer assessment — detailed rubrics, portfolios, oral check-ins — produces better evidence of learning. It also produces more hours you don’t have.
This isn’t an argument against depth. It’s an argument for honesty about the exchange rate. A five-minute exit ticket and a forty-minute portfolio conference both tell you something real about a student. They do not cost the same, and pretending otherwise is how burnout starts. For quick formative checks that protect your time without sacrificing useful data, the formative assessment approaches we’ve covered before are worth revisiting — they’re built for exactly this trade-off.
In practice, the fix isn’t cramming more hours into the week. It’s matching assessment weight to decision weight. A task that shapes tomorrow’s teaching deserves a fast, low-cost check. A task that reports a term’s learning to families deserves the time it takes. Treating every piece of work as if it needs the same depth is what makes the Sunday night pile unmanageable.
Trade-Off Two: Consistency
Here’s the uncomfortable one. Even careful, experienced markers are not as consistent as they believe themselves to be. A 2023 experimental study found that a student’s grade in one subject was significantly influenced by how that same student had performed in an entirely unrelated subject — the classic halo effect, and it held even after researchers controlled for gender and status (Schmidt et al., 2023). This isn’t a character flaw in individual teachers. It’s a known feature of how human judgement works under load, and it’s exactly what assessment for, as, and of learning tries to guard against when it separates the purpose of a judgement from the moment it’s made.
There’s a useful nuance here, though. Experienced teachers are better than novices at reading both the obvious “surface” signals — hands up, eyes down — and the harder “deep” signals, like a quiet student who might be disengaged or might just be thinking. That skill develops. It doesn’t develop by accident, and it doesn’t fully protect against bias on its own — structure still matters more than good intentions.
- Check the extremes first. Re-read your highest and lowest grades before anything else. That’s where halo effects and fatigue errors cluster.
- Mark one question across all scripts, not one script across all questions. This keeps your internal standard consistent instead of drifting as you tire.
- Run a five-minute calibration before you start. Grade one sample together with a colleague before marking solo. It costs almost nothing and catches drift before it happens, not after.
- Blind-mark when the stakes are high. Covering the name is a small step that removes a documented source of bias for borderline or subjective work.
This is also the territory AITSL Standard 5.3 points to directly — making consistent and comparable judgements isn’t a bonus skill, it’s a named professional standard. Worth remembering next time moderation feels like an add-on rather than the job itself.
Trade-Off Three: Depth of Data
The third cost is the one we talk about least, possibly because it’s the hardest to see from inside a busy term: what a grade actually tells you about a student’s understanding.
Recent research has raised real doubt about that link. Comparisons between teacher-awarded grades and external, course-aligned tests have found meaningful mismatches — evidence that a grade can look confident and still not accurately describe what a student understands. That gap widens when assessment rewards compliance — submitting on time, following the format, filling every box — over genuine thinking. AJ Juliani’s phrase for this, “doing school,” captures it well: students becoming skilled at producing what a task asks for without the underlying understanding ever being tested.
The fix isn’t abandoning efficient assessment formats. It’s building in occasional deeper checks — a follow-up question, a brief verbal explanation, a second look at borderline work — before a grade goes into the record permanently. Even one extra “tell me how you got there” question can surface whether a correct answer reflects understanding or a lucky guess. This connects directly to performance of worth: tasks that ask students to apply learning somewhere real are far harder to complete through compliance alone.
Choosing Your Trade-Off on Purpose
No single assessment strategy solves time, consistency, and depth simultaneously. A quick multiple-choice check is fast and reasonably consistent, but shallow. A rich portfolio conference is deep and genuinely revealing, but slow. A shared rubric with clear criteria protects consistency, but only if you also protect the time to moderate it properly.
That’s not a flaw in your planning. It’s the actual shape of the problem. The teachers who manage assessment well aren’t the ones who’ve found a loophole — they’re the ones who’ve stopped expecting one, and instead ask a simple question before choosing a strategy: for this task, which cost can I afford, and which one matters most?
- Low stakes, shapes tomorrow’s teaching? Choose speed. A fast formative check beats a slow, perfect one you never get to use.
- Reported to families or used for placement? Choose consistency. Build in moderation time even if it means assessing less often.
- Central to whether real understanding happened? Choose depth. Add one verbal or applied check before the grade is locked in.
That’s the honest version of assessment design. Not a search for the perfect tool, but a series of deliberate choices about which cost you’re willing to carry, task by task, term by term.
References
Australian Education Union. (2023). The heavy hours: Work intensification for teachers and school leaders. aeufederal.org.au
Australian Institute for Teaching and School Leadership (AITSL). (2018). Australian Professional Standards for Teachers — Standard 5.3: Make consistent and comparable judgements.
Juliani, A. J. (2026). The game of school is out of control. ajjuliani.beehiiv.com.
NSW Department of Education. (2024). Assessment practice — consistent teacher judgement. education.nsw.gov.au
Schmidt, F. T. C., Kaiser, A., & Retelsdorf, J. (2023). Halo effects in grading: An experimental approach. Educational Psychology, 43(2–3), 246–262.








