Assessment Design

Reliability

20 min

Simple Explanation

Reliability is the consistency of assessment outcomes. A reliable assessment produces the same result regardless of who marks it, when it is taken, or which version is used.

Professional Definition

Reliability refers to the precision and consistency of assessment results. It encompasses inter-rater reliability (across markers), test-retest reliability (across occasions), parallel forms reliability (across versions), and internal consistency.

Explain It to a 10-Year-Old

Imagine three driving examiners test the same person on the same day. If one passes, one fails, and one is unsure — the test is unreliable. A reliable test gives the same result no matter who marks it.

Why It Matters

Unreliable assessment is unfair. Two equally competent learners should get the same result. If results vary by marker, location, or day, the qualification loses credibility and regulatory standing.

Where It Fits in the Process

This topic belongs to Assessment Design in the assessment development journey.

Assessment Strategy Assessment Methods Blueprint Question Writing Validity Reliability

Step-by-Step Process

1
Identify reliability threats

Where might inconsistency arise?

2
Create detailed mark schemes

Explicit criteria reduce marker variation

3
Conduct standardisation

Calibrate markers before live marking

4
Implement moderation

Sample-check marking during and after the process

5
Use statistical checks

Monitor inter-rater agreement and mark distributions

Worked Example

Improving reliability in Digital Administration:
Problem: Two markers give different marks for the same practical task.
Solution: Detailed mark scheme with exemplar responses. Standardisation meeting to calibrate judgments. Moderation sampling to catch drift.
Result: Marking consistency improves from 70% to 95% agreement.

Common Mistakes

MistakeProblemBetter Approach
Assuming clear questions guarantee reliability Even clear questions can be marked inconsistently Invest in mark scheme quality and marker training
Ignoring version equivalence Different paper versions must produce comparable results Use statistical equating or common items between versions

Checklist

of 4 complete

Related Tool

Apply what you learned with a practical tool.

Open Blueprint Builder

Interview Connection

You may be asked about this topic in assessment developer interviews.

Practice Interview Questions

Knowledge Check (3 Questions)

You scored %

Summary

  • Reliability = consistency of assessment outcomes
  • Inter-rater: same result from different markers
  • Test-retest: same result on different occasions
  • Parallel forms: same result from different versions
  • Reliability is necessary but not sufficient for validity