All posts Assessment Best Practices

Comparing Two Batches That Did Not Sit the Same Test

Two parallel batches, same subject, same faculty, same fortnight. Batch A averaged 70%. Batch B averaged 64%.

The meeting reaches the obvious conclusion in about ninety seconds: Batch B is behind and needs extra classes. It is the wrong conclusion, and the numbers that produced it were all correct.

The two papers were not the same paper

This is the part that gets skipped. Parallel batches almost never sit an identical paper - they sit at different times, or on different days, or the faculty setting each paper wrote their own. So the two percentages being compared were produced by two different measuring instruments.

BatchPaper totalAverage marksAverage percentage
Batch A4028.070.0%
Batch B5032.064.0%

Converting to percentages feels like it solves the problem. It does solve one problem - the papers had different totals, and percentages remove that. What it does not remove is difficulty, because difficulty is not in the denominator. A percentage says what fraction of this paper a student got right. It says nothing about how hard this paper was compared to that one.

What a shared question set says instead

Now suppose both papers contained the same five questions, worth ten marks in total, deliberately placed in both. Score only those.

BatchWhole paperShared 10 marksShared as a percentage
Batch A70.0%6.262.0%
Batch B64.0%7.474.0%
Gap6.0 points to A12.0 points to B

On the only questions both batches actually answered, Batch B is ahead by twelve points. The comparison did not narrow. It reversed, by eighteen percentage points.

The arithmetic behind that is not mysterious. Take the shared marks out and Batch A scored 21.8 of the remaining 30 on its own paper - about 73% - while Batch B scored 24.6 of its remaining 40, about 62%. Batch B's paper was harder outside the shared section, and every percentage built on the whole paper carried that difficulty into the comparison as though it were a property of the students.

Why this matters more in coaching centres and universities than in schools

A school class usually sits one paper set by one teacher, so this problem stays small. The two places it does real damage are the two places parallel cohorts are normal.

Coaching centres

Weekly or fortnightly test series, several batches per subject, and a ranking that goes home to parents. When batch averages are compared without an anchor, a batch can be told it is behind because its paper was harder that week - and a batch can be told it is ahead for the same reason. Over a term of rankings the noise compounds, and the batch that gets the extra classes may not be the batch that needed them.

Universities and colleges

Multiple sections of the same course, often multiple faculty setting their own internal assessments, and a moderation meeting at the end of the semester that has to decide whether Section C's low average is the students or the paper. Without a shared component there is nothing in the data that can separate those two explanations, and the meeting resolves it by seniority instead of evidence.

The same problem appears across time as well as across sections. This semester's cohort is weaker than last year's is a claim about students that is usually a claim about a paper.

Three ways to make the comparison honest

  1. Carry an anchor set. Reserve 20 to 25% of every parallel paper for the same questions, set once and used in all versions. Compare batches only on that section. It costs nothing, it does not compromise the rest of the paper, and it is the only method here that measures the two groups with one instrument.
  2. Compare distributions, not means, when there is no anchor. If the papers genuinely share nothing, a mean is not comparable but a shape is. Where does the median sit relative to the top quartile? Is the batch tightly bunched or spread? A batch with a long tail needs different help from a batch that is uniformly mid, and that difference survives a change of paper where a mean does not.
  3. Record which paper each batch sat, permanently. This sounds like bookkeeping and it is the thing that makes the other two possible later. A result stored as a number with no link to the paper that produced it cannot be re-examined in the moderation meeting three months on.

What not to do

The obvious fix - give both batches the identical paper - works only if they sit it at the same moment. Parallel batches at a coaching centre are usually separated by hours or a day, and the paper reaches the second batch before the second batch does. That is not a reason to avoid a shared anchor: five questions leaking is a much smaller problem than a whole paper leaking, and the anchor is being used to compare cohorts rather than to rank individuals.

The other tempting fix is scaling - multiplying Batch B's marks by some factor until the averages agree. That assumes the answer in order to compute it. If the two batches genuinely differ in ability, scaling erases exactly the finding the comparison was for.

The question worth asking first

Before comparing two batches at all, it is worth being clear about what the comparison is for. If it decides who gets extra classes, it needs to be right, and it needs an anchor. If it is going into a ranking that parents see, it needs to be right and defensible, because somebody will eventually ask how it was calculated.

And if the honest answer is that the two papers shared nothing and the comparison cannot be made, that is a finding rather than a failure. It is considerably better than the ninety-second meeting that sends the stronger batch to remedial class.

Teemly stores every result against the paper that produced it and against the questions inside that paper, so a shared question set can be compared across batches, sections or semesters without re-entering anything. Question-level performance is visible across cohorts, which is what separates a hard question from a weak batch, and period comparisons keep the same student record intact when the cohort moves up a year. If you are working through the arithmetic that sits downstream of this, our guide to calculating a term percentage when tests have different total marks covers combining assessments with different maximums, and how negative marking changes a batch average covers a second way two batches can be made to look different without being different. You can see how the reporting works on our assessment software for colleges and universities and assessment software for coaching institutes pages.

Free PDF checklist

Not ready for a demo yet?

Get the Assessment Readiness Checklist - 42 questions to work through before you trust your next exam to a digital platform.

Get the free checklist
Book a free demo Explore Teemly