A Class 10 physics unit test, thirty students, five questions. The report says two of the questions were answered correctly by 36.7% of the class - the joint worst on the paper. They look like the same problem. They are opposite problems, and treating them the same wastes the next lesson.
The paper, as the class actually answered it
Here is the full picture, counting each question two ways: out of the whole class, and out of only the students who actually put something down.
| Question | Attempted | Correct | % of class | % of those who attempted |
|---|---|---|---|---|
| Q1 Define resistance (MCQ) | 30 | 26 | 86.7% | 86.7% |
| Q2 Series vs parallel (MCQ) | 30 | 11 | 36.7% | 36.7% |
| Q3 Read the circuit diagram (MCQ) | 29 | 17 | 56.7% | 58.6% |
| Q4 Numerical, three marks, late in the paper | 12 | 11 | 36.7% | 91.7% |
| Q5 Explain the experiment (written) | 0 | 0 | 0% | not gradable yet |
Thirty students, 150 submitted answers. The percentages here were produced by running the answer records through the same code the platform uses to maintain its per-question counters, not calculated by hand.
Q2 and Q4 are identical in the column most reports lead with, and 55 percentage points apart in the column that matters. Nineteen students got Q2 wrong having tried it. Only one student got Q4 wrong - the other eighteen never wrote anything at all.
Q2 is a teaching problem. Q4 is a paper problem.
Q2 was attempted by everybody and failed by two thirds. That is a concept the class does not have. Series versus parallel goes back on the board, and it goes back on the next paper to check whether the reteaching worked.
Q4 was answered correctly by nearly everyone who reached it. There is nothing wrong with the question and nothing wrong with the class's grasp of it. Eighteen students did not reach it, which is a statement about the length of the paper and the position of a three-mark numerical at the end of it, not about whether they can do the numerical. Reteaching this topic would be an hour spent on something the class can already do.
Both of those conclusions come from the same underlying data. Only one of them survives if you rank questions by percentage of the class.
A blank is not a wrong answer
This is the rule the whole thing rests on, and it is worth stating plainly: a question a student left untouched is evidence about the student's time management, not evidence about the question.
It is an easy rule to break by accident. When a paper is submitted, an unanswered question is not correct, so the tempting thing is to record it as incorrect and move on. Do that and every question near the end of a long paper acquires a reputation it did not earn. The questions that look hardest become simply the questions that came last, and the report quietly turns into a measure of paper length.
The fix is a denominator, not a filter. Keep counting attempts and correct answers separately, and read the ratio between them. A question with a low attempt count is telling you something real - it is just telling you about pacing rather than about difficulty, and pacing has a different fix.
The written question nobody has marked yet
Q5 is the one that catches people out. Twenty-four students wrote an answer to it. The table records zero attempts, because none of those answers has been marked, and an unmarked answer is not yet evidence of anything.
Showing it as zero is uninformative but honest. The dangerous alternative is to treat "not marked correct" as "marked incorrect", which is what happens if an ungraded answer is booked as a failure at submission time. A written question handled that way reads as one that every student in the batch has failed, and it keeps reading that way after the marking is done, because the marking never goes back and corrects the count. The question then sits at the top of every "hardest questions" list permanently, on the strength of a grade nobody ever gave it.
So there are three states, not two: correct, incorrect, and no evidence yet. Blank answers and unmarked written answers both belong in the third, and the third must not be silently folded into the second.
Reading the four cases
Once attempts and correct answers are counted separately, every question on a paper falls into one of four cases, and each one has a different action.
| Attempted by | Correct among them | What it means | What to do |
|---|---|---|---|
| Most of the class | High | The class has this | Nothing. Consider retiring the question as too easy |
| Most of the class | Low | A genuine gap, or a badly worded question | Reteach, and read the question again to be sure it is not the wording |
| Few | High | They ran out of time, not out of understanding | Shorten the paper or move the question earlier |
| Few | Low | The least informative square on the paper | Find out which it is before acting: ask the class, or set it again early in the next test |
The bottom-right case deserves a note. Few attempts and few correct answers is genuinely ambiguous - it could be a hard concept, or a question so confusing that most students skipped it and the few who tried misread it. There is not enough information in the numbers to tell those apart, and the honest move is to gather more rather than to guess. Setting the same question early in the next paper resolves it in one test.
Doing this without a spreadsheet exercise
None of this needs a statistics background, but it does need the per-question counts to exist. On paper tests they usually do not - the marks are totalled per student, the script goes back in the cupboard, and the question-level record is gone. That is why the same three or four topics get re-taught every year on the strength of memory.
The practical minimum is three numbers per question, carried forward across tests: how many students attempted it, how many got it right, and how many are still waiting on a grade. Ranked by the second divided by the first, a paper sorts itself into the four cases above in about a minute. Tracked across a term, the questions that keep appearing near the bottom are the syllabus points the batch has genuinely not absorbed, and they are usually not the ones anybody would have guessed.
The counts also have to be a projection of the marks rather than a running tally of button presses. If a student retakes a test, or a written answer is regraded from wrong to right, the per-question counts have to follow the corrected record. A counter that only ever increments drifts away from the marks it is supposed to summarise, and it drifts in a predictable direction - toward incorrect, because the earlier attempt is usually the worse one.
What this is actually for
Question-level analysis is not about grading the class harder. It is about spending the next lesson on the thing that is actually missing. A batch average tells you whether to worry. The question-level breakdown tells you what to do on Monday, and it is the difference between "they did badly, push them harder" and "they do not have series versus parallel, and the paper is twenty minutes too long".
Teemly keeps attempted, correct and incorrect counts against every question in the bank, across every test that question has appeared in, and updates them from the stored result rather than from the act of submitting - so retakes and regrades correct the record instead of inflating it. Unanswered and ungraded answers are held out of the counts rather than booked as failures. If you are working through the numbers that sit around this one, how negative marking changes a batch average covers a second way a paper can be made to look harder than it was, and comparing two batches that did not sit the same test covers what a shared question set lets you say across cohorts. You can see how the question bank stores this on our question bank software page.