
- Title: Fostering teachers’ interpretations of formative assessment results using clustered heatmaps
- Authors: Sarah Bez, Sebastian Wurster, Martin J. Tomasik & Samuel Merk
- Access the original paper here
- Watch a video overview:
Paper summary
Two online experiments in Germany tested whether a different way of displaying assessment results helps teachers make sense of them. A clustered heatmap is a colour-coded grid of students and questions that software reorders, so students with similar patterns of right and wrong answers sit next to each other. In Study 1, 202 trainee teachers using it grouped students more accurately and faster, and wrote better learning goals, than those given a plain table in random order. In Study 2, 262 experienced teachers compared it with a colour-coded grid in alphabetical order, using real classroom results. Here accuracy was no better and the time saving was small. Study 1 also included a pattern where some students got questions wrong only because of how the questions were set out. Very few teachers spotted it, whichever display they used.
If teachers remember one thing from this study, it should be…
When results show a class struggling on one question, check how that question was asked before you reteach the whole topic. In this study, very few trainee teachers noticed students failing purely because of a question’s format, and a clearer display of the results made no difference.
*** PAPER DEEP DIVE ***
What are the key technical terms used in the paper?
- Formative assessment: checking students’ understanding to decide what to teach next.
- Clustered heatmap: a colour-coded grid of results, reordered by software so students with similar answer patterns sit together.
- Misinterpreted task type: a question format students misunderstand, so they get it wrong despite knowing the content.
- Pre-service teachers: trainee teachers.
What are the characteristics of the participants in the study?
Study 1 involved 202 trainee primary and secondary teachers at two German universities, 62% female, mostly finishing a bachelor’s or starting a master’s. Study 2 involved 262 practising teachers in Germany, all teaching German, 71% female, 46% with over 20 years’ experience. Both studies randomly assigned participants.
What does this paper add to the current field of research?
Earlier studies show teachers find it hard to spot patterns in assessment data, and much dashboard research has relied on teachers rating how easy a display is to use. The authors believe these are the first experiments testing whether clustered heatmaps change how accurately and quickly teachers read results.
What are the key implications for teachers in the classroom?
- Before reteaching a topic, look closely at how the question was asked. The researchers built a pattern into their made-up results. The lowest-attaining group could identify nouns, verbs and adjectives when ticking the right word type from a list, yet got the same content wrong when asked to underline or circle the word. Very few trainee teachers noticed this, and seeing the results as a heatmap made no difference. The authors suggest a display reaches its limit when spotting a pattern depends on knowing the subject and the questions. This will be familiar to anyone who has filled in a colour-coded question-level analysis. Question 3 on adding fractions is red for half the class, so the natural response is to reteach adding fractions. Before you do, read question 3 again. Was it worded unusually, laid out in an unfamiliar way, or did it rely on another skill first? Then compare it with other questions on the same content. If students did well on those, the topic may be secure and the question format is what needs attention, either by teaching students to recognise that format or by rewording the question next time.
- Arrange your results so that students with similar answers sit next to each other. Trainee teachers given the clustered heatmap grouped students more accurately than those given a plain black-and-white table in random order, and they did it faster. The gap was biggest with the most complex data, a set of made-up decathlon scores across ten events. They also wrote better learning goals for the groups they formed. A mark sheet in surname order tells you nothing about attainment, so it isn’t far from the random table in this study (although the researchers shuffled the question order too). You can move your own spreadsheet towards the heatmap. Sort students by total score, put questions on the same topic in neighbouring columns, add a total for each student and each question, and use conditional formatting to colour right and wrong answers. This isn’t exactly what the study tested. The software grouped students by their pattern of answers rather than their total, and the heatmap also added colour and totals, so we can’t tell which feature made the difference. If your school’s assessment platform has an option to group or cluster students, try it. Keeping questions on the same topic side by side also makes the comparison in implication one much easier.
- Expect newer teachers to gain most from a sorted display. In Study 2, experienced teachers compared the clustered heatmap with a colour-coded grid in alphabetical order, using real results from Swiss classrooms. Their accuracy was the same with both, and the clustered version saved only about 20 seconds. They also left slightly more students out of any group when using it. The authors suggest experienced teachers have routines built around the layouts they already know, and an unfamiliar one can get in the way. The real data were also messier than the made-up data in Study 1. If you already colour-code your results and have years of practice reading them, reordering may add little. It is more likely to help trainees and early career teachers who are still learning what to look for. If you mentor a trainee, sit down with a sorted, colour-coded version of a recent assessment and talk through which students cluster together, then look at the questions to work out the reasons.
- Use the patterns to plan your next lessons, and be wary of turning them into fixed ability groups. The researchers asked teachers to sort students into groups because it was a convenient way to check whether they had spotted the patterns. The authors state clearly that they are not recommending grouping by ability, and point to research showing it can have downsides for lower-attaining and disadvantaged students. Use what the grid shows to decide which content to revisit with the whole class, which questions to reword, and which students need a quick check-in during the next lesson. If you do group students, keep the groups short-lived and tied to a specific gap the data revealed.
Why might teachers exercise caution before applying these findings in their classroom?
Study 1’s large benefits came from made-up data with three obvious groups, compared against a shuffled black-and-white table. With experienced teachers and real data, the advantage largely disappeared. Simply sorting a spreadsheet by score was not tested, and every participant was judging students they had never taught.
What is a single quote that summarises the key findings from the paper?
“The result that nearly none of the pre-service teachers noticed this pattern in the assessment results regardless of whether they saw the data displayed in a table or in a clustered heatmap suggests that visualizations can come to a limit when content or context knowledge is needed to notice and interpret patterns in visualized data.“








