Part of Evidence of Educative Game Design. The four principles referenced here are set out in four principles that make a game teach.
How to tell whether a game teaches
A game commissioned for the first time has no delivery history to point to. What an evaluator can check instead is the process that built it, and that process has a name in the research literature.
Two gates, not one, and both before the end
The useful checks happen during development, not after delivery. The first gate is accuracy: does the content the game presents hold up against subject-matter expertise. The second gate is separate and equally necessary: does the target group of learners actually learn from it. A game can pass the first gate and fail the second, because expert review checks whether the content is correct, not whether the design teaches it.
Weston and colleagues compared feedback from subject experts against feedback from the learners the material was built for. They found that learner feedback improved the material more than expert feedback did. The finding is described here qualitatively rather than with an exact figure, since that figure could not be independently confirmed. The direction of the finding is what matters for practice: expert review alone is not a substitute for testing with real learners.
Playtesting is formative evaluation
Tessmer's work on formative evaluation supplies the vocabulary bridge that makes this concrete: playtesting, in a learning game, is formative evaluation. It is not a separate quality-assurance step bolted onto game design practice. It is the same activity the instructional design literature already has a name and a method for, applied to a game.
The Design-Based Research Collective set out build-test-revise as a recognised research method in its own right, not an informal shortcut taken before the real evaluation. This is the most useful citation in the article for a partner assessing a proposal. A structured, repeated cycle of building, testing with real learners, and revising is not a corner cut. It is a documented method with its own standing in the literature.
What to test for
Wouters and colleagues found no significant motivational advantage for games over conventional instruction on average. The consequence for testing is to treat engagement and enjoyment as a design outcome to verify, not an assumption to carry in on the strength of the format. Ask playtesters directly whether they wanted to continue, not only whether they completed what was asked of them.
Fullerton's playtesting practice sets a further standard worth holding to: test with people who are not the design team, and do it more than once. Revise between rounds, rather than running one round and calling it validation.
A number to avoid
Nielsen's rule says five users are enough to surface most usability problems. It is sometimes carried over into playtesting because it sounds like the same activity. It is not. Nielsen's finding concerns interface usability, where problems repeat quickly across users. Whether a game teaches specific content is a different question, and there is no equivalent shortcut for it. This article does not cite a minimum sample size, because the evidence above does not support one.
What this means for an evaluator
Ask what the test group looked like, how many rounds of testing happened, and what changed between them. A single test with the design team standing in for learners answers neither gate above. Repeated testing with the actual target group, followed by real revision, is what the evidence in this article supports as evaluation. It is also what the four principles depend on to land as intended, rather than as assumed.
References
- Design-Based Research Collective. (2003). Design-based research: An emerging paradigm for educational inquiry. Educational Researcher, 32(1), 5–8.
- Fullerton, T. (2024). Game design workshop: A playcentric approach to creating innovative games (5th ed.). CRC Press.
- Tessmer, M. (1993). Planning and conducting formative evaluations: Improving the quality of education and training. Kogan Page.
- Weston, C., Le Maistre, C., McAlpine, L., & Bordonaro, T. (1997). The influence of participants in formative evaluation on the improvement of learning from instructional materials. Instructional Science, 25, 369–386.
- Wouters, P., van Nimwegen, C., van Oostendorp, H., & van der Spek, E. D. (2013). A meta-analysis of the cognitive and motivational effects of serious games. Journal of Educational Psychology, 105(2), 249–265.