COGNITIKA

Part of Evidence of Educative Game Design. The four principles referenced here are set out in four principles that make a game teach.

How to tell whether a game teaches

A game commissioned for the first time has no delivery history to point to. What an evaluator can check instead is the process that built it, and that process has a name in the research literature.

By Ilija Bojović, Founder / Lead game designer · 16 August 2026

Two gates, not one, and both before the end

The useful checks happen during development, not after delivery. The first gate is accuracy: does the content the game presents hold up against subject-matter expertise. The second gate is separate and equally necessary: does the target group of learners actually learn from it. A game can pass the first gate and fail the second, because expert review checks whether the content is correct, not whether the design teaches it.

Weston and colleagues compared feedback from subject experts against feedback from the learners the material was built for. They found that learner feedback improved the material more than expert feedback did. The finding is described here qualitatively rather than with an exact figure, since that figure could not be independently confirmed. The direction of the finding is what matters for practice: expert review alone is not a substitute for testing with real learners.

Playtesting is formative evaluation

Tessmer's work on formative evaluation supplies the vocabulary bridge that makes this concrete: playtesting, in a learning game, is formative evaluation. It is not a separate quality-assurance step bolted onto game design practice. It is the same activity the instructional design literature already has a name and a method for, applied to a game.

The Design-Based Research Collective set out build-test-revise as a recognised research method in its own right, not an informal shortcut taken before the real evaluation. This is the most useful citation in the article for a partner assessing a proposal. A structured, repeated cycle of building, testing with real learners, and revising is not a corner cut. It is a documented method with its own standing in the literature.

What to test for

Wouters and colleagues found no significant motivational advantage for games over conventional instruction on average. The consequence for testing is to treat engagement and enjoyment as a design outcome to verify, not an assumption to carry in on the strength of the format. Ask playtesters directly whether they wanted to continue, not only whether they completed what was asked of them.

Fullerton's playtesting practice sets a further standard worth holding to: test with people who are not the design team, and do it more than once. Revise between rounds, rather than running one round and calling it validation.

A number to avoid

Nielsen's rule says five users are enough to surface most usability problems. It is sometimes carried over into playtesting because it sounds like the same activity. It is not. Nielsen's finding concerns interface usability, where problems repeat quickly across users. Whether a game teaches specific content is a different question, and there is no equivalent shortcut for it. This article does not cite a minimum sample size, because the evidence above does not support one.

What this means for an evaluator

Ask what the test group looked like, how many rounds of testing happened, and what changed between them. A single test with the design team standing in for learners answers neither gate above. Repeated testing with the actual target group, followed by real revision, is what the evidence in this article supports as evaluation. It is also what the four principles depend on to land as intended, rather than as assumed.

Ilija Bojović is Cognitika's founder and lead game designer, with nearly five years designing and leading commercial mobile games. More about the team.

References

Back to articles