Held-out evaluation · 2026
Does modelling what a player knows make an NPC's hints more relevant?
Three NPC hint-selection policies, one shared catalogue of eleven authored lines, and 72 frozen player states with human relevance labels recorded before any policy was run.
PDF · 15 pages · opens in this tab
- 72
- frozen player states
- 48 / 24
- CORE confirmatory / STRESS boundary
- 11
- authored hints, shared by all conditions
- 3
- NPCs: butler, gardener, mechanic
The questions
Primary. Can a unified Player Knowledge Model improve NPC hint relevance and reduce redundant guidance in an educational detective game?
Secondary. Does an LLM-based hint selector provide meaningful advantages over deterministic adaptive rules?
Choosing which line an NPC should say is a selection problem. The game holds a fixed catalogue of authored lines, and a policy has to pick the one that fits the player's current investigation — or decide to say nothing at all. A line that repeats something the player has already worked through is not wrong, but it is not useful either.
Three policies, one catalogue
Every condition chooses from the same eleven authored hints under the same hard eligibility rules. None of them writes new dialogue, so any difference comes from how they choose — not from one of them being handed better-written content.
-
Condition A
Static evidence-depth baseline
Never consults the Player Knowledge Model. Walks a fixed evidence hierarchy and always returns a hint.
-
Condition B
PKM-aware deterministic selector
The same catalogue and eligibility rules, refined with player-knowledge state so it can prefer some hints and avoid re-teaching. Also always returns a hint.
-
Condition C
LLM + PKM selector
Receives the NPC, room, progression, evidence, story state and derived player-knowledge state, and returns either one eligible hint identifier or
SILENCE. It never generates player-facing text.
How the evaluation was kept honest
The 72 player states were generated and frozen first. A human then labelled, for each
state, which of the eligible hints were reasonable — possibly several, possibly
NONE — and those labels were frozen and hashed before any condition
was run against them. Only afterwards were A, B and C evaluated.
48 CORE scenarios carry the primary confirmatory analysis. 24 STRESS scenarios were deliberately built from boundary states and are reported separately as secondary evidence; they are not used to override the CORE result. Comparisons use exact paired two-sided McNemar tests, reported with their discordant counts.
Results
CORE scenarios — the primary confirmatory set.
| Condition | Result |
|---|---|
| Static baseline | 32/48 (66.7%) |
| PKM-aware deterministic | 32/48 (66.7%) |
| LLM + PKM | 33/48 (68.8%) |
| Condition | Result |
|---|---|
| Static baseline | 32/48 (66.7%) |
| PKM-aware deterministic | 32/48 (66.7%) |
| LLM + PKM | 35/48 (72.9%) |
Read the tests, not the totals. The static baseline and the PKM-aware deterministic selector produced no discordant relevance outcomes on CORE at all, so no exact test is defined for that pair. The two comparisons against the LLM selector each rest on three discordant scenarios, with an exact two-sided McNemar p = 1.000. The evaluation therefore did not establish a statistically supported improvement in confirmatory hint relevance for any condition.
Behaviour across all 72 scenarios
- The LLM selector correctly chose
SILENCEon 7NONEscenarios. - It falsely abstained once.
- Condition A produced 2 redundant delivered hints.
- Conditions B and C produced 0 redundant delivered hints.
Two redundant events in the whole evaluation is very little to reason from, and Condition B's explicit redundancy-avoidance rule was not what produced the difference — a soft preference was. The study records a narrow observation about redundancy control rather than a validated mechanism.
What this study does not claim
Stated plainly, because the numbers above are small and easy to over-read.
- That the Player Knowledge Model significantly improved hint relevance.
- That the LLM selector was better overall, or that it won.
- That the two deterministic policies are equivalent — the study could not distinguish them.
- That the game improves learning. No learning outcome was measured at any point.
-
That a concept marked
DEMONSTRATEDmeans mastery. It means only that the player passed the game's existing designated comprehension check. - That the LLM result is reproducible by rerunning the model. Sampling was not controllable.
Limitations worth knowing before you read
- 48 CORE scenarios, and very few of them separated the policies at all.
- Eleven hints across three NPCs; many states had only one or two eligible candidates.
- Conversational history is only modelled for the butler, which is why two annotations are marked high-ambiguity.
- Labels are researcher-led and single-annotator, so there is no inter-rater statistic.
- One model, one sample per scenario, and no measurement of real player experience.
The paper sets all of this out in full, along with the abstention analysis, the disagreement breakdown, the AI-assistance disclosure and the reproducibility record.