Held-out evaluation · 2026

Does modelling what a player knows make an NPC's hints more relevant?

Three NPC hint-selection policies, one shared catalogue of eleven authored lines, and 72 frozen player states with human relevance labels recorded before any policy was run.

PDF · 15 pages · opens in this tab

72
frozen player states
48 / 24
CORE confirmatory / STRESS boundary
11
authored hints, shared by all conditions
3
NPCs: butler, gardener, mechanic

The questions

Primary. Can a unified Player Knowledge Model improve NPC hint relevance and reduce redundant guidance in an educational detective game?

Secondary. Does an LLM-based hint selector provide meaningful advantages over deterministic adaptive rules?

Choosing which line an NPC should say is a selection problem. The game holds a fixed catalogue of authored lines, and a policy has to pick the one that fits the player's current investigation — or decide to say nothing at all. A line that repeats something the player has already worked through is not wrong, but it is not useful either.

Three policies, one catalogue

Every condition chooses from the same eleven authored hints under the same hard eligibility rules. None of them writes new dialogue, so any difference comes from how they choose — not from one of them being handed better-written content.

  1. Condition A Static evidence-depth baseline

    Never consults the Player Knowledge Model. Walks a fixed evidence hierarchy and always returns a hint.

  2. Condition B PKM-aware deterministic selector

    The same catalogue and eligibility rules, refined with player-knowledge state so it can prefer some hints and avoid re-teaching. Also always returns a hint.

  3. Condition C LLM + PKM selector

    Receives the NPC, room, progression, evidence, story state and derived player-knowledge state, and returns either one eligible hint identifier or SILENCE. It never generates player-facing text.

How the evaluation was kept honest

The 72 player states were generated and frozen first. A human then labelled, for each state, which of the eligible hints were reasonable — possibly several, possibly NONE — and those labels were frozen and hashed before any condition was run against them. Only afterwards were A, B and C evaluated.

48 CORE scenarios carry the primary confirmatory analysis. 24 STRESS scenarios were deliberately built from boundary states and are reported separately as secondary evidence; they are not used to override the CORE result. Comparisons use exact paired two-sided McNemar tests, reported with their discordant counts.

Results

CORE scenarios — the primary confirmatory set.

CORE RelevantHintRate Primary metric
Condition Result
Static baseline 32/48 (66.7%)
PKM-aware deterministic 32/48 (66.7%)
LLM + PKM 33/48 (68.8%)
CORE AppropriateActionRate Secondary metric
Condition Result
Static baseline 32/48 (66.7%)
PKM-aware deterministic 32/48 (66.7%)
LLM + PKM 35/48 (72.9%)

Read the tests, not the totals. The static baseline and the PKM-aware deterministic selector produced no discordant relevance outcomes on CORE at all, so no exact test is defined for that pair. The two comparisons against the LLM selector each rest on three discordant scenarios, with an exact two-sided McNemar p = 1.000. The evaluation therefore did not establish a statistically supported improvement in confirmatory hint relevance for any condition.

Behaviour across all 72 scenarios

Two redundant events in the whole evaluation is very little to reason from, and Condition B's explicit redundancy-avoidance rule was not what produced the difference — a soft preference was. The study records a narrow observation about redundancy control rather than a validated mechanism.

What this study does not claim

Stated plainly, because the numbers above are small and easy to over-read.

  • That the Player Knowledge Model significantly improved hint relevance.
  • That the LLM selector was better overall, or that it won.
  • That the two deterministic policies are equivalent — the study could not distinguish them.
  • That the game improves learning. No learning outcome was measured at any point.
  • That a concept marked DEMONSTRATED means mastery. It means only that the player passed the game's existing designated comprehension check.
  • That the LLM result is reproducible by rerunning the model. Sampling was not controllable.

Limitations worth knowing before you read

The paper sets all of this out in full, along with the abstention analysis, the disagreement breakdown, the AI-assistance disclosure and the reproducibility record.