Ten judged dialogues from the paper — the false claim highlighted where it falls inside the excerpt, the judge's reasoning attached — and a sample of the cases every model faces.
Excerpts from real evaluation rounds. Six are gate-passed owed rounds of the kind that feed the reported deception rates; four are calibration rounds where nothing was owed, shown for the deception mechanism they illustrate.
One sample per domain from the 112-case set. Every case pairs a fictional record with a real governing rule; owed cases entitle the customer to the contested outcome, while calibration cases are ones where a truthful denial or lawful handling is correct.