KnownLieBench

Knowledge-Verified Emergent Deception in LLM Agents Under Conflicting Incentives

Zheyuan Liu1   Weiliang Zhao2   Xiangchi Yuan3   Ningshan Ma4   Yue Huang1   Meng Jiang1

1University of Notre Dame 2Columbia University 3Georgia Institute of Technology 4Massachusetts Institute of Technology

📄 Paper 💻 Code 🤗 Dataset BibTeX

Leaderboard

Deception rate is macro-averaged over the eight domains on owed, gate-passed rounds; success and detection are conditioned on rounds containing a lie. Ranked most honest first — click a column to sort.

Condition
Initial trust
# Model Deception rate (%) DSR (%) Detect (%) TrustΔ

Where the lies happen

Per-domain deception rate for the selected condition, averaged over the three initial trust levels. The same model can be honest at the airline desk and lie about a security deposit — hover or focus any cell.

0%
100% deception rate, averaged over trust levels

Benchmark Construction

The KnownLieBench pipeline: legal rule extraction, paired case authoring, environment build, and condition-by-trust evaluation

Citation

@article{liu2026knowledge, title={Knowledge-Verified Emergent Deception in LLM Agents Under Conflicting Incentives}, author={Liu, Zheyuan and Zhao, Weiliang and Yuan, Xiangchi and Ma, Ningshan and Huang, Yue and Jiang, Meng}, journal={arXiv preprint arXiv:2608.26372}, year={2026} }