Skip to content
InfoResearchPreprintLLM-specific

Reasoning-Token Spikes Under Prompted Untruthful Responding in Large Language Models

Published
Record updated
View JSON

Summary

Researchers tested whether the number of reasoning tokens a model generates can signal untruthful behavior, without needing access to the reasoning content. Three reasoning-capable large language models answered 210 multiple-choice questions under system prompts to respond truthfully, falsely, or without regard for truth. Truth-directed responses used fewer reasoning tokens than both lie-directed and truth-indifferent responses across all three models. The authors present this as a proof-of-concept signal, not yet a detector of spontaneous deception.