Lies We Can See: Joint Verbal and Non-Verbal Deception by VLM Agents in Embodied Social Interactions
Abstract
MineAmongUs introduces a 3D multimodal Among Us environment and the ARIA harness to study embodied VLM-agent deception through verbal and non-verbal actions, revealing non-verbal channels as key to winning.
Strategic deception by LLM and VLM agents has emerged as a central AI alignment and safety concern. Social-deduction games (where each player holds a hidden role and communicates with others to deduce identities) serve as the canonical testbed, particularly in multi-agent settings. Existing testbeds, however, are text-only and run on a single fixed agent configuration, missing the non-verbal sensorimotor channels treated as core by deception taxonomies and leaving it ambiguous whether an observed behavior reflects the underlying model or the surrounding harness. We introduce MineAmongUs, a 3D multimodal Among Us sandbox where imposter agents must deceive crewmates through joint verbal and non-verbal action. We also propose ARIA, a configurable VLM-agent harness that exposes five cognitive-component ablation axes; and an atom- and arc-level annotation scheme grounded in deception taxonomies and operationalized at scale by an LLM-as-a-Judge reaching near-human atom-labeling agreement. Empirical results show that VLM agents pursue imposter wins through joint verbal and non-verbal deception, with non-verbal channels emerging as the more decisive winning contributors across both harness ablation and cross-VLM evaluation. Taken together, our work opens a new path for embodied VLM-agent alignment research.
Community
“Can an agent lie with its body, not just its words?”
The blind spot: Deception follows the action space
- Chatbots could deceive only in conversation: a false claim in the transcript. Digital agents (e.g., code agents) now hold real permissions and act on our behalf, and deception has kept pace: faked test results, quietly disabled oversight. Each time agents gained a new way to act, deception followed.
- A "body" is the next action space. There, deception targets not the record but other agents' eyes. No evaluation today watches that channel. So we built a world where it can be watched, counted, and scored.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Belief Without Behavior: Measuring the Translation of Theory of Mind into Coordinated Social Action in Vision-Language Models (2026)
- MeetingToM: Evaluating Multimodal LLMs on Theory-of-Mind Reasoning in Multi-Party Meetings (2026)
- CaM-Wolf: Causal-Aware Multimodal Agents for Social Deduction Games (2026)
- MafiaScope: Non-Invasive, Time-Resolved Belief Probing for LLM Agents in Social Deduction Games (2026)
- AffAdapt: AFFect-driven ADAPTive AI Personas for Seamless Conversations (2026)
- Knowledge-Verified Emergent Deception in LLM Agents Under Conflicting Incentives (2026)
- Can Agents Deceive? Evaluating Reasoning and Deception in ParliamentBench using a Social Deduction Game (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2608.30428 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper
