
MeetingToM: A Benchmark for the Social Skill AI Keeps Failing — Telling Real Agreement From Fake
A new multimodal benchmark tests whether AI can detect 'pseudo-consensus' in meetings — apparent agreement masking private dissent — and finds today's models still can't read the room.
As AI agents move from solo tasks into meetings, negotiations and multi-party collaboration, a new benchmark asks whether they can do the thing humans do instinctively and models do badly: infer what other people actually believe, as opposed to what they say. On MeetingToM, released to arXiv this week, the answer is a clear not yet.
Testing theory of mind where it's hardest
Theory of Mind — the ability to attribute beliefs, intentions and knowledge to others — is central to social interaction and a long-standing weakness of multimodal LLMs. MeetingToM stresses it in the setting where it's most demanding: naturalistic multi-party meetings, where the cues that reveal someone's real state of mind are scattered across speech, tone and behavior rather than stated outright.
The benchmark is organized hierarchically, evaluating reasoning at increasing social granularity: subject-level mental-state prediction (what does this individual believe?), dyadic-level addressee understanding (who is this remark actually aimed at?), and group-level consensus reasoning (does the room truly agree?).
The pseudo-consensus problem
MeetingToM's signature target is a phenomenon anyone who has sat in a status meeting will recognize: pseudo-consensus, where visible agreement masks private dissent that participants suppress under social pressure. Everyone nods; not everyone agrees. Detecting the gap requires integrating non-verbal signals — a hesitation, an averted glance, a too-quick concession — with the semantic content of what's said.
The paper's systematic evaluation of representative multimodal models finds persistent failures on exactly this integration: models struggle to fuse non-verbal cues, to infer hidden attitudes, and to distinguish genuine consensus from the performed kind.
Why it matters for agentic AI
The result is a pointed rebuttal to a comfortable assumption baked into the current agent boom — that models good at tasks are ready for teams. An AI that cannot tell real agreement from social capitulation is dangerous precisely in the roles it's being hired for: summarizing meetings, tracking decisions, facilitating group work, or acting as an autonomous participant. It will confidently record a consensus that does not exist and act on it.
A culturally loaded frontier
There is a dimension MeetingToM opens that the field will have to confront: pseudo-consensus is not culturally uniform. Indirectness, face-saving and the gap between spoken and intended meaning are weighted very differently across cultures — a live issue for the Asian enterprises and governments deploying meeting-agent AI at scale. A model trained to read Western-style directness may misjudge a room in Tokyo, Seoul or Jakarta in the opposite direction, hearing dissent as agreement or vice versa. Benchmarks like MeetingToM are the first step toward measuring a social competence that, unlike code correctness, does not have a single universal answer key.
Newsletter
Get Lanceum in your inbox
Weekly insights on AI and technology in Asia.
More in Research

Kimi K3's Weights Are Out: 2.8T Parameters, a 1.4TB Download, and an Escalation Nobody Can Undo

SciCodePile: A 128GB Corpus and a Brutal Benchmark Where Top Models Solve 12% of Scientific Code Tasks
