Asia Chip Stocks Snap Back: KOSPI Jumps 2.8%, Nikkei Reclaims 63,000CXMT Soars 466% on Shanghai Debut, Becoming China's Most Valuable Listed CompanyCyera to Buy Oasis Security for $1 Billion as AI Agent Protection Becomes Cybersecurity's Hottest MarketThe Security Split: How One Breach Redrew the AI Industry's Battle LinesChina's Embodied AI Machine: Record Robotics Funding Meets an IPO Assembly LineA 557% Profit Surge That Counts as a Miss: What SK Hynix's Quarter Reveals About the AI TradeCORE Biomedicine Raises $21M Across Boston, Tokyo and Suzhou for AI-Guided Precision OncologyFish Audio Reels In $52M Seed a Year After Starting in a Bedroom — With $21M ARR and 8M UsersAsia Chip Stocks Snap Back: KOSPI Jumps 2.8%, Nikkei Reclaims 63,000CXMT Soars 466% on Shanghai Debut, Becoming China's Most Valuable Listed CompanyCyera to Buy Oasis Security for $1 Billion as AI Agent Protection Becomes Cybersecurity's Hottest MarketThe Security Split: How One Breach Redrew the AI Industry's Battle LinesChina's Embodied AI Machine: Record Robotics Funding Meets an IPO Assembly LineA 557% Profit Surge That Counts as a Miss: What SK Hynix's Quarter Reveals About the AI TradeCORE Biomedicine Raises $21M Across Boston, Tokyo and Suzhou for AI-Guided Precision OncologyFish Audio Reels In $52M Seed a Year After Starting in a Bedroom — With $21M ARR and 8M UsersAsia Chip Stocks Snap Back: KOSPI Jumps 2.8%, Nikkei Reclaims 63,000CXMT Soars 466% on Shanghai Debut, Becoming China's Most Valuable Listed CompanyCyera to Buy Oasis Security for $1 Billion as AI Agent Protection Becomes Cybersecurity's Hottest MarketThe Security Split: How One Breach Redrew the AI Industry's Battle LinesChina's Embodied AI Machine: Record Robotics Funding Meets an IPO Assembly LineA 557% Profit Surge That Counts as a Miss: What SK Hynix's Quarter Reveals About the AI TradeCORE Biomedicine Raises $21M Across Boston, Tokyo and Suzhou for AI-Guided Precision OncologyFish Audio Reels In $52M Seed a Year After Starting in a Bedroom — With $21M ARR and 8M Users
Abstract illustration of AI research on social reasoning
Skycrumbs
Research

MeetingToM: A Benchmark for the Social Skill AI Keeps Failing — Telling Real Agreement From Fake

A new multimodal benchmark tests whether AI can detect 'pseudo-consensus' in meetings — apparent agreement masking private dissent — and finds today's models still can't read the room.

D
Daniel ParkAI Correspondent
3 min read

As AI agents move from solo tasks into meetings, negotiations and multi-party collaboration, a new benchmark asks whether they can do the thing humans do instinctively and models do badly: infer what other people actually believe, as opposed to what they say. On MeetingToM, released to arXiv this week, the answer is a clear not yet.

Testing theory of mind where it's hardest

Theory of Mind — the ability to attribute beliefs, intentions and knowledge to others — is central to social interaction and a long-standing weakness of multimodal LLMs. MeetingToM stresses it in the setting where it's most demanding: naturalistic multi-party meetings, where the cues that reveal someone's real state of mind are scattered across speech, tone and behavior rather than stated outright.

The benchmark is organized hierarchically, evaluating reasoning at increasing social granularity: subject-level mental-state prediction (what does this individual believe?), dyadic-level addressee understanding (who is this remark actually aimed at?), and group-level consensus reasoning (does the room truly agree?).

The pseudo-consensus problem

MeetingToM's signature target is a phenomenon anyone who has sat in a status meeting will recognize: pseudo-consensus, where visible agreement masks private dissent that participants suppress under social pressure. Everyone nods; not everyone agrees. Detecting the gap requires integrating non-verbal signals — a hesitation, an averted glance, a too-quick concession — with the semantic content of what's said.

The paper's systematic evaluation of representative multimodal models finds persistent failures on exactly this integration: models struggle to fuse non-verbal cues, to infer hidden attitudes, and to distinguish genuine consensus from the performed kind.

Why it matters for agentic AI

The result is a pointed rebuttal to a comfortable assumption baked into the current agent boom — that models good at tasks are ready for teams. An AI that cannot tell real agreement from social capitulation is dangerous precisely in the roles it's being hired for: summarizing meetings, tracking decisions, facilitating group work, or acting as an autonomous participant. It will confidently record a consensus that does not exist and act on it.

A culturally loaded frontier

There is a dimension MeetingToM opens that the field will have to confront: pseudo-consensus is not culturally uniform. Indirectness, face-saving and the gap between spoken and intended meaning are weighted very differently across cultures — a live issue for the Asian enterprises and governments deploying meeting-agent AI at scale. A model trained to read Western-style directness may misjudge a room in Tokyo, Seoul or Jakarta in the opposite direction, hearing dissent as agreement or vice versa. Benchmarks like MeetingToM are the first step toward measuring a social competence that, unlike code correctness, does not have a single universal answer key.

Newsletter

Get Lanceum in your inbox

Weekly insights on AI and technology in Asia.

Share

More in Research

Lanceum

Independent coverage of AI and technology across Asia. We go beyond headlines to explain what matters.

Colophon

Typeset in Space Grotesk & DM Serif Display. Built with Nuxt & Tailwind. Powered by curiosity.

© 2026 Lanceum. All rights reserved.

Independent • Rigorous • Asia-Focused