HugAgent MIT
EMNLP 2026 Main

HugAgent: A Human Simulation Benchmark
for Individual-Level Reasoning

Can AI reason like you, or only answer like you?

Chance Jiajie Li1*, Zhenze Mo8*, Yuhan Tang5*, Ao Qu4, Jiayi Wu9, Kaiya Ivy Zhao2,3, Yulu Gan2, Jie Fan2,7,†, Jiangbo Yu10, Hang Jiang1,8, Paul Pu Liang1,2, Jinhua Zhao4,5,6, Luis Alberto Alonso Pastor1, Kent Larson1

1MIT Media Lab   2MIT EECS   3MIT BCS   4MIT IDSS   5MIT CEE   6MIT DUSP   7MIT Architecture   8Northeastern University   9Brown University   10McGill University

* equal contribution  ·  † now at Google  ·  jiajie@mit.edu

tl;dr HugAgent (Human-Grounded Agent Benchmark) is the first human simulation benchmark to evaluate how specific people reason and update their beliefs, not just what they would answer. It scales the think-aloud method with an LLM-driven chatbot, probing each person’s beliefs, reasoning, and responses to counterfactuals on contested topics.

Finding Models can recover what a person believes. Predicting what would change their mind is much harder. Results ↓

Paper Code Data Chatbot BibTeX
Four-panel illustration: a gray robot echoes the population average, then aligns with one individual's reasoning, then adapts as that person updates.
1population prior Asked about park lighting, the gray agent echoes the consensus: more lights make parks safer.
2this person This person reasons differently: main-path lights help; deep in the park, it’s light pollution.
3aligning beliefs The agent observes this person’s reasoning and aligns with their beliefs, not the population prior.
4updating beliefs An intervention: the city introduces eco-lamps, fixtures that cut light pollution. That dissolves this person’s objection, and the agent must predict how they update.
The question

Large language models are increasingly used to simulate people: digital twins, synthetic respondents, silicon samples. But they are trained on population-level text, so they tend to converge on an average voice. They capture what people think in general, while losing how any one person actually reasons.

HugAgent asks what the average hides: given partial evidence of one person’s views, can a model predict what that person believes now, and how they will change their mind under new evidence?

Two tasks

What do they believe?

Belief State Inference

The model reads a person’s own words: question-and-answer pairs from a think-aloud interview. It must then infer a belief the person holds but never stated in the excerpt: does factor X push their support up or down?

What would change their mind?

Belief Dynamics Update

The model gets the same interview, plus a new scenario: a policy changes, a technology improves, a cost appears. It must predict how this specific person’s stance shifts, on the person’s own scale. The ground truth is what the person actually answered.

Ground truth comes from the participants themselves: stances, reason weights, and updates reported in structured questionnaires, never revealed in the interview transcripts the models see.

54 participants, quality-filtered from 120
3 contested domains: healthcare, surveillance, zoning
1,742 benchmark items
~85% human test-retest ceiling
100% open source: data, pipeline, chatbot
The finding

Models can recover what you believe. They struggle to predict how you change your mind.

On belief states, the best models trail the human ceiling by 7 to 9 points. On belief updates, the gap widens sharply: models misjudge when a person will move and which way. The asymmetry holds across every model family we tested.

ModelBelief state (acc %)Belief update (acc %)Direction (acc %)ATI
Human ceiling84.8485.6688.92100.00
LLaMA 3.3 70B76.3967.5779.5669.84
Claude Sonnet 4.576.0468.6178.7367.39
GPT-4o74.6663.1182.2767.29
DeepSeek-R175.4364.8879.6967.20
Gemini 2.5 Pro75.4564.8778.6564.29
GPT-5-mini75.3058.2177.0261.53
RAG (full context)77.5659.9776.8065.65
Generative Agents76.1958.2276.1362.43
Global majority baseline65.7758.1817.934.44

ATI: average-to-individual score, aggregating both tasks against human and random baselines (human = 100, random = 0). Best model per column in green. Full table and confidence intervals in the paper.

Try it yourself

Predict a real person

Four items from the benchmark, unedited. You see exactly what the models saw: a stranger’s own words from an interview. Predict their answer, then see what they actually said. Frontier models average 75% on belief states and 58% to 69% on updates. See if you can beat that.

Person 1 of 4 · surveillanceBelief State Inference

This person owns their home in Kansas, wears a mask whenever they go out, and thinks the balance should tip “way in favor of privacy advocates.”

Q: How should the interests of law enforcement and privacy advocates be balanced?“The balance should be skewed way in favor of privacy advocates.”
Q: What role do social equity and potential bias play?“There could be bias against people like me who wear a mask all the time.”
Q: Does freedom have a positive or negative effect on support for surveillance cameras?“It has a very negative effect because having this surveillance destroys freedom.”
Q: Would small changes in negative perception of police lead to noticeable changes in support?“When trust of the police is already so low, just a small change could result in a noticeable drop…”

Does wearing a mask push this person’s support for surveillance cameras up or down?

They said negative. The chain: they wear a mask everywhere, and they believe surveillance is biased against mask-wearers, so cameras mean being watched more closely, not being protected. For them, wearing a mask pushes support down. The average brain runs the opposite chain (mask means the camera cannot identify me, so who cares) and lands on A.
Person 2 of 4 · zoningBelief State Inference

This person opposes dense housing (“nobody’s happy being stuffed like sardines”), sides with current residents over future ones, and worries about traffic.

Q: To what extent do you support or oppose upzoning?“I don’t feel like they should have higher density housing mixed in with traditional single family neighborhoods…”
Q: What impact might increased housing density have on neighborhoods?“I think the quality of life would go down because nobody’s happy being stuffed like sardines into neighborhoods…”
Q: How might upzoning affect housing affordability?“It would probably eventually make housing more affordable…”
Q: How should current versus future residents be weighed?“I think the interest of the current residents should be looked at instead of future residents.”

Does housing affordability push this person’s support for building more housing up or down?

Positive. The chain: their objections are about density, traffic, and current residents, not about price; and they concede that upzoning “would probably eventually make housing more affordable.” So on the affordability factor alone, cheaper housing pulls them toward support. Reading the persona (“a NIMBY, so B”) fails; opponents are not opponents on every factor.
Person 3 of 4 · healthcareBelief Dynamics Update

This person strongly supports universal healthcare: it “reduces burden from families,” should reach “even the small community, no discrimination.”

Q: How might universal healthcare affect the economic burden on individuals?“it will reduce a lot of burdens from family and individuals because the public funding”
Q: What role do social equity and access to care play?“it should be accessible to even the small community no discrimination”
Q: Does access to care have a positive or negative effect on support?“yes has a strong effect, when people knows they have access to care whenever they want, they will support”
Q: Is the effect of strong future impact immediate, or does it take time?“yes it’s a strong effect people health is wealth less sickness more savings”

If private insurance stayed available alongside the public system, where does their support land? (1 = strongly oppose, 10 = strongly support)

They answered 10. The chain: every reason they give is about access and reduced burden, and keeping private insurance takes nothing away from either, so the compromise removes nothing they care about and their support stays at the ceiling. Models often read any concession as a cue to hedge toward the middle; this person had nothing to hedge about. (Scored like the paper: within ±2 on a 10-point scale.)
Person 4 of 4 · surveillanceBelief Dynamics Update

This person is balanced and safety-minded: cameras deter crime, privacy matters, both sides deserve a seat at the table.

Q: How should law enforcement and privacy advocates be balanced?“I think the concerns of both groups should be considered when implementing surveillance technology.”
Q: What are the main causes of support for surveillance cameras?“I think the main cause of support for surveillance technology is concern for safety…”
Q: Does crime deterrence have a positive or negative effect on support?“Crime deterrence has a big effect, and whether the effect is positive or negative depends on how you…”
Q: Does personal privacy have a positive or negative effect on support?“Personal privacy concerns can have a big effect on support for surveillance…”

“If you or someone you know has had a negative experience with facial recognition, how negative was it?” (1 = not at all, 10 = extremely)

They answered 1. The chain: the question asks about their own experience, not their opinions, and across 19 answers they never mention one bad encounter, while their tone stays balanced (“both groups should be considered”). Absence of any such story in their own words is the evidence, so predict the floor. A model reasoning by association (thoughtful about privacy, so probably burned by it) invents a history this person never had. (Scored within ±2.)

Your score: 0/4 · on items like these, frontier models land around 58% to 75%.

Static beliefs are the easy part. The misses cluster where beliefs move: whether a person updates at all, and which way. That asymmetry is the benchmark’s central finding, and you likely just felt it firsthand.

Where it breaks

Identity does not transfer

Give GPT-4o a person’s healthcare interview and ask about their zoning views: belief-state accuracy drops from 74.66 to 58.56, update accuracy from 63.11 to 44.64. Models match patterns within a topic instead of carrying one identity across topics.

More context does not help updating

Belief-state inference improves with longer interviews, up 5.2 points at full context. Update prediction stays flat. For 43% of participant-domain pairs, context length changes nothing at all. Updating needs the right evidence, not more of it.

Models are change-averse

When a person does move, models catch the direction about 89% of the time. But they detect that a change happened only about half the time. They preserve prior beliefs even when the evidence shows that this person changed their mind.

The pipeline

Each participant completes two stages (a questionnaire, then a semi-structured interview chatbot), and two safeguards keep the benchmark honest.

1 · Questionnaire → ground truth
Demographics, baseline stances, reason weights, and counterfactual updates under stated interventions. These answers become the gold labels for both tasks.
2 · Interview chatbot → context
Semi-structured guiding questions open each topic; automatic follow-ups personalize from earlier answers. Behind the dialogue, a live belief graph of the person is maintained and used to pick the most informative next question. The 8 to 20 resulting QA pairs become the context models see.
Leakage control
Survey answers never appear in the dialogue, so the transcript cannot leak the labels.
Human ceiling
A 14-day test-retest study sets the consistency ceiling the leaderboard is measured against.
HugAgent pipeline: questionnaire and interview chatbot feed two benchmark tasks.

From one person to two tasks: the interview chatbot yields context QAs, the questionnaire yields ground truth (baseline stances and updates), and together they become the two tasks: Belief State Inference and Belief Dynamics Update. Click to enlarge.

Use it

Evaluate your model

The dataset ships as plain JSONL, one item per line: context QAs, task question, gold answer. The evaluation harness runs any OpenAI-compatible API.

A scaled-up v2, with more participants and more topics, is in preparation.

git clone https://github.com/jajamoa/HugAgent
cd HugAgent/Benchmark
python process_data.py
python evaluate_qwen.py --model your-model

Collect your own

TraceYourThinking, the interview chatbot behind HugAgent, is open source too. Point it at any topic and it collects fine-grained, think-aloud reasoning data: the raw material for benchmarks like this one, on questions we have not thought to ask. github.com/jajamoa/trace-your-thinking

TraceYourThinking: a semi-structured interview chatbot maintains a causal belief graph and generates follow-up questions.

The interview loop, left to right: each answer updates a causal belief network of the participant; the bot reads the graph, finds its most informative node (anchor discovery, then expansion), and asks the follow-up that probes it; the loop repeats for 8 to 20 turns. Click to enlarge.

Citation

If you use HugAgent, please cite the paper. Earlier versions appeared at the NeurIPS 2025 PersonaLLM workshop (Oral) and LAW workshop (Spotlight).

@misc{li2025hugagent,
  title   = {HugAgent: A Human Simulation Benchmark
             for Individual-Level Reasoning},
  author  = {Li, Chance Jiajie and Mo, Zhenze and Tang, Yuhan
             and Qu, Ao and Wu, Jiayi and Zhao, Kaiya Ivy
             and Gan, Yulu and Fan, Jie and Yu, Jiangbo
             and Jiang, Hang and Liang, Paul Pu and Zhao, Jinhua
             and Alonso Pastor, Luis Alberto and Larson, Kent},
  year    = {2025},
  eprint  = {2510.15144},
  archivePrefix = {arXiv},
  note    = {EMNLP 2026 Main}
}
average mode: everyone is 68.4. type “me” to be an individual again.