EMNLP’26 papers

I’m excited to share two papers that will be presented at EMNLP 2026 in Budapest later this month, on user simulation and natural language user profiles.

  • “Act Like a 5th Grader” is Not Enough: Bounding Knowledge in LLM-Based User Simulators (Findings paper, with A. M. Bakken) — LLM-based simulators tend to be “superhuman”: even when prompted to act like a 5th grader, they ace reading comprehension tests. Using over 71k responses from 2,359 Norwegian primary-school students, we propose the Cognitively Bounded User Simulator (CBUS), which models restricted working memory through an episodic bottleneck and substantially reduces the gap to real student behavior.
  • A Survey on Natural Language User Profiles for Recommendation: Methods, Datasets, and Metrics (main conference paper, with M. Arustashvili) — This survey reviews the literature on the two coupled tasks of profile generation and profile-based recommendation, proposes a taxonomy for organizing evaluation metrics, and outlines open challenges, including the tension between transparency and optimization.

I’ll be attending the conference in person. If you’re there, please come say hello!

SIGIR’26 and ICTIR’26 papers

I’m happy to share two papers on conversational recommender systems that were presented at SIGIR and ICTIR this summer. As always, the code and data are publicly available!

  • A Standardized Re-evaluation of Conversational Recommender Systems on the ReDial Dataset (SIGIR reproducibility paper, with I. Kostric) — This work re-evaluates seven prominent CRS methods on ReDial under standardized conditions. We find that nearly half of the reported accuracy comes from “repetition shortcuts,” that performance gains often stem from stronger language model backbones rather than architectural innovations, and that user-centric utility metrics frequently contradict recall-based evaluation. [resources]
  • RecQuest: Towards Estimating User Domain Knowledge in Conversational Recommender Systems (ICTIR full paper, with I. Kostric and U. Gadiraju) — Conversational recommender systems often implicitly treat every user as an expert. This paper introduces the task of estimating user domain knowledge from dialogues, along with RecQuest, a game-with-a-purpose data collection protocol, and a dataset of 515 dialogues across five product domains. [resources]

EACL paper featured on the Google Research Blog

I’m excited to share that our EACL 2026 paper has been featured on the Google Research Blog!

We explore how to move beyond simple performance metrics to ensure simulated users actually behave like real ones and introduce a unique dual-agent data collection protocol that enables counterfactual validation. We also publicly release a new dataset of 4k+ human-AI shopping conversations.

Read the full deep-dive here: https://research.google/blog/convapparel-measuring-and-bridging-the-realism-gap-in-user-simulators/

CACM Opinion piece available online

I’m happy to share that our latest opinion piece, “The Indispensable Role of User Simulation in the Pursuit of AGI,” is now available in Communications of the ACM.

In this article, we argue that the path to Artificial General Intelligence (AGI) is currently blocked by two major bottlenecks: the lack of scalable evaluation and the scarcity of high-quality interaction data. We propose that user simulation is not just a helpful tool, but a critical catalyst for overcoming these challenges.

Read the full piece here: https://cacm.acm.org/opinion/the-indispensable-role-of-user-simulation-in-the-pursuit-of-agi/

EACL’26 and ECIR’26 papers

I’m excited to share some recent research we’ve been doing in the areas of user simulation, recommender systems, and explainability. The following papers will be presented at the upcoming EACL and ECIR conferences. Importantly, all these papers come with publicly available resources!