Presents M³Exam benchmark for multimodal conversational memory in user-agent settings and M³Proctor method that raises accuracy 13% while cutting construction time and tokens over 70%.
arXiv preprint arXiv:2510.13291 , year=
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
years
2026 2verdicts
UNVERDICTED 2representative citing papers
PGHS fuses policy-guided LLM reasoning and ML fitting to simulate group user behavior with 8.8% error on Meituan data from 101 merchants and 26k trajectories, beating pure reasoning and fitting baselines by 45.8% and 40.9%.
citing papers explorer
-
M$^3$Exam: Benchmarking Multimodal Memory for Realistic User-Agent Interactions
Presents M³Exam benchmark for multimodal conversational memory in user-agent settings and M³Proctor method that raises accuracy 13% while cutting construction time and tokens over 70%.
-
Meituan Merchant Business Diagnosis via Policy-Guided Dual-Process User Simulation
PGHS fuses policy-guided LLM reasoning and ML fitting to simulate group user behavior with 8.8% error on Meituan data from 101 merchants and 26k trajectories, beating pure reasoning and fitting baselines by 45.8% and 40.9%.