A new benchmark with real, time-ordered social-media posts shows all 14 tested LLMs over-retain past interests and miss emerging ones, with the best model reaching only ~52% recall.
InProceedings of the 34th ACM Inter- national Conference on Information and Knowledge Management, pages 5898–5906
2 Pith papers cite this work. Polarity classification is still indexing.
years
2026 2representative citing papers
UserGPT introduces a generative LLM framework with a behavior simulation engine, semantization module, and DF-GRPO post-training that scores 0.7325 on tag prediction and 0.7528 on summary generation on HPR-Bench while compressing records by up to 97.9%.
citing papers explorer
-
StreamProfileBench: A Benchmark for Fine-Grained User Profile Inference in Real-World Streaming Scenarios
A new benchmark with real, time-ordered social-media posts shows all 14 tested LLMs over-retain past interests and miss emerging ones, with the best model reaching only ~52% recall.
-
UserGPT Technical Report
UserGPT introduces a generative LLM framework with a behavior simulation engine, semantization module, and DF-GRPO post-training that scores 0.7325 on tag prediction and 0.7528 on summary generation on HPR-Bench while compressing records by up to 97.9%.