MobileMem introduces a multimodal, year-scale mobile memory benchmark generated from user priors, and its experiments show that current memory systems reach only 78 to 80 percent on simple recall and far less on temporal and visual reasoning.
Coloragent: Building a robust, personalized, and interactive os agent
3 Pith papers cite this work. Polarity classification is still indexing.
3
Pith papers citing it
representative citing papers
AgentProg reframes interaction history as a program with variables and control flow, plus a belief state for partial observability, achieving SOTA success rates on long-horizon GUI benchmarks while baselines degrade.
GROW decomposes trajectories into state-action samples to enable GRPO for multi-turn VLM agents and reports state-of-the-art results on more than 800 Minecraft tasks.
citing papers explorer
No citing papers match the current filters.