TAPO corrects credit misassignment in RL for multimodal search agents by using tool parameter similarity to share advantages across equivalent actions.
arXiv preprint arXiv:2509.06980 , year=
3 Pith papers cite this work. Polarity classification is still indexing.
years
2026 3representative citing papers
VISTAHOP, a 350-question multimodal benchmark, finds the best search agents still fail about 76% of long-horizon visual deep-search tasks.
SEARL uses a tool graph memory that integrates planning and execution to densify rewards and improve generalization in self-evolving agents on knowledge and math tasks.
citing papers explorer
-
TAPO: Tool-Aware Policy Optimization via Credit Transfer for Multimodal Search Agents
TAPO corrects credit misassignment in RL for multimodal search agents by using tool parameter similarity to share advantages across equivalent actions.
-
VistaHop: Benchmarking Long-Horizon Visual DeepSearch
VISTAHOP, a 350-question multimodal benchmark, finds the best search agents still fail about 76% of long-horizon visual deep-search tasks.
-
SEARL: Joint Optimization of Policy and Tool Graph Memory for Self-Evolving Agents
SEARL uses a tool graph memory that integrates planning and execution to densify rewards and improve generalization in self-evolving agents on knowledge and math tasks.