COVERT generates verifiable synthetic tool-use environments for RL by validated trajectory synthesis and oracle-preserving augmentations, improving tool-use accuracy on BFCL v3 and ACEBench while remaining complementary to SFT.
MAG-V: A multi-agent framework for synthetic data generation and verification
3 Pith papers cite this work. Polarity classification is still indexing.
years
2026 3verdicts
UNVERDICTED 3representative citing papers
A-MBER is a new benchmark for evaluating AI models on using interaction history to recognize and explain a user's present affective state across judgment, retrieval, and explanation tasks.
AgentMob is a training-free LLM-driven agent that formulates mobility prediction as adaptive evidence-controlled decision making and outperforms other training-free LLM methods on three datasets.
citing papers explorer
-
Controllable and Verifiable Tool-Use Data Synthesis for Agentic Reinforcement Learning
COVERT generates verifiable synthetic tool-use environments for RL by validated trajectory synthesis and oracle-preserving augmentations, improving tool-use accuracy on BFCL v3 and ACEBench while remaining complementary to SFT.
-
A-MBER: Affective Memory Benchmark for Emotion Recognition
A-MBER is a new benchmark for evaluating AI models on using interaction history to recognize and explain a user's present affective state across judgment, retrieval, and explanation tasks.
-
Towards Efficient and Evidence-grounded Mobility Prediction with LLM-Driven Agent
AgentMob is a training-free LLM-driven agent that formulates mobility prediction as adaptive evidence-controlled decision making and outperforms other training-free LLM methods on three datasets.