{"total":14,"items":[{"citing_arxiv_id":"2607.05391","ref_index":10,"ref_count":1,"confidence":0.9,"is_internal_anchor":true,"paper_title":"LLM-as-a-Verifier: A General-Purpose Verification Framework","primary_cat":"cs.AI","submitted_at":"2026-07-06T17:59:35+00:00","verdict":"CONDITIONAL","verdict_confidence":"HIGH","novelty_score":5.0,"formal_verification":"none","one_line_summary":"Expecting over scoring-token logits yields continuous, scalable verification that improves agent trajectory selection and dense RL rewards across coding, robotics, and medical benchmarks.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.32027","ref_index":6,"ref_count":1,"confidence":0.9,"is_internal_anchor":true,"paper_title":"Freeform Preference Learning for Robotic Manipulation","primary_cat":"cs.RO","submitted_at":"2026-06-30T17:54:02+00:00","verdict":"CONDITIONAL","verdict_confidence":"MODERATE","novelty_score":6.0,"formal_verification":"none","one_line_summary":"FPL trains a language-conditioned reward model from per-axis human preferences and a reward-conditioned policy, reporting 38-point average success gains over sparse-reward and binary-preference baselines on six manipulation tasks.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.28320","ref_index":36,"ref_count":1,"confidence":0.9,"is_internal_anchor":true,"paper_title":"WARP-RM: A Warp-Augmented Relative Progress Reward Model for Data Curation","primary_cat":"cs.RO","submitted_at":"2026-06-26T17:58:06+00:00","verdict":"CONDITIONAL","verdict_confidence":"MODERATE","novelty_score":6.0,"formal_verification":"none","one_line_summary":"Self-supervised relative progress from time-warped demos reweights BC action chunks, sustaining ~19/20 success and up to ~18× throughput on mixed-quality T-shirt folding where vanilla BC fails.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.24742","ref_index":6,"ref_count":1,"confidence":0.9,"is_internal_anchor":true,"paper_title":"World Value Models for Robotic Manipulation","primary_cat":"cs.RO","submitted_at":"2026-06-23T16:07:48+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":5.0,"formal_verification":"none","one_line_summary":"World Value Model (WVM) integrates world models with value estimation to achieve SOTA Value-Order Correlation on expert and suboptimal robotic data and improves downstream policy performance.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.24633","ref_index":14,"ref_count":1,"confidence":0.9,"is_internal_anchor":true,"paper_title":"Beyond Monotonic Progress: Retry-Supervised Value Learning for Robot Imitation","primary_cat":"cs.RO","submitted_at":"2026-06-23T14:27:10+00:00","verdict":"CONDITIONAL","verdict_confidence":"MODERATE","novelty_score":6.0,"formal_verification":"none","one_line_summary":"Sparse retry keypoints plus pairwise preference learning yield mistake-sensitive values that reweight mixed-quality demos and raise real-robot imitation success over progress-based baselines.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.22027","ref_index":31,"ref_count":1,"confidence":0.9,"is_internal_anchor":true,"paper_title":"RARM: Confidence-Gated Progress Reward Modeling for RL in Manipulation","primary_cat":"cs.RO","submitted_at":"2026-06-20T13:03:21+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"RARM is a lightweight visual comparator trained once on general videos that supplies dense progress rewards to RL by matching rollout clips to a reference demonstration and gating rewards on match confidence.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.13675","ref_index":42,"ref_count":1,"confidence":0.9,"is_internal_anchor":true,"paper_title":"Improving Robotic Generalist Policies via Flow Reversal Steering","primary_cat":"cs.RO","submitted_at":"2026-06-11T17:59:45+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":7.0,"formal_verification":"none","one_line_summary":"Flow Reversal Steering steers flow matching generalist policies by reversing suboptimal actions to nearby better modes, enabling improved zero-shot control, quick distillation, and RL bootstrapping in robotic manipulation.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.10305","ref_index":19,"ref_count":1,"confidence":0.9,"is_internal_anchor":true,"paper_title":"SARM2: Multi-Task Stage Aware Reward Modeling for Self Improving Robotic Manipulation","primary_cat":"cs.RO","submitted_at":"2026-06-09T01:46:23+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"SARM2 presents RM, a multi-task stage-aware reward model achieving 80% lower value-estimation MSE, which when used in SPIRAL boosts manipulation task success from ~50% to near-perfect on several benchmarks.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.00267","ref_index":120,"ref_count":1,"confidence":0.9,"is_internal_anchor":true,"paper_title":"StressDream: Steering Video World Models for Robust Policy Evaluation and Improvement","primary_cat":"cs.CV","submitted_at":"2026-05-29T18:57:57+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"StressDream optimizes initial noise in diffusion video world models using VLM semantic and plausibility objectives to steer generations toward specified high-impact outcomes for improved policy evaluation.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.30257","ref_index":4,"ref_count":1,"confidence":0.9,"is_internal_anchor":true,"paper_title":"Stable-Layers: Fine-Tuning Image Layer Decomposition Models with VLM-Scored Reinforcement Learning","primary_cat":"cs.CV","submitted_at":"2026-05-28T17:20:31+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"Stable-Layers applies Flow-GRPO with LoRA and a two-stage VLM scoring pipeline to improve layer decomposition without paired supervision, yielding stronger separation and lower reconstruction error on Crello.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.22123","ref_index":8,"ref_count":1,"confidence":0.9,"is_internal_anchor":true,"paper_title":"Beyond Pixels: Learning Invariant Rewards for Real-World Robotics From a Few Demonstrations","primary_cat":"cs.RO","submitted_at":"2026-05-21T07:55:35+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"A framework learns invariant symbolic reward functions from few demonstrations that generalize zero-shot to variations in robotic manipulation tasks.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.12369","ref_index":15,"ref_count":2,"confidence":0.9,"is_internal_anchor":true,"paper_title":"GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization","primary_cat":"cs.RO","submitted_at":"2026-05-12T16:38:40+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"GuidedVLA improves VLA generalization by supervising individual attention heads with manually defined auxiliary signals for three task-relevant factors.","context_count":1,"top_context_role":"background","top_context_polarity":"background","context_text":"interpretable and more generalizable. Limitations and Future Work.Our method relies on predefined factors, and automating factor discovery remains an open challenge, especially for continuous tasks where au- tomatic skill labeling is difficult. Promising directions include automatic skill discovery [103, 82] and the use of continuous progress signals as latent skill targets [15]. VII. ACKNOWLEDGMENT This work is supported by the National Natural Science Foundation of China (Grant No. 62521004), the Science and Technology Commission of Shanghai Municipality (No. 24511103100) and the New Cornerstone Science Foundation through the XPLORER PRIZE. This work is also in part sup- ported by Scientific Research Innovation Capability Support"},{"citing_arxiv_id":"2604.11751","ref_index":12,"ref_count":1,"confidence":0.9,"is_internal_anchor":true,"paper_title":"Grounded World Model for Semantically Generalizable Planning","primary_cat":"cs.RO","submitted_at":"2026-04-13T17:25:41+00:00","verdict":"CONDITIONAL","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"A vision-language-aligned world model turns visuomotor MPC into a language-following planner that reaches 87% success on 288 unseen semantic tasks where standard VLAs drop to 22%.","context_count":1,"top_context_role":"background","top_context_polarity":"background","context_text":"a1: Unifying understanding, generation and action for robotic manipulation, 2026. URL https://arxiv.org/abs/2601.02456. [11] Shirui Chen, Cole Harrison, Ying-Chun Lee, Angela Jin Yang, Zhongzheng Ren, Lillian J. Ratliff, Jiafei Duan, Dieter Fox, and Ranjay Krishna. Topreward: Token probabilities as hidden zero-shot rewards for robotics, 2026. URLhttps://arxiv.org/abs/2602.19313. [12] Tianxing Chen, Zanxin Chen, Baijun Chen, Zijian Cai, Yibin Liu, Zixuan Li, Qiwei Liang, Xianliang Lin, Yiheng Ge, Zhenyu Gu, Weiliang Deng, Yubin Guo, Tian Nian, Xuanbing Xie, Qiangyu Chen, Kailun Su, Tianling Xu, Guodong Liu, Mengkang Hu, Huan ang Gao, Kaixuan Wang, Zhixuan Liang, Yusen Qin, Xiaokang Yang, Ping Luo, and Yao Mu. Robotwin 2.0: A scalable data generator and benchmark with strong domain randomization for robust bimanual"},{"citing_arxiv_id":"2603.02115","ref_index":50,"ref_count":1,"confidence":0.9,"is_internal_anchor":true,"paper_title":"Robometer: Scaling General-Purpose Robotic Reward Models via Trajectory Comparisons","primary_cat":"cs.RO","submitted_at":"2026-03-02T17:38:58+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"Robometer combines intra-trajectory progress supervision with inter-trajectory preference supervision on a 1M-trajectory dataset to learn more generalizable robotic reward functions than prior methods.","context_count":1,"top_context_role":"background","top_context_polarity":"background","context_text":"gating the use of video-language models as behavior critics for catching undesirable agent behaviors,\" in Conference on Language Modeling (COLM), 2024. [49] J. Rocamonde, V . Montesinos, E. Nava, E. Perez, and D. Lindner, \"Vision-language models are zero-shot re- ward models for reinforcement learning,\" inInterna- tional Conference on Learning Representations (ICLR), 2024. [50] S. Chen, C. Harrison, Y .-C. Lee, A. J. Yang, Z. Ren, L. J. Ratliff, J. Duan, D. Foxet al., \"Topreward: Token probabilities as hidden zero-shot rewards for robotics,\" arXiv preprint arXiv:2602.19313, 2026. [51] L. Fan, G. Wang, Y . Jiang, A. Mandlekar, Y . Yang, H. Zhu, A. Tang, D.-A. Huanget al., \"Minedojo: Building open-ended embodied agents with internet-"}],"limit":50,"offset":0}