Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-18T21:03:31.606674Z
Paper Citation Record · LEDGER
As of 4 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 41 inbound Pith citation observations for arXiv:2508.19652.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-18T21:03:31.606674Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T18:53:02.707977Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-10T12:15:01.137692Z
30 of 30 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 293fc308-5472-4b24-a02f-8d6ba918037d · outbound
Self-Rewarding Vision-Language Model via Reasoning Decomposition Qwen2.5-VL Technical Report
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a0f187c4-8a2f-4b11-acc3-86415fc8e676 · outbound
Self-Rewarding Vision-Language Model via Reasoning Decomposition Demonstrating and Reducing Shortcuts in Vision-Language Representation Learning
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 27f72ccf-f6e7-47f0-8be2-cb46e75ae1dd · outbound
Self-Rewarding Vision-Language Model via Reasoning Decomposition Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ff9d0d59-9f98-4a59-a587-6a834c107b71 · outbound
Self-Rewarding Vision-Language Model via Reasoning Decomposition InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e1a50a99-75c1-4344-9b2e-db4716f94233 · outbound
Self-Rewarding Vision-Language Model via Reasoning Decomposition DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d009ac87-0d15-4c3b-a2ce-d4a802682aad · outbound
Self-Rewarding Vision-Language Model via Reasoning Decomposition GPT-4 Technical Report
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f1f2f615-d4f8-4f43-a686-c92943506c4b · outbound
Self-Rewarding Vision-Language Model via Reasoning Decomposition Reward shaping to mitigate reward hacking in rlhf
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ef2dc4f0-e265-4679-b6d4-e2dd59127887 · outbound
Self-Rewarding Vision-Language Model via Reasoning Decomposition Attention-Based Reward Shaping for Sparse and Delayed Rewards
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8d7105e9-bd1a-4901-b191-bf3c75b8f53e · outbound
Self-Rewarding Vision-Language Model via Reasoning Decomposition Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation cb772b72-8f4e-40bf-b00a-250d364a23e4 · outbound
Self-Rewarding Vision-Language Model via Reasoning Decomposition Describe-then-Reason: Improving Multimodal Mathematical Reasoning through Visual Comprehension Training
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b0093d23-a9f9-4116-a40b-b88bbc091e5f · outbound
Self-Rewarding Vision-Language Model via Reasoning Decomposition Reward Generation via Large Vision-Language Model in Offline Reinforcement Learning
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7ae09349-ea87-48c8-8ee1-e57bfa4b6b35 · outbound
Self-Rewarding Vision-Language Model via Reasoning Decomposition More Thinking, Less Seeing? Assessing Amplified Hallucination in Multimodal Reasoning Models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b5311bb5-c27b-48c5-8d9b-294a3bd75f35 · outbound
Self-Rewarding Vision-Language Model via Reasoning Decomposition A Survey on Hallucination in Large Vision-Language Models
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5b704c8c-fc4a-4656-9aa9-60b8aa3f330d · outbound
Self-Rewarding Vision-Language Model via Reasoning Decomposition Enhancing Rating-Based Reinforcement Learning to Effectively Leverage Feedback from Large Vision-Language Models
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b4282b8d-24b2-4d02-b756-5486ebf3b871 · outbound
Self-Rewarding Vision-Language Model via Reasoning Decomposition The Bell System Tech- nical Journal27(3), 379–423 (1948) https://doi.org/10.1002/j.1538-7305.1948
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e84b70b4-5be6-47c1-8bf3-377ccab63756 · outbound
Self-Rewarding Vision-Language Model via Reasoning Decomposition Spurious Rewards: Rethinking Training Signals in RLVR
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2bc5ef9e-81fc-4775-99e2-e3b867943254 · outbound
Self-Rewarding Vision-Language Model via Reasoning Decomposition DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0d7a78ad-feb8-40cf-8637-6f2f438b3313 · outbound
Self-Rewarding Vision-Language Model via Reasoning Decomposition Language prior is not the only shortcut: A benchmark for shortcut learning in vqa
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 00c1700a-e9b8-41b0-a2e4-0874cb7e3076 · outbound
Self-Rewarding Vision-Language Model via Reasoning Decomposition RLSR: Reinforcement Learning from Self Reward
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8a3bf153-0a8b-470d-b7ea-f08e298d5ef3 · outbound
Self-Rewarding Vision-Language Model via Reasoning Decomposition Post-Training Large Language Models via Reinforcement Learning from Self-Feedback
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7d132ff0-3378-46e0-82d5-13dc69765dc4 · outbound
Self-Rewarding Vision-Language Model via Reasoning Decomposition Measuring Multimodal Mathematical Reasoning with MATH-Vision Dataset
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b9bf60ec-d0b4-4ef7-a010-e50bf0495250 · outbound
Self-Rewarding Vision-Language Model via Reasoning Decomposition VisNumBench: Evaluating Number Sense of Multimodal Large Language Models
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 74f7f39e-8c50-4111-bd88-6f6e8c4d0c9a · outbound
Self-Rewarding Vision-Language Model via Reasoning Decomposition Perception-R1: Advancing multimodal reasoning capabilities of MLLMs via visual perception reward
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 114e4901-c09a-4c47-a19e-8362aa4a6701 · outbound
Self-Rewarding Vision-Language Model via Reasoning Decomposition Are Reasoning Models More Prone to Hallucination?
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b58b8565-ac9d-418f-baf7-a7c69adfc8d5 · outbound
Self-Rewarding Vision-Language Model via Reasoning Decomposition Self-Rewarding Language Models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 52c95987-247e-43bb-b0ff-d8df4fd4ccb4 · outbound
Self-Rewarding Vision-Language Model via Reasoning Decomposition MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 14aebc15-4c08-429c-a69f-337bf607fea9 · outbound
Self-Rewarding Vision-Language Model via Reasoning Decomposition MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8ac18930-0a43-4dde-8022-f2def321c5e8 · outbound
Self-Rewarding Vision-Language Model via Reasoning Decomposition R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ada79fa5-e828-46fa-b761-45f3a4215ebf · outbound
Self-Rewarding Vision-Language Model via Reasoning Decomposition Learning to Reason without External Rewards
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6c929866-b6dd-492c-98de-e715d255ec40 · outbound
Self-Rewarding Vision-Language Model via Reasoning Decomposition Calibrated Self-Rewarding Vision Language Models
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2eafd2e8-7aed-48ad-a766-753d8f68af19 · inbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Self-Rewarding Vision-Language Model via Reasoning Decomposition
Reference 243
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b36915be-0427-4923-9bf1-971eaa09ee47 · inbound
CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models Self-Rewarding Vision-Language Model via Reasoning Decomposition
Reference 2010
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d84ad3ba-d1df-450b-b707-2ee2ae46435e · inbound
DeFacto: Counterfactual Thinking with Images for Enforcing Evidence-Grounded and Faithful Reasoning Self-Rewarding Vision-Language Model via Reasoning Decomposition
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation bcd44945-aecd-4afd-8a2c-2b81dfa37b58 · inbound
DeFacto: Counterfactual Thinking with Images for Enforcing Evidence-Grounded and Faithful Reasoning Self-Rewarding Vision-Language Model via Reasoning Decomposition
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2c6590ce-6977-40c4-a451-9f9dda7cdd27 · inbound
VIDEOP2R: Video Understanding from Perception to Reasoning Self-Rewarding Vision-Language Model via Reasoning Decomposition
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation cd83fee7-a67b-43c8-a9ed-008127e27ec1 · inbound
EvoLMM: Self-Evolving Large Multimodal Models with Continuous Rewards Self-Rewarding Vision-Language Model via Reasoning Decomposition
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ea4dfa3-11d9-4158-9b50-b17e8074ccdc · inbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Self-Rewarding Vision-Language Model via Reasoning Decomposition
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d62a3ddf-6928-4790-86fb-d5dc65b530b8 · inbound
TANDEM: Temporal-Aware Neural Detection for Multimodal Hate Speech Self-Rewarding Vision-Language Model via Reasoning Decomposition
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b94b329-7fd3-4851-b138-f5d515deb13f · inbound
HART: High-Resolution Annotation-Free Reasoning Technique through a Closed-loop Framework Self-Rewarding Vision-Language Model via Reasoning Decomposition
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81a68ae1-201e-4c4b-aef5-0b97f88d6cca · inbound
Cultivating Forensic Reasoning for Generalizable Multimodal Manipulation Detection Self-Rewarding Vision-Language Model via Reasoning Decomposition
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 514393f7-012e-4e33-abb1-d27bd40f88ef · inbound
SABER: A Stealthy Agentic Black-Box Attack Framework for Vision-Language-Action Models Self-Rewarding Vision-Language Model via Reasoning Decomposition
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 632d624b-3df6-4891-b991-456f302384e8 · inbound
VERTIGO: Visual Preference Optimization for Cinematic Camera Trajectory Generation Self-Rewarding Vision-Language Model via Reasoning Decomposition
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 62c89ad5-fd68-4c0f-af40-27ecefbed93c · inbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Self-Rewarding Vision-Language Model via Reasoning Decomposition
Reference 170
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b2ca0511-8271-4181-8935-f98229babc5f · inbound
EVE: Verifiable Self-Evolution of MLLMs via Executable Visual Transformations Self-Rewarding Vision-Language Model via Reasoning Decomposition
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 4608f550-acd6-40c2-ad73-24abc60bcef4 · inbound
Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data Self-Rewarding Vision-Language Model via Reasoning Decomposition
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 25cba037-6055-41dc-8be2-e261c579c7e2 · inbound
Co-Evolving LLM Decision and Skill Bank Agents for Long-Horizon Tasks Self-Rewarding Vision-Language Model via Reasoning Decomposition
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0b8bdd2f-9b5a-4524-a485-fd9e7bacebc3 · inbound
Pest-Thinker: Learning to Think and Reason like Entomologists via Reinforcement Learning Self-Rewarding Vision-Language Model via Reasoning Decomposition
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 468708aa-c229-4320-8bd0-d3f3ad7d69ed · inbound
Mirror, Mirror on the Wall: Can VLM Agents Tell Who They Are at All? Self-Rewarding Vision-Language Model via Reasoning Decomposition
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 28ba0451-294d-4223-9158-aada18d5ab5d · inbound
Reinforcing Multimodal Reasoning Against Visual Degradation Self-Rewarding Vision-Language Model via Reasoning Decomposition
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ca736119-30ee-4cae-9c82-f376cb704189 · inbound
DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification Self-Rewarding Vision-Language Model via Reasoning Decomposition
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f74b90da-2ddc-45f2-b32c-3af7f623c520 · inbound
PDCR: Perception-Decomposed Confidence Reward for Vision-Language Reasoning Self-Rewarding Vision-Language Model via Reasoning Decomposition
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8e6cd266-c765-47e9-8415-ab40275e0e21 · inbound
Semantic-Enriched Latent Visual Reasoning Self-Rewarding Vision-Language Model via Reasoning Decomposition
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d0de086b-62db-40ae-92d4-842f5e7410dc · inbound
Semantic-Enriched Latent Visual Reasoning Self-Rewarding Vision-Language Model via Reasoning Decomposition
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a6bf4ee8-f21d-4db2-bb37-e88da72484fb · inbound
CaMo: Camera Motion Grounded Evaluation and Training for Vision-Language Models Self-Rewarding Vision-Language Model via Reasoning Decomposition
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 4362425c-f3d1-4614-8545-2e3e350b214b · inbound
RISE: Reliable Improvement in Self-Evolving Vision-Language Models Self-Rewarding Vision-Language Model via Reasoning Decomposition
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9fc96a1f-4e3c-40ab-b7ba-a66b812737d7 · inbound
RISE: Reliable Improvement in Self-Evolving Vision-Language Models Self-Rewarding Vision-Language Model via Reasoning Decomposition
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b8713252-d2d9-4d04-97b3-73766edbde3a · inbound
Visual-Advantage On-Policy Distillation for Vision-Language Models Self-Rewarding Vision-Language Model via Reasoning Decomposition
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 32aa04c6-5614-4ca7-a836-7414863dfabe · inbound
EvoVid: Temporal-Centric Self-Evolution for Video Large Language Models Self-Rewarding Vision-Language Model via Reasoning Decomposition
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7506bc9c-35d5-4bff-82ab-9a4724f08313 · inbound
Bridging the Detection-to-Abstention Gap in Reasoning Models under Insufficient Information Self-Rewarding Vision-Language Model via Reasoning Decomposition
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 03f0b52c-3524-409b-8772-9488da3ff4d8 · inbound
Stable-Layers: Fine-Tuning Image Layer Decomposition Models with VLM-Scored Reinforcement Learning Self-Rewarding Vision-Language Model via Reasoning Decomposition
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 44e5e864-636f-4cb6-aac4-f912c2b5abe2 · inbound
Decomposed On-Policy Distillation for Vision-Language Reasoning: Steering Gradients for Visual Grounding Self-Rewarding Vision-Language Model via Reasoning Decomposition
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation dd3849ca-ca49-4b81-ac7e-096a3b9bdfd2 · inbound
Scaling Participation in Modular AI Systems Self-Rewarding Vision-Language Model via Reasoning Decomposition
Reference 119
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 1afc69ce-5bca-4bf7-88dd-56d96514fbef · inbound
Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization Self-Rewarding Vision-Language Model via Reasoning Decomposition
Reference 284
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7b7ec798-908f-4efb-a064-20eabcfc5268 · inbound
Improving Reasoning in Vision-Language Models via Perception Verified Self-Training Self-Rewarding Vision-Language Model via Reasoning Decomposition
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation eb47cb2f-387e-4298-b916-725573a4bb5c · inbound
Personalizing MLLMs via Reinforced Multimodal Reference Game Self-Rewarding Vision-Language Model via Reasoning Decomposition
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3c179be6-3bc0-407a-a340-25139bc23a3f · inbound
Token-Sparse Medical Multimodal Reasoning via Dual-Stream Reinforcement Learning Self-Rewarding Vision-Language Model via Reasoning Decomposition
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a7aa0d15-ecfb-42d8-b703-5f3bd601bbc6 · inbound
LASER: A Corrective Lens for LVLMs via Visual Attention Preservation and Sink Suppression Self-Rewarding Vision-Language Model via Reasoning Decomposition
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3158a5c5-bf95-4637-8211-a8dc6661c89f · inbound
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Self-Rewarding Vision-Language Model via Reasoning Decomposition
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a164a77-f953-4015-a040-5be20571ce00 · inbound
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Self-Rewarding Vision-Language Model via Reasoning Decomposition
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3e0ded3-08e6-48e1-aa33-eefd0bad2c28 · inbound
Med-OPD: Improving Medical Vision-Language Models via Evidence-Aware On-Policy Distillation Self-Rewarding Vision-Language Model via Reasoning Decomposition
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f60e0dd2-6991-47d1-977c-6a66da0f9a18 · inbound
RP-OPSD: Resolution-Privileged On-Policy Self-Distillation for Multimodal Large Language Models Self-Rewarding Vision-Language Model via Reasoning Decomposition
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.