Pith. sign in

Paper Citation Record · LEDGER

Self-Rewarding Vision-Language Model via Reasoning Decomposition

As of 4 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 41 inbound Pith citation observations for arXiv:2508.19652.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.19652 v2

Coverage vector

measured 30 of 30 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-18T21:03:31.606674Z

measured 71 of 71 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 41 of 41 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T18:53:02.707977Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T12:15:01.137692Z

Reference resolution

30 of 30 outbound references displayed

  • verified exact12
  • verified fuzzy1
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch17

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 293fc308-5472-4b24-a02f-8d6ba918037d · outbound

This paper cites Qwen2.5-VL Technical Report.

Self-Rewarding Vision-Language Model via Reasoning Decomposition Qwen2.5-VL Technical Report

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T21:06:50.854009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T21:03:31.606674Z digest=sha256:e54772a1fcd6393081e842a955957e7270ee7514232ce7bac9552e4fecf0cfb1

Observation a0f187c4-8a2f-4b11-acc3-86415fc8e676 · outbound

This paper cites Demonstrating and Reducing Shortcuts in Vision-Language Representation Learning.

Self-Rewarding Vision-Language Model via Reasoning Decomposition Demonstrating and Reducing Shortcuts in Vision-Language Representation Learning

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-18T21:06:50.894430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T21:03:31.606674Z digest=sha256:8a6fa26dc0fff05f875d0257ebfd0488e78e08aac9b2eb89d796c7e3dcbbfc10

Observation 27f72ccf-f6e7-47f0-8be2-cb46e75ae1dd · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

Self-Rewarding Vision-Language Model via Reasoning Decomposition Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-18T21:06:50.842818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T21:03:31.606674Z digest=sha256:8c0e62b002c54a27f833e6c44bd1bbcd2e28779058e77ba6297bec3b86328497

Observation ff9d0d59-9f98-4a59-a587-6a834c107b71 · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

Self-Rewarding Vision-Language Model via Reasoning Decomposition InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T21:06:50.859408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T21:03:31.606674Z digest=sha256:39c348e14726d8b7c6ca220783d747e3b648c36851e71428b532f2cf0c6774ed

Observation e1a50a99-75c1-4344-9b2e-db4716f94233 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Self-Rewarding Vision-Language Model via Reasoning Decomposition DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-18T21:06:50.904071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T21:03:31.606674Z digest=sha256:09f8681ca7435ff1a1009af828784d67215470f08df0d21ab3a011540a4d2140

Observation d009ac87-0d15-4c3b-a2ce-d4a802682aad · outbound

This paper cites GPT-4 Technical Report.

Self-Rewarding Vision-Language Model via Reasoning Decomposition GPT-4 Technical Report

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T21:06:50.819432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T21:03:31.606674Z digest=sha256:b6649b27adf10fe300cd76f3f2e7de169005317e4322ebf3197fbf880ddbb852

Observation f1f2f615-d4f8-4f43-a686-c92943506c4b · outbound

This paper cites Reward shaping to mitigate reward hacking in rlhf.

Self-Rewarding Vision-Language Model via Reasoning Decomposition Reward shaping to mitigate reward hacking in rlhf

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T21:06:50.884482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T21:03:31.606674Z digest=sha256:cb3759d2d438a92e5c7eb500ccd74b12ec20709118f4354c0bb1991e06488c66

Observation ef2dc4f0-e265-4679-b6d4-e2dd59127887 · outbound

This paper cites Attention-Based Reward Shaping for Sparse and Delayed Rewards.

Self-Rewarding Vision-Language Model via Reasoning Decomposition Attention-Based Reward Shaping for Sparse and Delayed Rewards

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-18T21:06:50.924903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T21:03:31.606674Z digest=sha256:3a30a5a5986af7ca0d05bdbcf8bf82f1118445365e1998915b290d7a9818fd89

Observation 8d7105e9-bd1a-4901-b191-bf3c75b8f53e · outbound

This paper cites Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models.

Self-Rewarding Vision-Language Model via Reasoning Decomposition Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T21:06:50.870094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T21:03:31.606674Z digest=sha256:d70fe60a5b217e4bfb6b1cb09c1e01f0b9f98710af1e320b0d0061313672021d

Observation cb772b72-8f4e-40bf-b00a-250d364a23e4 · outbound

This paper cites Describe-then-Reason: Improving Multimodal Mathematical Reasoning through Visual Comprehension Training.

Self-Rewarding Vision-Language Model via Reasoning Decomposition Describe-then-Reason: Improving Multimodal Mathematical Reasoning through Visual Comprehension Training

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T21:06:50.889155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T21:03:31.606674Z digest=sha256:cf3912cee359488f560068d3de27b19923c7e5e01cb0c3c4d3df9c8e3145f083

Observation b0093d23-a9f9-4116-a40b-b88bbc091e5f · outbound

This paper cites Reward Generation via Large Vision-Language Model in Offline Reinforcement Learning.

Self-Rewarding Vision-Language Model via Reasoning Decomposition Reward Generation via Large Vision-Language Model in Offline Reinforcement Learning

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-18T21:06:50.909122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T21:03:31.606674Z digest=sha256:7bbeda2816f400321d768e959fea00089162e102ec3cc588cd4fcf8cbb462979

Observation 7ae09349-ea87-48c8-8ee1-e57bfa4b6b35 · outbound

This paper cites More Thinking, Less Seeing? Assessing Amplified Hallucination in Multimodal Reasoning Models.

Self-Rewarding Vision-Language Model via Reasoning Decomposition More Thinking, Less Seeing? Assessing Amplified Hallucination in Multimodal Reasoning Models

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T21:06:50.807141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T21:03:31.606674Z digest=sha256:ec3aeecf4389fa5233b2d222b725b9ff5bf72236a2c0ce66856023caa913f373

Observation b5311bb5-c27b-48c5-8d9b-294a3bd75f35 · outbound

This paper cites A Survey on Hallucination in Large Vision-Language Models.

Self-Rewarding Vision-Language Model via Reasoning Decomposition A Survey on Hallucination in Large Vision-Language Models

Reference 13

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T21:06:50.874660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T21:03:31.606674Z digest=sha256:669d68a8ba065f83003a5b84e050ff7d0fa8890b6d46a5e62141657f300dc927

Observation 5b704c8c-fc4a-4656-9aa9-60b8aa3f330d · outbound

This paper cites Enhancing Rating-Based Reinforcement Learning to Effectively Leverage Feedback from Large Vision-Language Models.

Self-Rewarding Vision-Language Model via Reasoning Decomposition Enhancing Rating-Based Reinforcement Learning to Effectively Leverage Feedback from Large Vision-Language Models

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-18T21:06:50.849423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T21:03:31.606674Z digest=sha256:ab1bce1c784fb4f8525933f9e664a382c15f940f771bb2f3a14df3b24c7a1e3e

Observation b4282b8d-24b2-4d02-b756-5486ebf3b871 · outbound

This paper cites The Bell System Tech- nical Journal27(3), 379–423 (1948) https://doi.org/10.1002/j.1538-7305.1948.

Self-Rewarding Vision-Language Model via Reasoning Decomposition The Bell System Tech- nical Journal27(3), 379–423 (1948) https://doi.org/10.1002/j.1538-7305.1948

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T21:06:50.351996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T21:03:31.606674Z digest=sha256:cd09fc48296f09b28944a3f896ff219698c7bc26c578c9c24eb1026be1a23fa3

Observation e84b70b4-5be6-47c1-8bf3-377ccab63756 · outbound

This paper cites Spurious Rewards: Rethinking Training Signals in RLVR.

Self-Rewarding Vision-Language Model via Reasoning Decomposition Spurious Rewards: Rethinking Training Signals in RLVR

Reference 16

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T21:06:50.837840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T21:03:31.606674Z digest=sha256:751f9bfdd9e49315d01b4492a8d274d2f2e5e3933e14a7f7e8490b7b4ce173f6

Observation 2bc5ef9e-81fc-4775-99e2-e3b867943254 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Self-Rewarding Vision-Language Model via Reasoning Decomposition DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 17

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T21:06:50.913499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T21:03:31.606674Z digest=sha256:74ce158de8cc560e9862b6fc3082784ac427c1e8aea8fb2d75fdca3baf7585b7

Observation 0d7a78ad-feb8-40cf-8637-6f2f438b3313 · outbound

This paper cites Language prior is not the only shortcut: A benchmark for shortcut learning in vqa.

Self-Rewarding Vision-Language Model via Reasoning Decomposition Language prior is not the only shortcut: A benchmark for shortcut learning in vqa

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T21:11:51.929614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T21:03:31.606674Z digest=sha256:98b4d5f8bb1a6372ad00e6a37b265fd20dc5a7fb524fd5cf8e5730057bc9e25e

Observation 00c1700a-e9b8-41b0-a2e4-0874cb7e3076 · outbound

This paper cites RLSR: Reinforcement Learning from Self Reward.

Self-Rewarding Vision-Language Model via Reasoning Decomposition RLSR: Reinforcement Learning from Self Reward

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-18T21:06:50.865244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T21:03:31.606674Z digest=sha256:37ae6e7a56f5e9e7c9f42e68c40de40057aa5f71c2d7f8b3d8629f6eb429d957

Observation 8a3bf153-0a8b-470d-b7ea-f08e298d5ef3 · outbound

This paper cites Post-Training Large Language Models via Reinforcement Learning from Self-Feedback.

Self-Rewarding Vision-Language Model via Reasoning Decomposition Post-Training Large Language Models via Reinforcement Learning from Self-Feedback

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-18T21:06:50.793278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T21:03:31.606674Z digest=sha256:0f24c7ab51db46c4c0352fce707a8265d4549c83bc3156164c09d164b83bd5ee

Observation 7d132ff0-3378-46e0-82d5-13dc69765dc4 · outbound

This paper cites Measuring Multimodal Mathematical Reasoning with MATH-Vision Dataset.

Self-Rewarding Vision-Language Model via Reasoning Decomposition Measuring Multimodal Mathematical Reasoning with MATH-Vision Dataset

Reference 21

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T21:06:50.824810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T21:03:31.606674Z digest=sha256:68a1893c126a5f89053cecdc9c73b5008d574af1c4e0162b22f98d7f0f2a5466

Observation b9bf60ec-d0b4-4ef7-a010-e50bf0495250 · outbound

This paper cites VisNumBench: Evaluating Number Sense of Multimodal Large Language Models.

Self-Rewarding Vision-Language Model via Reasoning Decomposition VisNumBench: Evaluating Number Sense of Multimodal Large Language Models

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-18T21:06:50.831375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T21:03:31.606674Z digest=sha256:f08ad33bc3ecc2a51a560f9960cb31f3c8c3394e456bfcaa072cbe8acc9cabc7

Observation 74f7f39e-8c50-4111-bd88-6f6e8c4d0c9a · outbound

This paper cites Perception-R1: Advancing multimodal reasoning capabilities of MLLMs via visual perception reward.

Self-Rewarding Vision-Language Model via Reasoning Decomposition Perception-R1: Advancing multimodal reasoning capabilities of MLLMs via visual perception reward

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T21:06:50.780707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T21:03:31.606674Z digest=sha256:8d5ce3f8ddebdc2873315b4da5dfe3265445abc233b31052737c2e286f9127f8

Observation 114e4901-c09a-4c47-a19e-8362aa4a6701 · outbound

This paper cites Are Reasoning Models More Prone to Hallucination?.

Self-Rewarding Vision-Language Model via Reasoning Decomposition Are Reasoning Models More Prone to Hallucination?

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-18T21:06:50.899623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T21:03:31.606674Z digest=sha256:f068ef24c47a92ba762c2d16fb8be01cbbae8e417af346aa233c091362bc5fac

Observation b58b8565-ac9d-418f-baf7-a7c69adfc8d5 · outbound

This paper cites Self-Rewarding Language Models.

Self-Rewarding Vision-Language Model via Reasoning Decomposition Self-Rewarding Language Models

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-18T21:06:50.919220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T21:03:31.606674Z digest=sha256:8911e91eb4bc1fa140a6de81c5c9b2fe7e4a1672c1b7bf5dbd516032dc5dc57d

Observation 52c95987-247e-43bb-b0ff-d8df4fd4ccb4 · outbound

This paper cites MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI.

Self-Rewarding Vision-Language Model via Reasoning Decomposition MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI

Reference 26

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T21:06:50.787426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T21:03:31.606674Z digest=sha256:9efd1113258a3430802f2cae5e6c4a199f9f648d9c44837d2ae18f7e2ec7524c

Observation 14aebc15-4c08-429c-a69f-337bf607fea9 · outbound

This paper cites MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark.

Self-Rewarding Vision-Language Model via Reasoning Decomposition MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark

Reference 27

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T21:06:50.800944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T21:03:31.606674Z digest=sha256:fb5cf6b217756ebd4031094d0fc150831442be2bd1006758b985bf353860a89f

Observation 8ac18930-0a43-4dde-8022-f2def321c5e8 · outbound

This paper cites R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization.

Self-Rewarding Vision-Language Model via Reasoning Decomposition R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization

Reference 28

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T21:06:50.812882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T21:03:31.606674Z digest=sha256:663e4b69004942f32ea3ba068355fb6304ac1bd3744fa79bc4258dbacd337e86

Observation ada79fa5-e828-46fa-b761-45f3a4215ebf · outbound

This paper cites Learning to Reason without External Rewards.

Self-Rewarding Vision-Language Model via Reasoning Decomposition Learning to Reason without External Rewards

Reference 29

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T21:06:50.879153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T21:03:31.606674Z digest=sha256:6ddbceb810d81eb05d4614457c26d28afe8f5dd68fcede7326d48dbbbd7f24b2

Observation 6c929866-b6dd-492c-98de-e715d255ec40 · outbound

This paper cites Calibrated Self-Rewarding Vision Language Models.

Self-Rewarding Vision-Language Model via Reasoning Decomposition Calibrated Self-Rewarding Vision Language Models

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-18T21:06:50.930423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T21:03:31.606674Z digest=sha256:282de0082ac9c06840dded8b89758504b2ed675ede2fb301fb1861cbc896ca81

Pith citing papers

Observation 2eafd2e8-7aed-48ad-a766-753d8f68af19 · inbound

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey cites this paper.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Self-Rewarding Vision-Language Model via Reasoning Decomposition

Reference 243

Resolution
verified exact
local_arxiv, observed 2026-05-18T19:21:48.797520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:38f10bdda9a04202cb98357929f695273393ee4a3231ea21a3a7ea4b57c0c052

Observation b36915be-0427-4923-9bf1-971eaa09ee47 · inbound

CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models cites this paper.

CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models Self-Rewarding Vision-Language Model via Reasoning Decomposition

Reference 2010

Resolution
unresolved
no resolver link, observed 2026-08-04T18:53:02.707977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:53:02.707977Z digest=sha256:f7055bb0f3779c1a060621b9357925cd51fd22f8b6033fa319915fdd713662de

Observation d84ad3ba-d1df-450b-b707-2ee2ae46435e · inbound

DeFacto: Counterfactual Thinking with Images for Enforcing Evidence-Grounded and Faithful Reasoning cites this paper.

DeFacto: Counterfactual Thinking with Images for Enforcing Evidence-Grounded and Faithful Reasoning Self-Rewarding Vision-Language Model via Reasoning Decomposition

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-21T22:35:42.650352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T22:35:36.136639Z digest=sha256:af22ae8ba3db53dbc3b630f59f65320047b69b309ec8eb33b4ab16017d0830bb

Observation bcd44945-aecd-4afd-8a2c-2b81dfa37b58 · inbound

DeFacto: Counterfactual Thinking with Images for Enforcing Evidence-Grounded and Faithful Reasoning cites this paper.

DeFacto: Counterfactual Thinking with Images for Enforcing Evidence-Grounded and Faithful Reasoning Self-Rewarding Vision-Language Model via Reasoning Decomposition

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-22T13:36:35.958576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T13:36:06.229352Z digest=sha256:11f67d46d98503893998f576bfdfd061ac729953a8fc998a098a1262e176f589

Observation 2c6590ce-6977-40c4-a451-9f9dda7cdd27 · inbound

VIDEOP2R: Video Understanding from Perception to Reasoning cites this paper.

VIDEOP2R: Video Understanding from Perception to Reasoning Self-Rewarding Vision-Language Model via Reasoning Decomposition

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-17T22:25:22.650022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T22:24:41.760120Z digest=sha256:9ebe3a9991ec5c50b8ee89062cf6b4a442738968f7411a477928e74ef707e5d4

Observation cd83fee7-a67b-43c8-a9ed-008127e27ec1 · inbound

EvoLMM: Self-Evolving Large Multimodal Models with Continuous Rewards cites this paper.

EvoLMM: Self-Evolving Large Multimodal Models with Continuous Rewards Self-Rewarding Vision-Language Model via Reasoning Decomposition

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T21:09:23.052942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:09:23.052942Z digest=sha256:7107ec6a984e6f778d3a533c0ffd0c0b0ed56ff3bf89756d326361d451120899

Observation 9ea4dfa3-11d9-4158-9b50-b17e8074ccdc · inbound

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs cites this paper.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Self-Rewarding Vision-Language Model via Reasoning Decomposition

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:04.725008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:04.725008Z digest=sha256:279da520074439caf6637b65dce71d46b460ba2b207153a168763296a4bd5c7f

Observation d62a3ddf-6928-4790-86fb-d5dc65b530b8 · inbound

TANDEM: Temporal-Aware Neural Detection for Multimodal Hate Speech cites this paper.

TANDEM: Temporal-Aware Neural Detection for Multimodal Hate Speech Self-Rewarding Vision-Language Model via Reasoning Decomposition

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T10:10:54.216618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T10:10:54.216618Z digest=sha256:988f66192c8b985929273a06ef4112493b16ca38d30cbdb5d007e3b36ba23c47

Observation 1b94b329-7fd3-4851-b138-f5d515deb13f · inbound

HART: High-Resolution Annotation-Free Reasoning Technique through a Closed-loop Framework cites this paper.

HART: High-Resolution Annotation-Free Reasoning Technique through a Closed-loop Framework Self-Rewarding Vision-Language Model via Reasoning Decomposition

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T20:19:14.531269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:19:14.531269Z digest=sha256:59ff3aed3b765bd246beba24ba1c1b05eb8991a8f5c24fbd4181f37995a3c5d1

Observation 81a68ae1-201e-4c4b-aef5-0b97f88d6cca · inbound

Cultivating Forensic Reasoning for Generalizable Multimodal Manipulation Detection cites this paper.

Cultivating Forensic Reasoning for Generalizable Multimodal Manipulation Detection Self-Rewarding Vision-Language Model via Reasoning Decomposition

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-21T11:20:02.501807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T11:16:30.368440Z digest=sha256:eb15d22fafa111f7123a59b750150d106351c980967b15ab676caeacb944b6a0

Observation 514393f7-012e-4e33-abb1-d27bd40f88ef · inbound

SABER: A Stealthy Agentic Black-Box Attack Framework for Vision-Language-Action Models cites this paper.

SABER: A Stealthy Agentic Black-Box Attack Framework for Vision-Language-Action Models Self-Rewarding Vision-Language Model via Reasoning Decomposition

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-15T01:08:25.727193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T01:06:42.421062Z digest=sha256:21072b41f971354717af14a79a235bbb1cadf9aafc5c8df6af3ce5507323dfa1

Observation 632d624b-3df6-4891-b991-456f302384e8 · inbound

VERTIGO: Visual Preference Optimization for Cinematic Camera Trajectory Generation cites this paper.

VERTIGO: Visual Preference Optimization for Cinematic Camera Trajectory Generation Self-Rewarding Vision-Language Model via Reasoning Decomposition

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-13T21:18:17.211861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T21:14:42.021240Z digest=sha256:42e322f6afb0647a71c06ea1d4d57038e8ad042b18c5b8840e8c6ec83fe7f746

Observation 62c89ad5-fd68-4c0f-af40-27ecefbed93c · inbound

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges cites this paper.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Self-Rewarding Vision-Language Model via Reasoning Decomposition

Reference 170

Resolution
verified exact
local_arxiv, observed 2026-05-10T14:00:28.148434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:3aea157c3cee12483af04874ca06241f5186c1ba8fa9a7b46e2c8f4a44802ecb

Observation b2ca0511-8271-4181-8935-f98229babc5f · inbound

EVE: Verifiable Self-Evolution of MLLMs via Executable Visual Transformations cites this paper.

EVE: Verifiable Self-Evolution of MLLMs via Executable Visual Transformations Self-Rewarding Vision-Language Model via Reasoning Decomposition

Reference 24

Resolution
metadata mismatch
local_arxiv, observed 2026-05-10T05:25:54.882531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T05:23:45.814042Z digest=sha256:c256b0d2ecedaf123c83bcdbe247824a3ab86fc25f538d25addccbad9e561315

Observation 4608f550-acd6-40c2-ad73-24abc60bcef4 · inbound

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data cites this paper.

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data Self-Rewarding Vision-Language Model via Reasoning Decomposition

Reference 44

Resolution
metadata mismatch
local_arxiv, observed 2026-05-10T05:36:02.006979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T05:32:23.972335Z digest=sha256:35166993c47edfb1aa39152af9e3bce85b2be1faf30beadde4d87c2f5cbd9e53

Observation 25cba037-6055-41dc-8be2-e261c579c7e2 · inbound

Co-Evolving LLM Decision and Skill Bank Agents for Long-Horizon Tasks cites this paper.

Co-Evolving LLM Decision and Skill Bank Agents for Long-Horizon Tasks Self-Rewarding Vision-Language Model via Reasoning Decomposition

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-05-10T00:14:46.441646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:14:07.017420Z digest=sha256:a5f7dbb1f169153d5641fc8bcb1c28d48ec8781654438e4686f5dfa454831a7a

Observation 0b8bdd2f-9b5a-4524-a485-fd9e7bacebc3 · inbound

Pest-Thinker: Learning to Think and Reason like Entomologists via Reinforcement Learning cites this paper.

Pest-Thinker: Learning to Think and Reason like Entomologists via Reinforcement Learning Self-Rewarding Vision-Language Model via Reasoning Decomposition

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-11T18:46:09.475805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T14:02:29.480442Z digest=sha256:40bad4b8e2f2f265ea595787f15155855724214449eca6579d65474722dba963

Observation 468708aa-c229-4320-8bd0-d3f3ad7d69ed · inbound

Mirror, Mirror on the Wall: Can VLM Agents Tell Who They Are at All? cites this paper.

Mirror, Mirror on the Wall: Can VLM Agents Tell Who They Are at All? Self-Rewarding Vision-Language Model via Reasoning Decomposition

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T03:36:18.582096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T03:35:39.884350Z digest=sha256:302082d0df825f4265bfe352dc06322a3da7aae6d1f57ea5dca2a39fdeec09ab

Observation 28ba0451-294d-4223-9158-aada18d5ab5d · inbound

Reinforcing Multimodal Reasoning Against Visual Degradation cites this paper.

Reinforcing Multimodal Reasoning Against Visual Degradation Self-Rewarding Vision-Language Model via Reasoning Decomposition

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-12T06:06:25.272467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T04:37:21.451146Z digest=sha256:48ce5cb59878d356b9ddb5462817790bb41e7221aeb1c9ec2d92f6591b434bcc

Observation ca736119-30ee-4cae-9c82-f376cb704189 · inbound

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification cites this paper.

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification Self-Rewarding Vision-Language Model via Reasoning Decomposition

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-12T06:01:24.251497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T04:41:44.833354Z digest=sha256:9e3c5e2908527e378e5aac234b1cf0b116eefc5983b67496d248893b5cef2cbd

Observation f74b90da-2ddc-45f2-b32c-3af7f623c520 · inbound

PDCR: Perception-Decomposed Confidence Reward for Vision-Language Reasoning cites this paper.

PDCR: Perception-Decomposed Confidence Reward for Vision-Language Reasoning Self-Rewarding Vision-Language Model via Reasoning Decomposition

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-14T19:22:50.506455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T19:20:32.435135Z digest=sha256:a87b358a24393b360149a875126211beeace753e71b9ac30803abc092770a8e8

Observation 8e6cd266-c765-47e9-8415-ab40275e0e21 · inbound

Semantic-Enriched Latent Visual Reasoning cites this paper.

Semantic-Enriched Latent Visual Reasoning Self-Rewarding Vision-Language Model via Reasoning Decomposition

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-20T06:43:05.841371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T06:40:55.537488Z digest=sha256:2631a50dcf77e9969980f6c7fc3409662253c095ba81f22f3cdc7af14b519a1d

Observation d0de086b-62db-40ae-92d4-842f5e7410dc · inbound

Semantic-Enriched Latent Visual Reasoning cites this paper.

Semantic-Enriched Latent Visual Reasoning Self-Rewarding Vision-Language Model via Reasoning Decomposition

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-06-30T18:55:00.783814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T18:48:04.370230Z digest=sha256:2df941795a161bb1f5b68229d853155e593a4065dfb0b3e1448dffe4798b8374

Observation a6bf4ee8-f21d-4db2-bb37-e88da72484fb · inbound

CaMo: Camera Motion Grounded Evaluation and Training for Vision-Language Models cites this paper.

CaMo: Camera Motion Grounded Evaluation and Training for Vision-Language Models Self-Rewarding Vision-Language Model via Reasoning Decomposition

Reference 22

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T05:28:04.551335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-20T05:27:30.938311Z digest=sha256:cddae767104fdb217b8dee60dbf0a4af68077d2e867646e01af600ab4690b15c

Observation 4362425c-f3d1-4614-8545-2e3e350b214b · inbound

RISE: Reliable Improvement in Self-Evolving Vision-Language Models cites this paper.

RISE: Reliable Improvement in Self-Evolving Vision-Language Models Self-Rewarding Vision-Language Model via Reasoning Decomposition

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-21T05:39:40.542640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T05:38:26.590720Z digest=sha256:063bc30c7970bfd07017bf88ddbd16a8685afbbf90989afe00a82b1d54697624

Observation 9fc96a1f-4e3c-40ab-b7ba-a66b812737d7 · inbound

RISE: Reliable Improvement in Self-Evolving Vision-Language Models cites this paper.

RISE: Reliable Improvement in Self-Evolving Vision-Language Models Self-Rewarding Vision-Language Model via Reasoning Decomposition

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-06-30T17:24:57.576003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T17:19:09.596950Z digest=sha256:b78b32c6c00c5a5e16db58b503509bdb9189946affcecc6c7350062e6edf094e

Observation b8713252-d2d9-4d04-97b3-73766edbde3a · inbound

Visual-Advantage On-Policy Distillation for Vision-Language Models cites this paper.

Visual-Advantage On-Policy Distillation for Vision-Language Models Self-Rewarding Vision-Language Model via Reasoning Decomposition

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-22T07:24:42.889722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T07:24:22.355378Z digest=sha256:256cc7330c2246787912bc959b71893ef14dfc0125633ad4065794ab2e5fa97f

Observation 32aa04c6-5614-4ca7-a836-7414863dfabe · inbound

EvoVid: Temporal-Centric Self-Evolution for Video Large Language Models cites this paper.

EvoVid: Temporal-Centric Self-Evolution for Video Large Language Models Self-Rewarding Vision-Language Model via Reasoning Decomposition

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-22T07:21:12.876596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T07:19:30.508843Z digest=sha256:a3eb3165e60546cd5b9f9101569857cee8a05afa1c930a0918e47ce61b523a7a

Observation 7506bc9c-35d5-4bff-82ab-9a4724f08313 · inbound

Bridging the Detection-to-Abstention Gap in Reasoning Models under Insufficient Information cites this paper.

Bridging the Detection-to-Abstention Gap in Reasoning Models under Insufficient Information Self-Rewarding Vision-Language Model via Reasoning Decomposition

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-06-29T12:43:25.856065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T12:37:16.138816Z digest=sha256:2722840ab00f54f44c05862c4854fae0647cf06edecc1ca155671939eefd1173

Observation 03f0b52c-3524-409b-8772-9488da3ff4d8 · inbound

Stable-Layers: Fine-Tuning Image Layer Decomposition Models with VLM-Scored Reinforcement Learning cites this paper.

Stable-Layers: Fine-Tuning Image Layer Decomposition Models with VLM-Scored Reinforcement Learning Self-Rewarding Vision-Language Model via Reasoning Decomposition

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-06-29T08:33:15.656607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T08:26:06.346326Z digest=sha256:378cf59bc4d2839d8407bcd78d629d197176c5e8d256c33d1823f9290e1589d6

Observation 44e5e864-636f-4cb6-aac4-f912c2b5abe2 · inbound

Decomposed On-Policy Distillation for Vision-Language Reasoning: Steering Gradients for Visual Grounding cites this paper.

Decomposed On-Policy Distillation for Vision-Language Reasoning: Steering Gradients for Visual Grounding Self-Rewarding Vision-Language Model via Reasoning Decomposition

Reference 16

Resolution
metadata mismatch
local_arxiv, observed 2026-06-28T19:32:34.996398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-28T19:26:37.281563Z digest=sha256:c68ebcbb8b7e36e2a13263f68da4a7ab82bc1293ac794f519ef1a9805ccabac4

Observation dd3849ca-ca49-4b81-ac7e-096a3b9bdfd2 · inbound

Scaling Participation in Modular AI Systems cites this paper.

Scaling Participation in Modular AI Systems Self-Rewarding Vision-Language Model via Reasoning Decomposition

Reference 119

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T18:57:16.640238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-27T21:49:27.042616Z digest=sha256:b4c37072a9fbb72259728d92ead7ee0dee625bde16744512829e528e61e11885

Observation 1afc69ce-5bca-4bf7-88dd-56d96514fbef · inbound

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization cites this paper.

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization Self-Rewarding Vision-Language Model via Reasoning Decomposition

Reference 284

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T01:37:30.255619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-27T16:26:34.918099Z digest=sha256:f1c8da41e4d1ef8415d6756c5f6dc6e80b3510e4edc45b7bf27f5d866185b3f7

Observation 7b7ec798-908f-4efb-a064-20eabcfc5268 · inbound

Improving Reasoning in Vision-Language Models via Perception Verified Self-Training cites this paper.

Improving Reasoning in Vision-Language Models via Perception Verified Self-Training Self-Rewarding Vision-Language Model via Reasoning Decomposition

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-07-01T09:35:39.626458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-01T06:29:51.635039Z digest=sha256:6fa455e0b12a1e4c68c5948206d226e1110bbcb3392a35fbcc46fbef4c0a494e

Observation eb47cb2f-387e-4298-b916-725573a4bb5c · inbound

Personalizing MLLMs via Reinforced Multimodal Reference Game cites this paper.

Personalizing MLLMs via Reinforced Multimodal Reference Game Self-Rewarding Vision-Language Model via Reasoning Decomposition

Reference 20

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T12:44:40.065540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T10:08:38.106708Z digest=sha256:9c871d4db89f05e9ef6b777833e1865635d72adbdcf945fca756d1bf216bfd6e

Observation 3c179be6-3bc0-407a-a340-25139bc23a3f · inbound

Token-Sparse Medical Multimodal Reasoning via Dual-Stream Reinforcement Learning cites this paper.

Token-Sparse Medical Multimodal Reasoning via Dual-Stream Reinforcement Learning Self-Rewarding Vision-Language Model via Reasoning Decomposition

Reference 51

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T10:15:45.188616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-07-01T05:36:43.609602Z digest=sha256:37c678934d909b11514ebf54f5877a24c32885a3e99b2495071ad21e46b3c0aa

Observation a7aa0d15-ecfb-42d8-b703-5f3bd601bbc6 · inbound

LASER: A Corrective Lens for LVLMs via Visual Attention Preservation and Sink Suppression cites this paper.

LASER: A Corrective Lens for LVLMs via Visual Attention Preservation and Sink Suppression Self-Rewarding Vision-Language Model via Reasoning Decomposition

Reference 32

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T16:58:42.352796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-03T16:53:43.427625Z digest=sha256:f5d5380628cb3d1870c1c5f800be8a55b165aee56612e07fdcbd4e994246268d

Observation 3158a5c5-bf95-4637-8211-a8dc6661c89f · inbound

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading cites this paper.

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Self-Rewarding Vision-Language Model via Reasoning Decomposition

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-13T05:28:45.311405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T05:28:45.311405Z digest=sha256:c36baa570e445916a8c692935639537e9f243ffbe820f01f1e179c57ed03d86b

Observation 3a164a77-f953-4015-a040-5be20571ce00 · inbound

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading cites this paper.

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Self-Rewarding Vision-Language Model via Reasoning Decomposition

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-14T15:24:58.589243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:24:58.589243Z digest=sha256:714f4c3127e30b82d10b1b80d38d8e22a14054ed191146978308c5c77a985604

Observation c3e0ded3-08e6-48e1-aa33-eefd0bad2c28 · inbound

Med-OPD: Improving Medical Vision-Language Models via Evidence-Aware On-Policy Distillation cites this paper.

Med-OPD: Improving Medical Vision-Language Models via Evidence-Aware On-Policy Distillation Self-Rewarding Vision-Language Model via Reasoning Decomposition

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T06:36:35.384825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:36:35.384825Z digest=sha256:7accb019aac562a3a2ba4af028e7043d08a58f9f13e1cdec46f6bd532d7ac593

Observation f60e0dd2-6991-47d1-977c-6a66da0f9a18 · inbound

RP-OPSD: Resolution-Privileged On-Policy Self-Distillation for Multimodal Large Language Models cites this paper.

RP-OPSD: Resolution-Privileged On-Policy Self-Distillation for Multimodal Large Language Models Self-Rewarding Vision-Language Model via Reasoning Decomposition

Reference 51

Resolution
unresolved
no resolver link, observed 2026-07-31T14:46:28.736812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T14:46:28.736812Z digest=sha256:827923e6bc53a686bca6c3941bd7b2fca098e1892c4e2f7f1ee65f47d1e94c8a