Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-22T07:24:22.355378Z
Paper Citation Record · LEDGER
As of 30 July 2026, this Paper Citation Record lists 33 of 33 outbound references and 1 inbound Pith citation observation for arXiv:2605.21924.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-22T07:24:22.355378Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-07-30T06:33:22.917629+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-30T05:51:18.037199Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-06-30T05:54:18.602535Z
33 of 33 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 57559873-2a57-4dc0-8753-5044f14c5556 · outbound
Visual-Advantage On-Policy Distillation for Vision-Language Models On-policy distillation of language models: Learning from self-generated mistakes
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-30T06:33:22.917629+00:00.
Observation 013e0229-9e32-47fc-af7c-5315f2ddd62b · outbound
Visual-Advantage On-Policy Distillation for Vision-Language Models Minillm: Knowledge distillation of large language models
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-30T06:33:22.917629+00:00.
Observation aa5096a4-7941-4869-95dd-7cba8386883a · outbound
Visual-Advantage On-Policy Distillation for Vision-Language Models Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-30T06:33:22.917629+00:00.
Observation 7c5fbd27-8657-4e24-8198-00f7d6280272 · outbound
Visual-Advantage On-Policy Distillation for Vision-Language Models DistiLLM: Towards Streamlined Distillation for Large Language Models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-30T06:33:22.917629+00:00.
Observation 888c5264-6fa5-44ce-986d-090b55c4b6e6 · outbound
Visual-Advantage On-Policy Distillation for Vision-Language Models Qwen3-VL Technical Report
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-30T06:33:22.917629+00:00.
Observation aee4b324-9c80-4536-ba4f-e7242736c636 · outbound
Visual-Advantage On-Policy Distillation for Vision-Language Models Inter-gps: Interpretable geometry problem solving with formal language and symbolic reasoning
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-30T06:33:22.917629+00:00.
Observation e01efe42-27d9-4c62-8c42-b61f8cee2516 · outbound
Visual-Advantage On-Policy Distillation for Vision-Language Models VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-30T06:33:22.917629+00:00.
Observation 8005a4f9-3d78-45ee-9715-6d428831ed0e · outbound
Visual-Advantage On-Policy Distillation for Vision-Language Models Distilling the Knowledge in a Neural Network
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-30T06:33:22.917629+00:00.
Observation 8be8682c-5803-4ea4-97b1-6ef7587c8a4a · outbound
Visual-Advantage On-Policy Distillation for Vision-Language Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-30T06:33:22.917629+00:00.
Observation 730a824b-9a15-4357-8ae9-130108c9edc3 · outbound
Visual-Advantage On-Policy Distillation for Vision-Language Models Perception-Aware Policy Optimization for Multimodal Reasoning
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-30T06:33:22.917629+00:00.
Observation 1b19c7b6-e00e-4fa5-af0a-58120a41237f · outbound
Visual-Advantage On-Policy Distillation for Vision-Language Models We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-30T06:33:22.917629+00:00.
Observation f76cbc4b-0cf9-4835-be48-443c2b0427ec · outbound
Visual-Advantage On-Policy Distillation for Vision-Language Models MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-30T06:33:22.917629+00:00.
Observation 5db27f69-3c9f-4d38-8b28-8cea840df618 · outbound
Visual-Advantage On-Policy Distillation for Vision-Language Models Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? InEuropean Conference on Computer Vision, pages 169–186
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-30T06:33:22.917629+00:00.
Observation 00a5a185-29ef-4b77-bced-94cd9a4bfbf6 · outbound
Visual-Advantage On-Policy Distillation for Vision-Language Models Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-30T06:33:22.917629+00:00.
Observation caf7152e-1ab1-48b8-b867-447713b8f377 · outbound
Visual-Advantage On-Policy Distillation for Vision-Language Models A diagram is worth a dozen images
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-30T06:33:22.917629+00:00.
Observation aa3e59f3-790a-4a9d-96cf-36596968c0fc · outbound
Visual-Advantage On-Policy Distillation for Vision-Language Models Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-30T06:33:22.917629+00:00.
Observation e3750704-5ce0-4b95-b317-7ad0c7351e45 · outbound
Visual-Advantage On-Policy Distillation for Vision-Language Models Are we on the right way for evaluating large vision- language models?Advances in Neural Information Processing Systems, 37:27056–27087
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-30T06:33:22.917629+00:00.
Observation ff1efb2c-c145-42b1-839e-568be0d64260 · outbound
Visual-Advantage On-Policy Distillation for Vision-Language Models Ocrbench: on the hidden mystery of ocr in large multimodal models.Science China Information Sciences, 67(12):220102
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-30T06:33:22.917629+00:00.
Observation c3462313-02a8-417a-911e-ae6131e2ce0a · outbound
Visual-Advantage On-Policy Distillation for Vision-Language Models Multi-modal hallucination control by visual information grounding
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-30T06:33:22.917629+00:00.
Observation e9832a10-f06f-496a-966b-f151860241c1 · outbound
Visual-Advantage On-Policy Distillation for Vision-Language Models Mitigating object hallucinations in large vision-language models through visual contrastive decoding
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-30T06:33:22.917629+00:00.
Observation 34958356-7e2e-4cce-951d-a757cadab9ac · outbound
Visual-Advantage On-Policy Distillation for Vision-Language Models Visual Description Grounding Reduces Hallucinations and Boosts Reasoning in LVLMs
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-30T06:33:22.917629+00:00.
Observation f803e542-6f41-4534-91c7-a63396146e47 · outbound
Visual-Advantage On-Policy Distillation for Vision-Language Models V-dpo: Mitigating hallucination in large vision language models via vision-guided direct preference optimization
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-30T06:33:22.917629+00:00.
Observation b8713252-d2d9-4d04-97b3-73766edbde3a · outbound
Visual-Advantage On-Policy Distillation for Vision-Language Models Self-Rewarding Vision-Language Model via Reasoning Decomposition
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-30T06:33:22.917629+00:00.
Observation e246103d-1dfb-489a-a35f-d02c0f80c8f6 · outbound
Visual-Advantage On-Policy Distillation for Vision-Language Models Mitigating hallucinations in large vision-language models with instruction contrastive decoding
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-30T06:33:22.917629+00:00.
Observation 5194f307-16ab-49f9-940c-f4bfb160999b · outbound
Visual-Advantage On-Policy Distillation for Vision-Language Models Token preference optimization with self-calibrated visual-anchored rewards for hallucination mitigation
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-30T06:33:22.917629+00:00.
Observation 921e3ae4-68e3-47cc-a0e5-8cc3f47e821c · outbound
Visual-Advantage On-Policy Distillation for Vision-Language Models Spotlight on token perception for multimodal reinforcement learning
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-30T06:33:22.917629+00:00.
Observation 5d371482-c469-4a26-8dab-dbb0b4d7fe08 · outbound
Visual-Advantage On-Policy Distillation for Vision-Language Models Rethinking token-level policy optimization for multimodal chain-of-thought
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-30T06:33:22.917629+00:00.
Observation af1aefb4-1f19-4ef4-a049-9c6da5593db9 · outbound
Visual-Advantage On-Policy Distillation for Vision-Language Models Noisyrollout: Reinforcing visual reasoning with data augmentation
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-30T06:33:22.917629+00:00.
Observation 1a57432b-5303-4b59-a980-a052d79d3125 · outbound
Visual-Advantage On-Policy Distillation for Vision-Language Models Rethinking kullback-leibler divergence in knowledge distillation for large language models
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-30T06:33:22.917629+00:00.
Observation 9d9cf732-09dd-4463-9940-65546755eadc · outbound
Visual-Advantage On-Policy Distillation for Vision-Language Models Entropy-Aware On-Policy Distillation of Language Models
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-30T06:33:22.917629+00:00.
Observation 1b7bc637-9aed-4d75-b525-4857ce2f19a8 · outbound
Visual-Advantage On-Policy Distillation for Vision-Language Models Speculative Knowledge Distillation: Bridging the Teacher-Student Gap Through Interleaved Sampling
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-30T06:33:22.917629+00:00.
Observation 6fc31c0c-aadc-4534-9ba9-eb0689295ab3 · outbound
Visual-Advantage On-Policy Distillation for Vision-Language Models On-policy distillation
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-30T06:33:22.917629+00:00.
Observation ba133aa1-9274-48c1-862d-c427a72288bd · outbound
Visual-Advantage On-Policy Distillation for Vision-Language Models Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-30T06:33:22.917629+00:00.
Observation 6ec24019-afb0-4680-98cf-7d0137eb90e3 · inbound
DOPD: Dual On-policy Distillation Visual-Advantage On-Policy Distillation for Vision-Language Models
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-30T06:33:22.917629+00:00.