Pith. sign in

Paper Citation Record · LEDGER

OPD-V: Visual On-Policy Self-Distillation with Modality Balance

As of 18 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 1 inbound Pith citation observation for arXiv:2608.05131.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.05131 v2

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T16:47:41.888680Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:30:56.200896Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-16T00:30:57.011169Z

Reference resolution

52 of 52 outbound references displayed

  • verified exact1
  • verified fuzzy2
  • unresolved49
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ed2759aa-4eb8-4ab2-b4d3-8af106b060cb · outbound

This paper cites Beyond nl2code: A structured survey of multimodal code intelligence,.

OPD-V: Visual On-Policy Self-Distillation with Modality Balance Beyond nl2code: A structured survey of multimodal code intelligence,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T16:47:41.713547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:47:41.713547Z digest=sha256:26c3de2087eeae3f216aff94264327cc624b48e332f9e5bac46278189f8ac208

Observation a957e4b6-8496-453d-bd30-96c355e8b680 · outbound

This paper cites Self-evolving multi-agent systems via textual backpropagation.

OPD-V: Visual On-Policy Self-Distillation with Modality Balance Self-evolving multi-agent systems via textual backpropagation

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T16:47:41.722245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:47:41.722245Z digest=sha256:71d30d475a0542ef01b8f738b5379a1b1b27f8da5cdd845435b4c06c0db1e965

Observation 2b7d2f76-763a-4cae-aa54-c90bf61db022 · outbound

This paper cites SPOT! Revisiting Video-Language Models for Event Understanding.

OPD-V: Visual On-Policy Self-Distillation with Modality Balance SPOT! Revisiting Video-Language Models for Event Understanding

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T16:47:41.729486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:47:41.729486Z digest=sha256:dbed395a60a861e47eb0b02852623b4f93cb1be4f4d0fbf3aa9725632153b726

Observation 254e64fe-9d1e-48f9-bc76-a3401a47e5ae · outbound

This paper cites Loong: 10 Synthesize long chain-of-thoughts at scale through verifiers, 2025.

OPD-V: Visual On-Policy Self-Distillation with Modality Balance Loong: 10 Synthesize long chain-of-thoughts at scale through verifiers, 2025

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T16:47:42.561478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-08T16:47:41.733162Z digest=sha256:05f5c152da1be2dd45ae7a2c34c7f632961cf4766ecff7a0d4aac935fe7703b2

Observation a06cf162-82a8-49af-88e0-2ab2085af74d · outbound

This paper cites On-policy distillation of language models: Learning from self-generated mistakes.

OPD-V: Visual On-Policy Self-Distillation with Modality Balance On-policy distillation of language models: Learning from self-generated mistakes

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T16:47:41.737267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:47:41.737267Z digest=sha256:0dd2a9b90d5faec6736ff3c63ff51d29ab88c8de00d16b36294d779e8e3feeec

Observation e501289c-cf77-407c-adf0-54fc0584ba2e · outbound

This paper cites On-policy distillation.Thinking Machines Lab: Connectionism,.

OPD-V: Visual On-Policy Self-Distillation with Modality Balance On-policy distillation.Thinking Machines Lab: Connectionism,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T16:47:41.740648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:47:41.740648Z digest=sha256:d75362ff6f63d5f1c627686c3df45585d04098024f3586346abda50b65610a75

Observation 3feb1583-11b1-4f7c-bc09-36a6128c8d97 · outbound

This paper cites Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models.

OPD-V: Visual On-Policy Self-Distillation with Modality Balance Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T16:47:41.747694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:47:41.747694Z digest=sha256:ba9fdd40f8731cf17bce2223fb427374271d919ad668a32e47dd4f0e22ffb039

Observation 055d4105-15f7-4174-b43c-ced02e7f4f16 · outbound

This paper cites DOPD: Dual On-policy Distillation.

OPD-V: Visual On-Policy Self-Distillation with Modality Balance DOPD: Dual On-policy Distillation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T16:47:41.751943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:47:41.751943Z digest=sha256:36e7c4e43321c4ccae801f97772837f36368ae548bfb32a9267bf5d42b3c1198

Observation ca766d31-089c-4d82-a559-29b1bcc5222c · outbound

This paper cites ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning.

OPD-V: Visual On-Policy Self-Distillation with Modality Balance ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-08T16:47:42.408371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-08T16:47:41.755414Z digest=sha256:2d0688f9c3d32ab2105b788b438bc69f6e1676b4738fdd9286dccaf27791970e

Observation b8f55094-0812-40f8-84d9-ebf2da359fcc · outbound

This paper cites Vision-OPD: Learning to See Fine Details for Multimodal LLMs via On-Policy Self-Distillation.

OPD-V: Visual On-Policy Self-Distillation with Modality Balance Vision-OPD: Learning to See Fine Details for Multimodal LLMs via On-Policy Self-Distillation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T16:47:41.759043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:47:41.759043Z digest=sha256:0d8798ef4067e8b935e645863f80ab280789cae71333497afe15fba7553a98ce

Observation aad56867-fcd7-48c0-a791-7ac824617cab · outbound

This paper cites Visual-OPSD: Cross-Modal On-Policy Self-Distillation for Efficient Unified Multimodal Reasoning.

OPD-V: Visual On-Policy Self-Distillation with Modality Balance Visual-OPSD: Cross-Modal On-Policy Self-Distillation for Efficient Unified Multimodal Reasoning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T16:47:41.762461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:47:41.762461Z digest=sha256:1c19cc2f40bc893e49820948109ac4736c7887f36088496849bf96ed933a3385

Observation 9be022a1-7d4b-4b42-b393-b600bfac4269 · outbound

This paper cites Visual-Advantage On-Policy Distillation for Vision-Language Models.

OPD-V: Visual On-Policy Self-Distillation with Modality Balance Visual-Advantage On-Policy Distillation for Vision-Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T16:47:41.766048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:47:41.766048Z digest=sha256:085be561892bbe205337e3c1b054a49fdd13fccdc55eaec49cce9ba52a6acf5f

Observation 5b746a6a-f9b9-496a-b6b6-55ac0388ecf3 · outbound

This paper cites Visual Contrastive Self-Distillation.

OPD-V: Visual On-Policy Self-Distillation with Modality Balance Visual Contrastive Self-Distillation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T16:47:41.769576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:47:41.769576Z digest=sha256:eda0b8def25a6f520a959d35ba359be624790dc4876f81d8c014913431af1287

Observation 060ef9cb-da57-4525-aacc-0a5d3c2dfd47 · outbound

This paper cites LLaVA steering: Visual instruction tuning with 500x fewer parameters through modality linear representation- steering.

OPD-V: Visual On-Policy Self-Distillation with Modality Balance LLaVA steering: Visual instruction tuning with 500x fewer parameters through modality linear representation- steering

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T16:47:41.772811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:47:41.772811Z digest=sha256:12f6a36b0bf087a76c12e61cab15956745ffe3587eeffb6535b8df88b9468945

Observation 93dde378-685d-40f5-8ca2-430199bcc29a · outbound

This paper cites an unresolved cited work.

OPD-V: Visual On-Policy Self-Distillation with Modality Balance Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-08T16:47:42.512743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-08T16:47:41.776381Z digest=sha256:2da57743139b15f065f74cc8294df64d10bd6c343d5a099cc9fc8793e60d5b63

Observation cda22931-260a-4b4a-8793-d93c19c0840f · outbound

This paper cites Evaluating and steering modality preferences in multimodal large language model.arXiv preprint arXiv:2505.20977, 2025.

OPD-V: Visual On-Policy Self-Distillation with Modality Balance Evaluating and steering modality preferences in multimodal large language model.arXiv preprint arXiv:2505.20977, 2025

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T16:47:41.779515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:47:41.779515Z digest=sha256:2420c92fa42eb4559e0d6f8a6492df876fd568490539312d2ef391b8536fe2b8

Observation 9a6c8aec-8f49-4b46-bc1b-d65693c338d4 · outbound

This paper cites Instruction Anchor: Dissecting the Mechanistic Dynamics of Modality Arbitration.

OPD-V: Visual On-Policy Self-Distillation with Modality Balance Instruction Anchor: Dissecting the Mechanistic Dynamics of Modality Arbitration

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T16:47:41.782679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:47:41.782679Z digest=sha256:802c40de6afe2fb75340d0b5566a562b87590fd39fc0ea4cc97667ef282befe4

Observation 27326db0-6506-484b-8172-e2d398e169b6 · outbound

This paper cites Qwen3.5: Towards native multimodal agents, February 2026.

OPD-V: Visual On-Policy Self-Distillation with Modality Balance Qwen3.5: Towards native multimodal agents, February 2026

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T16:47:42.503922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-08T16:47:41.785879Z digest=sha256:1a9b5e33a9b1258d8082462748c5dfb25a023ba323955abacd90a174292a9211

Observation 721b0ae7-c96f-480c-aa05-6888ba63d76f · outbound

This paper cites Qwen3-VL Technical Report.

OPD-V: Visual On-Policy Self-Distillation with Modality Balance Qwen3-VL Technical Report

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T16:47:41.789313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:47:41.789313Z digest=sha256:f48d33bcaf2216813fce45a60eed7658cee6edd474eb393230f326427066be7c

Observation 4a5d6c25-b766-49b4-af7a-d8a675397542 · outbound

This paper cites V*: Guided visual search as a core mechanism in multimodal llms.

OPD-V: Visual On-Policy Self-Distillation with Modality Balance V*: Guided visual search as a core mechanism in multimodal llms

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T16:47:41.792448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:47:41.792448Z digest=sha256:2fc5e409d73b47422d74b56809b438bc637c7313e4e0fe36cce01ae562f7ffe7

Observation 628510d3-0a09-402a-99ce-e7b255d5067d · outbound

This paper cites Zooming without zooming: Region-to-image distillation for fine-grained multimodal perception.arXiv preprint arXiv:2602.11858, 2026.

OPD-V: Visual On-Policy Self-Distillation with Modality Balance Zooming without zooming: Region-to-image distillation for fine-grained multimodal perception.arXiv preprint arXiv:2602.11858, 2026

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T16:47:41.795303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:47:41.795303Z digest=sha256:5bdbfbc733f9e372b72ba81240ae9cdff691c714ef243910849e86e1d1e1a5f6

Observation bcf628de-e8a5-4926-9151-7830395329d9 · outbound

This paper cites Divide, conquer and combine: A training-free framework for high-resolution image perception in multimodal large language models.

OPD-V: Visual On-Policy Self-Distillation with Modality Balance Divide, conquer and combine: A training-free framework for high-resolution image perception in multimodal large language models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T16:47:41.798498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:47:41.798498Z digest=sha256:6a21a5c17974092fe146f25cc1d8f8df6ab1afad9421451b90c33fbc9f69b188

Observation 40158c7f-09e5-4552-9fba-b28ee91b0b2b · outbound

This paper cites MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?.

OPD-V: Visual On-Policy Self-Distillation with Modality Balance MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T16:47:41.802360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:47:41.802360Z digest=sha256:d6d3ad8cb9fe4b75ec73046d8ccb3356a015cc9c903965b4a29bd55c09edd006

Observation ab24a133-da4e-45bc-b676-7c79dcc5b530 · outbound

This paper cites Gemini 3.1 pro.https://deepmind.google/models/model-cards/gemini-3-1-pro/, 2026.

OPD-V: Visual On-Policy Self-Distillation with Modality Balance Gemini 3.1 pro.https://deepmind.google/models/model-cards/gemini-3-1-pro/, 2026

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T16:47:41.806120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:47:41.806120Z digest=sha256:59d7e514b320c1f92deb9bf5f4fe24f2748ac2318b92dc3b194730b8784399c4

Observation 2773812a-a9c9-4fa1-bb77-da9ed4ad9f51 · outbound

This paper cites Gemini 3.5 flash.

OPD-V: Visual On-Policy Self-Distillation with Modality Balance Gemini 3.5 flash

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T16:47:41.809291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:47:41.809291Z digest=sha256:edd67e983354e030b8fe5d745846d32ed5af1f5a2041d1266ae100f6806e3c70

Observation ca9e340c-f0d7-439e-8e83-26070ff7c657 · outbound

This paper cites Introducing gpt-5.4.https://openai.com/index/introducing-gpt-5-4/, 2026.

OPD-V: Visual On-Policy Self-Distillation with Modality Balance Introducing gpt-5.4.https://openai.com/index/introducing-gpt-5-4/, 2026

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T16:47:41.812775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:47:41.812775Z digest=sha256:efecd7b0bd83625827ce8844bbb4d02592e1433133e65c29be6857e43a1cc683

Observation 27fc40de-414f-47f5-bf54-476a9b14687a · outbound

This paper cites Introducing gpt-5.2.https://openai.com/index/introducing-gpt-5-2/, 2025.

OPD-V: Visual On-Policy Self-Distillation with Modality Balance Introducing gpt-5.2.https://openai.com/index/introducing-gpt-5-2/, 2025

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T16:47:41.816002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:47:41.816002Z digest=sha256:19e6fd014cf0df40a8ab304c5f32e816f7bd0c4257b545aacc932b22f6cdc82a

Observation a5e74056-a2b3-43a1-a194-f107f6c8bbdc · outbound

This paper cites Thyme: Think Beyond Images.

OPD-V: Visual On-Policy Self-Distillation with Modality Balance Thyme: Think Beyond Images

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T16:47:41.819444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:47:41.819444Z digest=sha256:fdf1fcafaf44c70abb9cba1b978a58aff7c6dcc651633cfa38d0113e99a04f2d

Observation 954c6a81-cf98-4f6d-9f58-34af42d9f9e4 · outbound

This paper cites DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning.

OPD-V: Visual On-Policy Self-Distillation with Modality Balance DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T16:47:41.823170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:47:41.823170Z digest=sha256:70e9af80db513835b2f8e3e6263d02baf24c4464b4a5089da7619db76f03f39c

Observation a028e577-74e6-4dc8-bb15-c040df5b32f3 · outbound

This paper cites DeepEyesV2: Toward Agentic Multimodal Model.

OPD-V: Visual On-Policy Self-Distillation with Modality Balance DeepEyesV2: Toward Agentic Multimodal Model

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T16:47:41.826547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:47:41.826547Z digest=sha256:f36fe769c601ebd911b719f5b94733c4e6266c37458037f599b52c6ea24cd404

Observation ebcee3e8-f82b-45e2-ad3c-d79c23b1e969 · outbound

This paper cites MiMo-VL Technical Report.

OPD-V: Visual On-Policy Self-Distillation with Modality Balance MiMo-VL Technical Report

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T16:47:41.830087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:47:41.830087Z digest=sha256:e90f1adfa7f12929a5ee70a8a65b260ccbe4cd14b0962b5d759e3051e5589bfc

Observation 22d24745-bb26-4331-9009-9623b51ffed3 · outbound

This paper cites Sensenova-mars: Empowering multimodal agentic reasoning and search via reinforcement learning.arXiv preprint arXiv:2512.24330, 2025.

OPD-V: Visual On-Policy Self-Distillation with Modality Balance Sensenova-mars: Empowering multimodal agentic reasoning and search via reinforcement learning.arXiv preprint arXiv:2512.24330, 2025

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-08T16:47:41.833295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:47:41.833295Z digest=sha256:898ccec2dbf0c3adbde49c5de2cbe06e96fccf09512b122eff03c985e3ea88fe

Observation 4d9cd7de-c772-4acc-9483-641d3c458f37 · outbound

This paper cites MiniCPM-V 4.5: Cooking Efficient MLLMs via Architecture, Data, and Training Recipe.

OPD-V: Visual On-Policy Self-Distillation with Modality Balance MiniCPM-V 4.5: Cooking Efficient MLLMs via Architecture, Data, and Training Recipe

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T16:47:41.836012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:47:41.836012Z digest=sha256:b003031de3199f39869d1c427a9338fdf721ca2984b0df43658c1c3b94c8aa1b

Observation b12c61fa-0861-4964-b96d-ef8818118568 · outbound

This paper cites GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning.

OPD-V: Visual On-Policy Self-Distillation with Modality Balance GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T16:47:41.839109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:47:41.839109Z digest=sha256:37dd41ac160f6333fc6a788ae4c434619572fb484e17310686e926fd6a2d1625

Observation 938cea84-7f31-4bf3-8013-9df19ce2a5c9 · outbound

This paper cites Kimi k2.6: From code to creation, from one to many.https://www.kimi.com/ai-models/ kimi-k2-6/, 2026.

OPD-V: Visual On-Policy Self-Distillation with Modality Balance Kimi k2.6: From code to creation, from one to many.https://www.kimi.com/ai-models/ kimi-k2-6/, 2026

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-08T16:47:41.842795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:47:41.842795Z digest=sha256:1f702d046a2d159a67cb8d92f6a47a25260e765c6a035d2535bf572f92858e6b

Observation 3d79de2b-d3b6-4fa2-8119-71a5949e1554 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

OPD-V: Visual On-Policy Self-Distillation with Modality Balance DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-08T16:47:41.845963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:47:41.845963Z digest=sha256:512ba148cc2f4f358519f89d466b04aaf3810bd5adbd2efefe085c6f68132267

Observation 8e24a82e-fd46-47e4-a818-60ce8542ff9d · outbound

This paper cites When More is Less: Understanding Chain-of-Thought Length in LLMs.

OPD-V: Visual On-Policy Self-Distillation with Modality Balance When More is Less: Understanding Chain-of-Thought Length in LLMs

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-08T16:47:41.849279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:47:41.849279Z digest=sha256:7dc37e168d19090d2f2b23fbde20fbcb47d4bff7e4686eed615fc628d7bd423b

Observation 18919af5-3d33-4cf9-ba2d-292157395f06 · outbound

This paper cites Don’t overthink it.

OPD-V: Visual On-Policy Self-Distillation with Modality Balance Don’t overthink it

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T16:47:41.852570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:47:41.852570Z digest=sha256:cf801a5e055cf3bb7c5d593b61840f0d499e340d1d431a611ef24a9f9ef0a7aa

Observation cfeaea17-b462-46fd-8be8-6090e63fc58a · outbound

This paper cites Distilling the Knowledge in a Neural Network.

OPD-V: Visual On-Policy Self-Distillation with Modality Balance Distilling the Knowledge in a Neural Network

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-08T16:47:41.855996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:47:41.855996Z digest=sha256:9af932a9a76592424e5cfe5fdb5f412ddc6b040ec726c64132a7ec87fa34a54b

Observation d2995293-38ed-4ca9-b24b-f8d0e7c717b7 · outbound

This paper cites Sequence-level knowledge distillation.

OPD-V: Visual On-Policy Self-Distillation with Modality Balance Sequence-level knowledge distillation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-08T16:47:41.859547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:47:41.859547Z digest=sha256:a88620e8c99c293a569704adac7693baa9f3cef9c86862cf2fb87452b72e1a76

Observation d0421358-6d0a-41bc-a84d-e9e81061e875 · outbound

This paper cites CoT-Kinetics: A Theoretical Modeling Assessing LRM Reasoning Process.

OPD-V: Visual On-Policy Self-Distillation with Modality Balance CoT-Kinetics: A Theoretical Modeling Assessing LRM Reasoning Process

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-08T16:47:41.862715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:47:41.862715Z digest=sha256:65c72d3170aa3c0b52f71db387407e93216f3b1c7b9067187be360103112875a

Observation f9e318d0-216f-48d2-9da8-063db2e87e94 · outbound

This paper cites EchoRL: Reinforcement learning via rollout echoing.

OPD-V: Visual On-Policy Self-Distillation with Modality Balance EchoRL: Reinforcement learning via rollout echoing

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-08T16:47:41.865808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:47:41.865808Z digest=sha256:05b8a68a02bf97ec9b69525a4c05d47fb6c3fe53aa8e4b24fbfef589082e5efd

Observation d0961669-044a-4ccd-be6c-6128c5833ef8 · outbound

This paper cites Seeing Before Reasoning: Decoupling Perception and Reasoning for Shortcut-Resilient Multimodal On-Policy Self-Distillation.

OPD-V: Visual On-Policy Self-Distillation with Modality Balance Seeing Before Reasoning: Decoupling Perception and Reasoning for Shortcut-Resilient Multimodal On-Policy Self-Distillation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-08T16:47:41.868728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:47:41.868728Z digest=sha256:28b8416e53c9013fe7ff1228369c5b067ecfe42f9861c164bae87b8a9ad0120e

Observation b9ef025d-63b8-40d7-814f-e6f10cb4c6f6 · outbound

This paper cites PRISM: Self-Pruning Intrinsic Selection Method for Training-Free Multimodal Data Selection.

OPD-V: Visual On-Policy Self-Distillation with Modality Balance PRISM: Self-Pruning Intrinsic Selection Method for Training-Free Multimodal Data Selection

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-08T16:47:41.871888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:47:41.871888Z digest=sha256:3ec86f19f756cd9d4a8595bd41afb2e1be84742761c36d6ca060389fe4b8b3cc

Observation 91c9a79a-ae49-4731-a22b-707d9a1f3a32 · outbound

This paper cites Can visual input be compressed? a visual token compression benchmark for large multimodal models, 2025.

OPD-V: Visual On-Policy Self-Distillation with Modality Balance Can visual input be compressed? a visual token compression benchmark for large multimodal models, 2025

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-08T16:47:41.879240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:47:41.879240Z digest=sha256:35b337269b0711b3d83a9c29e6e5d6575e0694a4711de0935470cf76ed527c11

Observation 09b0b6dc-c9de-40ff-8713-0cfd24cae192 · outbound

This paper cites Ascd: Attention-steerable contrastive decoding for reducing hallucination in mllm.Proceedings of the AAAI Conference on Artificial Intelligence, 40(12): 10306–10314, Mar.

OPD-V: Visual On-Policy Self-Distillation with Modality Balance Ascd: Attention-steerable contrastive decoding for reducing hallucination in mllm.Proceedings of the AAAI Conference on Artificial Intelligence, 40(12): 10306–10314, Mar

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-08T16:47:41.881958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:47:41.881958Z digest=sha256:3455b48e8dd9efe230db635b3478c3f223a55ca5f8a0871c8af4256b7c0ed41c

Observation 3960ff7c-0418-4ebe-a215-3d0b33d472ee · outbound

This paper cites MINED: Probing and Updating with Multimodal Time-Sensitive Knowledge for Large Multimodal Models.

OPD-V: Visual On-Policy Self-Distillation with Modality Balance MINED: Probing and Updating with Multimodal Time-Sensitive Knowledge for Large Multimodal Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-08T16:47:41.885039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:47:41.885039Z digest=sha256:9c8824710b50d86f4a3450443069effa15dd63f0b974c7d2150ba2a6f9f099c2

Observation 2b07d251-cb74-4932-b2aa-62eeeac04ac8 · outbound

This paper cites an unresolved cited work.

OPD-V: Visual On-Policy Self-Distillation with Modality Balance Unresolved cited work

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-08T16:47:41.875763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:47:41.875763Z digest=sha256:8c1be458b1c431a1f031ca71821742118eb46dcfd3c2085cb1ae7acc971ee020

Observation 2e658217-1d48-4b4b-b506-bca35bec248c · outbound

This paper cites KORE: Enhancing Knowledge Injection for Large Multimodal Models via Knowledge-Oriented Controls.

OPD-V: Visual On-Policy Self-Distillation with Modality Balance KORE: Enhancing Knowledge Injection for Large Multimodal Models via Knowledge-Oriented Controls

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-08T16:47:41.888680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:47:41.888680Z digest=sha256:31d659894d693e97b33243478b7d9dff8b8a38b673be8dfcd895036f3ca3da4c

Observation 1e974154-47f1-47f6-b8ea-890b31dc9ab4 · outbound

This paper cites an unresolved cited work.

OPD-V: Visual On-Policy Self-Distillation with Modality Balance Unresolved cited work

Reference 483

Resolution
unresolved
no resolver link, observed 2026-08-08T16:47:41.725852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:47:41.725852Z digest=sha256:965bda1012d75c5f6af6db12a8506f689837b87258c3a404fa0088d6ce275da2

Observation 8d93984f-861c-4542-a334-2b3f02ea1762 · outbound

This paper cites https://thinkingmachines.ai/blog/on-policy-distillation.

OPD-V: Visual On-Policy Self-Distillation with Modality Balance https://thinkingmachines.ai/blog/on-policy-distillation

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-08T16:47:41.744100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:47:41.744100Z digest=sha256:11481fcb5846fa2cc7cba70b0b1bef807b0a31da865719b648ef29c774c93f5e

Observation 9a0e09c0-0eba-4c85-84b6-bb1f5ed476ed · outbound

This paper cites Beyond NL2Code: A Structured Survey of Multimodal Code Intelligence.

OPD-V: Visual On-Policy Self-Distillation with Modality Balance Beyond NL2Code: A Structured Survey of Multimodal Code Intelligence

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-08T16:47:41.718003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:47:41.718003Z digest=sha256:47af2130fd89e2e09d3acbb77454ce7dc290d03715be97cb16c976425cfb7f6f

Pith citing papers

Observation 2befa88c-35d5-4c60-9168-be9980f6f945 · inbound

Asymptotic Risk Calibration for Selective Question Answering cites this paper.

Asymptotic Risk Calibration for Selective Question Answering OPD-V: Visual On-Policy Self-Distillation with Modality Balance

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-16T00:30:57.037536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:30:56.200896Z digest=sha256:72cd1268d7db3e2b18c5a24cab9aa3f855e91e17b3ad892b4f58bcfc70bd8983