Pith. sign in

Paper Citation Record · LEDGER

Silkie: Preference Distillation for Large Visual Language Models

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 26 inbound Pith citation observations for arXiv:2312.10665.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2312.10665 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 26 of 26 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:55:12.792022Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

3
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 43b959f8-b5da-4dbb-89c9-e9adc90adbeb · inbound

PMC-VQA: Visual Instruction Tuning for Medical Visual Question Answering cites this paper.

PMC-VQA: Visual Instruction Tuning for Medical Visual Question Answering Silkie: Preference Distillation for Large Visual Language Models

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-15T23:08:21.802896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T23:08:21.673251Z digest=sha256:80df3b260f3892c6dd43bb631eb070effb463d938fe958f730b21e0a4606e9d8

Observation 33f6ea4b-dd79-4aba-b544-0a251b110e92 · inbound

A Survey on Multimodal Large Language Models cites this paper.

A Survey on Multimodal Large Language Models Silkie: Preference Distillation for Large Visual Language Models

Reference 117

Resolution
verified exact
arxiv_id, observed 2026-05-16T02:56:41.820282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T02:56:41.658658Z digest=sha256:8a5f27dd25ea10e398efe6fa8e51a79183f0145fa221eb56fe98cc40af3ebf3b

Observation a9f82327-f41d-408d-a9eb-0d0f0314a681 · inbound

Aligning Modalities in Vision Large Language Models via Preference Fine-tuning cites this paper.

Aligning Modalities in Vision Large Language Models via Preference Fine-tuning Silkie: Preference Distillation for Large Visual Language Models

Reference 161

Resolution
verified exact
arxiv_id, observed 2026-05-17T10:58:53.398463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-17T10:58:53.215887Z digest=sha256:de19e5a152a8bb436dd6925d391730fd57336e40268c4fa5148c3748e1c86f3f

Observation 2c650b40-4f4a-4d41-bb7a-cc3b06e5423c · inbound

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models cites this paper.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models Silkie: Preference Distillation for Large Visual Language Models

Reference 106

Resolution
verified exact
arxiv_id, observed 2026-05-23T22:20:21.522610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:10539be866a6eff59dd284a3ab7f0b4e0985469c832cfd5d3047106fdf340dd7

Observation c39c5c11-79f8-4f80-8e00-22f304f144db · inbound

A Survey on Knowledge Distillation of Large Language Models cites this paper.

A Survey on Knowledge Distillation of Large Language Models Silkie: Preference Distillation for Large Visual Language Models

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T23:31:11.672692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-17T23:31:11.213552Z digest=sha256:ab8dd3546d59cb1543f44bc9afd03ab6e64814da0989534134e3539ca2a37180

Observation 233cdd86-cc0c-4292-a444-a54f32b7ff39 · inbound

OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments cites this paper.

OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments Silkie: Preference Distillation for Large Visual Language Models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:19:32.561482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:19:32.406859Z digest=sha256:e0954f38447f8f5272f66f40a4f8bd0de408d4c17481404c08651564ca518ac5

Observation acdea6cc-2661-40b4-829c-5593e707556c · inbound

Hallucination of Multimodal Large Language Models: A Survey cites this paper.

Hallucination of Multimodal Large Language Models: A Survey Silkie: Preference Distillation for Large Visual Language Models

Reference 105

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:33:33.939949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T12:33:32.631346Z digest=sha256:e8b460fa08e6efc53005e6e3b1422509a48d745a37bb2ef0453fc92910371423

Observation 13a8d59a-a86e-4d9e-90e9-2a40123a5b4b · inbound

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output cites this paper.

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output Silkie: Preference Distillation for Large Visual Language Models

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-17T10:46:28.649205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T10:46:28.447347Z digest=sha256:d5a18e758f04f639cf778448239d2b43377203e2a17c33c544a3a988ae49e7d2

Observation 60a72d31-2619-46f2-9ead-8271ae6e6c22 · inbound

Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization cites this paper.

Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization Silkie: Preference Distillation for Large Visual Language Models

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:16:17.256298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T09:16:17.150383Z digest=sha256:018d4a74855d5d7ae2fa09f23021bb6c60b2177c2bf509a91c839f23d9364737

Observation 0eb33f72-d27a-4a76-930a-3606de24e0fc · inbound

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking cites this paper.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Silkie: Preference Distillation for Large Visual Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:12.792022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:12.792022Z digest=sha256:46951ffcc849115989511667b9a2640b3b31688d588bdfa99ef0cd0f210126cc

Observation 2851502b-aa0c-47f8-bb4b-5bbecad3798b · inbound

DPO Learning with LLMs-Judge Signal for Computer Use Agents cites this paper.

DPO Learning with LLMs-Judge Signal for Computer Use Agents Silkie: Preference Distillation for Large Visual Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:11:47.407418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:11:47.407418Z digest=sha256:7e6a497e5ea719ec87fcec5c1f125dac45d1c4a72ba7385a28cbf05d10dfedd1

Observation 96db0a19-af0d-4def-b013-5187b34ca691 · inbound

LeanPO: Lean Preference Optimization for Likelihood Alignment in Video-LLMs cites this paper.

LeanPO: Lean Preference Optimization for Likelihood Alignment in Video-LLMs Silkie: Preference Distillation for Large Visual Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:49.027723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:49.027723Z digest=sha256:e076185b3cbc1c66e4d435157baa8d7b0009fc1eb00bc281ef8ddcf201c604b7

Observation c32024ec-ed48-4892-8f1a-7e24be6b056e · inbound

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs cites this paper.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Silkie: Preference Distillation for Large Visual Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T23:27:19.350818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:27:19.350818Z digest=sha256:7b2549b196446e8e5de2a7cea9f4bfa0637b94717609b8eb07e2a4a81a53ae80

Observation 197ee628-b76b-4860-9504-86827db0aa01 · inbound

Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities cites this paper.

Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities Silkie: Preference Distillation for Large Visual Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:35.238949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:35.238949Z digest=sha256:031a8a1919ea51b1421d9d071f6c189e66dfab9ebcdbabb1a951f4e166b33300

Observation a5c9280e-2c3a-4799-be94-002e5bd50402 · inbound

Mitigating Object Hallucinations via Sentence-Level Early Intervention cites this paper.

Mitigating Object Hallucinations via Sentence-Level Early Intervention Silkie: Preference Distillation for Large Visual Language Models

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-25T08:35:32.553471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-25T08:31:24.173135Z digest=sha256:169d5a731e144ecc9b301854223b51c764108e863a18cd0b074697fdc41dc3ff

Observation da879e89-1fa4-4893-b6cd-b2135a2a2a45 · inbound

Empowering Multimodal LLMs with External Tools: A Comprehensive Survey cites this paper.

Empowering Multimodal LLMs with External Tools: A Comprehensive Survey Silkie: Preference Distillation for Large Visual Language Models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-05T20:28:48.508649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:28:48.508649Z digest=sha256:4b818618990a8b5ee825c7e9bb8cee612105d6197b3e8c65cbc256ef0ad07c23

Observation 82399be2-2ebe-412b-b0e7-3220ae1b890c · inbound

Improving Large Vision and Language Models by Learning from a Panel of Peers cites this paper.

Improving Large Vision and Language Models by Learning from a Panel of Peers Silkie: Preference Distillation for Large Visual Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.380128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.380128Z digest=sha256:5ec0c820ba899bbabe3025c6266ee60e2eab5f55c89c39b04c7b760f5765a04a

Observation bcfcfc53-0e48-4a26-963a-ef15269b830b · inbound

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs cites this paper.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Silkie: Preference Distillation for Large Visual Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:54.615173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:54.615173Z digest=sha256:c6ea3b42c03127281a318fcca28b54472f555546cbee3537b2399e5892c51481

Observation 8fc7bf25-77a8-4a10-837c-680aa97badec · inbound

Topo-R1: Detecting Topological Anomalies via Vision-Language Models cites this paper.

Topo-R1: Detecting Topological Anomalies via Vision-Language Models Silkie: Preference Distillation for Large Visual Language Models

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T11:45:32.848017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T11:41:27.021776Z digest=sha256:7b411add49c6f59fd9c3d2921440183fa9d79c52114235c7b8ba7563b7e0c820

Observation e06f9429-3404-4e78-868f-c91be94d3765 · inbound

You Only Judge Once: Multi-response Reward Modeling in a Single Forward Pass cites this paper.

You Only Judge Once: Multi-response Reward Modeling in a Single Forward Pass Silkie: Preference Distillation for Large Visual Language Models

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T09:20:59.335015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T16:06:57.084550Z digest=sha256:deed3a28817d147547dcfdacc1d75a99296ffeee6d9e4808ff798fe6e1025d12

Observation 9a57f30b-4b80-4ea1-ab84-b05383e431fc · inbound

Visual Preference Optimization with Rubric Rewards cites this paper.

Visual Preference Optimization with Rubric Rewards Silkie: Preference Distillation for Large Visual Language Models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:56:00.855516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:45:52.980881Z digest=sha256:0b943a3ee6fcfe5693b777346ccf18b54371147d710c214c406c8045f4306165

Observation 529af750-587c-4766-b290-e9a46de366ea · inbound

SignDPO: Multi-level Direct Preference Optimisation for Skeleton-based Gloss-free Sign Language Translation cites this paper.

SignDPO: Multi-level Direct Preference Optimisation for Skeleton-based Gloss-free Sign Language Translation Silkie: Preference Distillation for Large Visual Language Models

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:01:06.492898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T04:17:32.036750Z digest=sha256:c105c5966860480986d3affa30bb2cb3a03ed4cb683a9c14a2e4ca162213de65

Observation bcfe1c6d-7f72-4077-a0f3-a9edb50883ab · inbound

Online Self-Calibration Against Hallucination in Vision-Language Models cites this paper.

Online Self-Calibration Against Hallucination in Vision-Language Models Silkie: Preference Distillation for Large Visual Language Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:16:10.669193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-09T20:20:27.931679Z digest=sha256:46214d72d760490bcd5fd03733591969a33b97e9293d9f618d456648c933ccc7

Observation d2a7ace1-8370-4b84-a5ae-b6f5fe6768e0 · inbound

Deep Pre-Alignment for VLMs cites this paper.

Deep Pre-Alignment for VLMs Silkie: Preference Distillation for Large Visual Language Models

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-19T16:27:39.101866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-19T16:26:41.094936Z digest=sha256:506d38c2959e0c5e1e2e4c674a260aee8b01a36bf3aa96b859032c8ca1f539d4

Observation 792dae05-a8e5-47e6-8fdd-9ba7f06406cf · inbound

Toward Native Multimodal Modeling: A Roadmap cites this paper.

Toward Native Multimodal Modeling: A Roadmap Silkie: Preference Distillation for Large Visual Language Models

Reference 171

Resolution
verified exact
arxiv_id, observed 2026-06-29T23:04:01.994211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T22:58:38.610609Z digest=sha256:1c529676458503353f2d84c9df615993722644aa16469e4cdcb5183581f39d80

Observation 34a67d23-2985-4dfa-943f-853f97b59a81 · inbound

MAPL: Multi-Objective Preference Learning for Robot Locomotion cites this paper.

MAPL: Multi-Objective Preference Learning for Robot Locomotion Silkie: Preference Distillation for Large Visual Language Models

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-04T19:40:06.232152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-25T21:11:00.753949Z digest=sha256:a4734377aa6ba278e7a8352a3107df8438403f9f1cb22885736f7f8f97c84363