Pith. sign in

Paper Citation Record · LEDGER

Reconstructive Visual Instruction Tuning

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 14 inbound Pith citation observations for arXiv:2410.09575.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.09575 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T20:32:46.962180Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T14:28:32.245115Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 8805349b-7f10-43ce-95fe-b6f8bde3ef67 · inbound

VGR: Visual Grounded Reasoning cites this paper.

VGR: Visual Grounded Reasoning Reconstructive Visual Instruction Tuning

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-19T09:12:14.436682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T09:11:00.295700Z digest=sha256:33360347028d9a9020d60510df8bba41d612365708fd5b17f5c7fcc3affb4552

Observation 961c6b19-704e-4cf5-8b3d-43130886e1a4 · inbound

ReconVLA: Reconstructive Vision-Language-Action Model as Effective Robot Perceiver cites this paper.

ReconVLA: Reconstructive Vision-Language-Action Model as Effective Robot Perceiver Reconstructive Visual Instruction Tuning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-05T20:32:46.962180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:32:46.962180Z digest=sha256:9cf4a6becdf0ee24f279053ef8393842e41b3fabfc6ad7a03c7f7e805cf0a7cb

Observation 88e3bb72-4031-4ae4-bb77-97095bef5fbe · inbound

Reconstruction Alignment Improves Unified Multimodal Models cites this paper.

Reconstruction Alignment Improves Unified Multimodal Models Reconstructive Visual Instruction Tuning

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-04T22:36:08.181930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T22:36:08.181930Z digest=sha256:c1de9ca82744b8e1b3b1d132772247bb615116c8503c349bcae6ba968a1d43ca

Observation 9c34c403-4ceb-4df0-a2ec-b05a656ffa54 · inbound

SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models cites this paper.

SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models Reconstructive Visual Instruction Tuning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T09:54:26.467952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:54:26.467952Z digest=sha256:ab583311735797303d8f495a548a7c5ed654e408c2379dae8ee20a2ba39311c7

Observation a6ff7103-f3bd-4daa-a8d2-eb8018ed8ca8 · inbound

DriveVLA-W0: World Models Amplify Data Scaling Law in Autonomous Driving cites this paper.

DriveVLA-W0: World Models Amplify Data Scaling Law in Autonomous Driving Reconstructive Visual Instruction Tuning

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-17T06:48:01.100732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T06:48:00.943591Z digest=sha256:3fc61b825cf5fc32fbe12d17050d6d0ce4d8dab1815e4054a7440d1d76e5cdec

Observation 4f0c3670-1400-4d15-8359-c45d93d6fba8 · inbound

LaRe: Latent Refocusing for Multimodal Reasoning cites this paper.

LaRe: Latent Refocusing for Multimodal Reasoning Reconstructive Visual Instruction Tuning

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-04T00:15:19.047232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:15:19.047232Z digest=sha256:3975e8a2cf35115d8bf3e01729334278c48b66e42c4c36e5467f8bdee55b54a5

Observation 727ea4d3-504e-4b96-b9c4-822755bbed99 · inbound

Generation Enhances Understanding in Unified Multimodal Models via Multi-Representation Generation cites this paper.

Generation Enhances Understanding in Unified Multimodal Models via Multi-Representation Generation Reconstructive Visual Instruction Tuning

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-03T07:03:13.944749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T07:03:13.944749Z digest=sha256:0e347ec60b250ee10408073d8bf5df642767ded9006b10b16c17a347377f13e6

Observation f122d2ab-b75b-451d-9b42-eafd08f36b78 · inbound

Latent Denoising Improves Visual Alignment in Large Multimodal Models cites this paper.

Latent Denoising Improves Visual Alignment in Large Multimodal Models Reconstructive Visual Instruction Tuning

Reference 86

Resolution
verified exact
arxiv_id, observed 2026-05-09T23:09:26.735275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-09T23:07:54.806529Z digest=sha256:0794571b5245f233bd91320a09892ff25ccbe5f4546cdb9d0c6b16caa78089b9

Observation bacff6a8-81a5-4502-80e1-c01ddd75dceb · inbound

See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding cites this paper.

See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding Reconstructive Visual Instruction Tuning

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:13:16.191326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T12:10:54.874012Z digest=sha256:74af870990ebc4a11dbf23748475c56f730a330bb7d275f29d794419a248746a

Observation 09bfac4b-1a20-4343-92f9-1c01c18a326f · inbound

Vision Inference Former: Sustaining Visual Consistency in Multimodal Large Language Models cites this paper.

Vision Inference Former: Sustaining Visual Consistency in Multimodal Large Language Models Reconstructive Visual Instruction Tuning

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:38:14.438803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T11:38:01.947806Z digest=sha256:dcf017b6306d5adad5f90a6bbc2b56ab2b9f79600a2857029fc1bb0cb653c171

Observation db4553d1-d21c-444d-9b2a-fa253baf28eb · inbound

Vision Inference Former: Sustaining Visual Consistency in Multimodal Large Language Models cites this paper.

Vision Inference Former: Sustaining Visual Consistency in Multimodal Large Language Models Reconstructive Visual Instruction Tuning

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:45:00.162571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T18:42:46.422756Z digest=sha256:e4841a5c27710f5e3981795ab92c351e177cb0c3a78299994fe0a714198cd974

Observation c1c5f902-0a47-4c48-a955-f265a42a3164 · inbound

Semantic Generative Tuning for Unified Multimodal Models cites this paper.

Semantic Generative Tuning for Unified Multimodal Models Reconstructive Visual Instruction Tuning

Reference 61

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T11:33:14.237180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T11:32:24.007847Z digest=sha256:276aab62692d30e85ba330c77c84f0218fc1b33ad6e47aa401ab71fbddd31f1e

Observation fb0388fb-fd3c-4570-b15b-345d8c0f86fc · inbound

Semantic Generative Tuning for Unified Multimodal Models cites this paper.

Semantic Generative Tuning for Unified Multimodal Models Reconstructive Visual Instruction Tuning

Reference 61

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T18:35:00.342775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:09c9e6536dc40e7ec2909edcd8e54aaa77e3b2f81b5235f997aa0e6602566cea

Observation 6b3391d9-473d-4a90-a54f-0e402270f3bb · inbound

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers cites this paper.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers Reconstructive Visual Instruction Tuning

Reference 135

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T14:28:32.246590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:f7570893f4ca9187174a3718c981163e23dbecb8cf30668efee6f6aee4bc4135