Pith. sign in

Paper Citation Record · LEDGER

Planting a SEED of Vision in Large Language Model

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 28 inbound Pith citation observations for arXiv:2307.08041.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2307.08041 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 28 of 28 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T12:37:38.362436Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

10
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d3fe3931-6c77-4fb0-ae5c-027a8e27dca6 · inbound

SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension cites this paper.

SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension Planting a SEED of Vision in Large Language Model

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T16:59:50.525873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-12T16:59:50.495335Z digest=sha256:fd58d4bee9f287fb62edb3579d8b5517be381727d0a829b8543b8eedd2ae4eb0

Observation 26d03b63-a1bc-4499-97de-22bfc5ebf43f · inbound

DragNUWA: Fine-grained Control in Video Generation by Integrating Text, Image, and Trajectory cites this paper.

DragNUWA: Fine-grained Control in Video Generation by Integrating Text, Image, and Trajectory Planting a SEED of Vision in Large Language Model

Reference 258

Resolution
verified exact
arxiv_id, observed 2026-05-20T13:03:58.025865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-20T13:03:57.828598Z digest=sha256:0b895cff8d0d95521c2b7d75e237d7967ff0f8e5c55524f6cf2fbb73e8e0e48a

Observation d8eff912-36a9-4665-8a80-3d3ee5cdaeb6 · inbound

Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models cites this paper.

Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models Planting a SEED of Vision in Large Language Model

Reference 48

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T07:44:47.448796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-17T07:44:47.355960Z digest=sha256:8c3c5a6b91d7d414c24b4b563875416f983512f3dca6602797375f4a92b72268

Observation 8bc9297e-4349-43c7-891b-fde40426852c · inbound

SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation cites this paper.

SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation Planting a SEED of Vision in Large Language Model

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T22:48:36.350446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T22:48:36.010306Z digest=sha256:1a5beaa8339a66fcf4f8dae685b8db06e1f2bbce8187984f088d5c9f94c56191

Observation b9fe7687-ffd6-4a6c-b006-a12f68c18305 · inbound

Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs cites this paper.

Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs Planting a SEED of Vision in Large Language Model

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T00:05:03.785210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-17T00:05:03.547664Z digest=sha256:cffe6ae2d312d0bfdb9b3549597d0a6856fa84a7f325c7c0a55e9fffc38a49ae

Observation e61c6e61-1175-439f-9f01-4d34cc6d17e3 · inbound

VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation cites this paper.

VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation Planting a SEED of Vision in Large Language Model

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T00:26:21.394182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-16T00:26:21.313005Z digest=sha256:3a071e4ec71ea7cede529d8411e9d5ab6656fde0563b1246f3a6b34e4be4eae9

Observation 66db182b-a9ef-4077-b2f4-0460af504fe6 · inbound

Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation cites this paper.

Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation Planting a SEED of Vision in Large Language Model

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:09:16.472580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T22:09:16.001309Z digest=sha256:a2f3ca26d99e00938f5828d7bb9e5175486d3843a2cb648901a259634ff78f48

Observation de316236-edad-4bf0-86f1-c444ee35cb11 · inbound

MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding cites this paper.

MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Planting a SEED of Vision in Large Language Model

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T12:37:38.362436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:37:38.362436Z digest=sha256:dc7b58381b3296bde8abaa091370817f6bb085ada271097b326065a88ac4a90e

Observation 9154674f-9add-4bb6-9b97-678903d4b16d · inbound

Liquid: Language Models are Scalable and Unified Multi-modal Generators cites this paper.

Liquid: Language Models are Scalable and Unified Multi-modal Generators Planting a SEED of Vision in Large Language Model

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T21:35:40.237779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:35:40.237779Z digest=sha256:99aaae24d50e7316ee2e5368858e6fe20f4a402963d429c41480c5b5c8e0f06e

Observation 61e1d6ab-c3e5-4ca6-b4f9-94f8caa32e1e · inbound

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation cites this paper.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Planting a SEED of Vision in Large Language Model

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:08.253576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:08.253576Z digest=sha256:02afe2eb761075b3f1974a1c67571ccd3149c9d96956c675780eb0eab24ff9de

Observation 7ba20b08-59d5-4ff9-ae81-01a4de02db27 · inbound

EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios cites this paper.

EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios Planting a SEED of Vision in Large Language Model

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:00.130011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:00.130011Z digest=sha256:d9603de6b10323feb7cc811b0ba7123189cb0a2ec33940b7aafeea826423f0db

Observation c83293ac-b572-44e4-a03e-2a2e4eec0f1a · inbound

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models cites this paper.

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models Planting a SEED of Vision in Large Language Model

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:43.785719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:31:43.785719Z digest=sha256:e9811c41095bc462597faa98f1eb0534a1c181bf8f753d69f55649590c1bbf79

Observation ef9d18dc-8bc5-4634-bde8-aa445cfd9309 · inbound

MetaMorph: Multimodal Understanding and Generation via Instruction Tuning cites this paper.

MetaMorph: Multimodal Understanding and Generation via Instruction Tuning Planting a SEED of Vision in Large Language Model

Reference 216

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T07:51:13.261881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-17T07:51:12.953777Z digest=sha256:cb94c0db22dece83fc900c4d19f2e35ad91dabd613207178f4574492aee55f0a

Observation 8662debd-c011-4ec7-bd3e-628f45358232 · inbound

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey cites this paper.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey Planting a SEED of Vision in Large Language Model

Reference 131

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.691759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.691759Z digest=sha256:65ed9e23bca89994685ce828a713ccff0e65be6087d60f65e80dccf4cae5cf86

Observation 1cf65f1b-7790-4aa7-8ee8-aa3128b26611 · inbound

Dual Diffusion for Unified Image Generation and Understanding cites this paper.

Dual Diffusion for Unified Image Generation and Understanding Planting a SEED of Vision in Large Language Model

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T23:02:07.652832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:02:07.652832Z digest=sha256:9e4635b72c4fa3c20dc9492380c6c87a1701083b2be6bec9a2312dfe4e859a4a

Observation cef15181-dc25-4773-93a6-0df94599faa0 · inbound

HealthGPT: A Medical Large Vision-Language Model for Unifying Comprehension and Generation via Heterogeneous Knowledge Adaptation cites this paper.

HealthGPT: A Medical Large Vision-Language Model for Unifying Comprehension and Generation via Heterogeneous Knowledge Adaptation Planting a SEED of Vision in Large Language Model

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T20:21:30.400783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T20:21:30.400783Z digest=sha256:7e109123d1140bf6e81d261f2e1643f7bdd7c09c9d5577088fd9e4f5c96af375

Observation 7f3c46e5-a305-4c8f-8f23-2cdc7a32c81f · inbound

Transfer between Modalities with MetaQueries cites this paper.

Transfer between Modalities with MetaQueries Planting a SEED of Vision in Large Language Model

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T22:49:23.162924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-14T22:49:23.074271Z digest=sha256:dab7da22c0900c7dc70f231cd5bbfe1346d130c566317e4a6a1421d87856b180

Observation 3d2fe74f-9855-4096-83cc-2da854771b37 · inbound

Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation cites this paper.

Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation Planting a SEED of Vision in Large Language Model

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T07:24:04.821007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-17T07:24:04.460276Z digest=sha256:35fabb82e880557b2225624f56f975115091e82f3353aea7fbc93ee96728811c

Observation 9121c834-8d28-4554-bafa-d53507f3dc9c · inbound

FUDOKI: Discrete Flow-based Unified Understanding and Generation via Kinetic-Optimal Velocities cites this paper.

FUDOKI: Discrete Flow-based Unified Understanding and Generation via Kinetic-Optimal Velocities Planting a SEED of Vision in Large Language Model

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:04:54.247781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:04:54.247781Z digest=sha256:19683e598042af27c029a6860b0d0cfa92d903158d3a88f42a3b1580056792f4

Observation a314e9b2-206b-4574-adc3-c2b4999361bd · inbound

From Standalone LLMs to Integrated Intelligence: A Survey of Compound Al Systems cites this paper.

From Standalone LLMs to Integrated Intelligence: A Survey of Compound Al Systems Planting a SEED of Vision in Large Language Model

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:52:16.301701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-19T11:49:36.574471Z digest=sha256:0c15b9c782030863fb4e26ee72bd3b60b5b6888fa0ec2d542e61a522bfd56f8c

Observation 5244ed3a-2046-4a49-a4ee-77feb14e17cc · inbound

Instella-T2I: Pushing the Limits of 1D Discrete Latent Space Image Generation cites this paper.

Instella-T2I: Pushing the Limits of 1D Discrete Latent Space Image Generation Planting a SEED of Vision in Large Language Model

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T22:43:02.997268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:43:02.997268Z digest=sha256:ac73fdfc70c32ac156417ac3b34818335f017a5528dda2baddc5a9a13c619b02

Observation ff4062fc-f51c-4b6a-8b42-aeb20bd31974 · inbound

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs cites this paper.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs Planting a SEED of Vision in Large Language Model

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:15.540771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:15.540771Z digest=sha256:466ebec3f91bdd59b1ef937c2436a49d99212f763a3925c623c253692155f774

Observation 81ff7b25-87d0-4a82-9adf-ca1f202ec84a · inbound

Colon-Bench: An Agentic Workflow for Scalable Dense Lesion Annotation in Full-Procedure Colonoscopy Videos cites this paper.

Colon-Bench: An Agentic Workflow for Scalable Dense Lesion Annotation in Full-Procedure Colonoscopy Videos Planting a SEED of Vision in Large Language Model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-13T18:06:18.381075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T18:06:18.381075Z digest=sha256:c30ee69887adc98c6a0c6bdb2c8d6f20fe5719468b6d9e1a2be65d954cd02e0e

Observation 4c277ecd-dbf2-4959-a401-7b8862c8793e · inbound

DeepSight: Long-Horizon World Modeling via Latent States Prediction for End-to-End Autonomous Driving cites this paper.

DeepSight: Long-Horizon World Modeling via Latent States Prediction for End-to-End Autonomous Driving Planting a SEED of Vision in Large Language Model

Reference 206

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:31:26.428162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-12T04:13:37.421188Z digest=sha256:fc44ef73905afd4d846f0c24faf228d169d2ebabdacd8ac5aedb1bd8517fd0ae

Observation adfc6a0b-763b-49c7-8317-ce7147a39521 · inbound

Learning to See What You Need: Gaze Attention for Multimodal Large Language Models cites this paper.

Learning to See What You Need: Gaze Attention for Multimodal Large Language Models Planting a SEED of Vision in Large Language Model

Reference 59

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T20:22:56.374615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-14T20:13:18.813131Z digest=sha256:c251b0de2ad0a465a5688a3758e33c68daafc37a627ce49522f7c2beee32d23f

Observation 24696ebf-5e9e-48c5-9253-40d51874cfd6 · inbound

When Recovery Matters: The Blind Spot of Surrogate Privacy in MLLM Editing cites this paper.

When Recovery Matters: The Blind Spot of Surrogate Privacy in MLLM Editing Planting a SEED of Vision in Large Language Model

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T17:07:13.132709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-27T22:08:57.792229Z digest=sha256:0bc053903f5f031238d3d4a0aa46ef9441a09c5acc16b820feba5a719da40f8b

Observation 4df426a2-fa2a-467a-9f90-1df208e42540 · inbound

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes cites this paper.

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes Planting a SEED of Vision in Large Language Model

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T11:55:27.864554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:55:27.864554Z digest=sha256:96bbc3c307c302ed07da44099e286c9ebe5b7e445575d13fd7f4c5af8a0939ab

Observation c015a76d-6c53-43ef-b39f-fca19bfc22c3 · inbound

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes cites this paper.

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes Planting a SEED of Vision in Large Language Model

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T17:08:54.589720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:08:54.589720Z digest=sha256:9c1b105a50f7b267c0063fee6ea05622847d962ec700fbc151985ccb463560d4