Pith. sign in

Paper Citation Record · LEDGER

Planting a SEED of Vision in Large Language Model

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 21 inbound Pith citation observations for arXiv:2307.08041.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2307.08041 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 21 of 21 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T17:08:54.589720Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

10
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d3fe3931-6c77-4fb0-ae5c-027a8e27dca6 · inbound

SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension cites this paper.

SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension Planting a SEED of Vision in Large Language Model

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T16:59:50.525873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T16:59:50.495335Z digest=sha256:b58ccb8cdbd01a7e2485e3c3784e8a849dcd8d6c53415cfebfdd63ca9f1c121c

Observation 26d03b63-a1bc-4499-97de-22bfc5ebf43f · inbound

DragNUWA: Fine-grained Control in Video Generation by Integrating Text, Image, and Trajectory cites this paper.

DragNUWA: Fine-grained Control in Video Generation by Integrating Text, Image, and Trajectory Planting a SEED of Vision in Large Language Model

Reference 258

Resolution
verified exact
arxiv_id, observed 2026-05-20T13:03:58.025865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-20T13:03:57.828598Z digest=sha256:f9cb877b3b33b806caf03a20562b2bb35b022b9d639400b0b5aca1f34976f4f7

Observation d8eff912-36a9-4665-8a80-3d3ee5cdaeb6 · inbound

Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models cites this paper.

Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models Planting a SEED of Vision in Large Language Model

Reference 48

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T07:44:47.448796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T07:44:47.355960Z digest=sha256:d30723e0e6570f0c092fc94a5f5c336a561db8d3a8e4750a422799833b2cd9ce

Observation 8bc9297e-4349-43c7-891b-fde40426852c · inbound

SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation cites this paper.

SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation Planting a SEED of Vision in Large Language Model

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T22:48:36.350446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T22:48:36.010306Z digest=sha256:cdb942383c7bed8f50f918a3271f9cefd9a94ecbd66d02693300f6cc932cb019

Observation b9fe7687-ffd6-4a6c-b006-a12f68c18305 · inbound

Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs cites this paper.

Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs Planting a SEED of Vision in Large Language Model

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T00:05:03.785210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T00:05:03.547664Z digest=sha256:009a01b8352bb36545a2cbecb09111694a20d554066ea965671920aba242b5fc

Observation e61c6e61-1175-439f-9f01-4d34cc6d17e3 · inbound

VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation cites this paper.

VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation Planting a SEED of Vision in Large Language Model

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T00:26:21.394182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T00:26:21.313005Z digest=sha256:7a95a01231e753c792fcda84952984e1d6b9693cdefc1ce210cbdd57ba554bff

Observation 66db182b-a9ef-4077-b2f4-0460af504fe6 · inbound

Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation cites this paper.

Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation Planting a SEED of Vision in Large Language Model

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:09:16.472580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T22:09:16.001309Z digest=sha256:ca7ac8f413cf5dc42ab07388e4fc59944ea8d9c5fcaec3447c211050d85cba82

Observation ef9d18dc-8bc5-4634-bde8-aa445cfd9309 · inbound

MetaMorph: Multimodal Understanding and Generation via Instruction Tuning cites this paper.

MetaMorph: Multimodal Understanding and Generation via Instruction Tuning Planting a SEED of Vision in Large Language Model

Reference 216

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T07:51:13.261881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-17T07:51:12.953777Z digest=sha256:03058faf7d136964611421d5e4873aded407baba070bf181527feb7699730082

Observation cef15181-dc25-4773-93a6-0df94599faa0 · inbound

HealthGPT: A Medical Large Vision-Language Model for Unifying Comprehension and Generation via Heterogeneous Knowledge Adaptation cites this paper.

HealthGPT: A Medical Large Vision-Language Model for Unifying Comprehension and Generation via Heterogeneous Knowledge Adaptation Planting a SEED of Vision in Large Language Model

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T20:21:30.400783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T20:21:30.400783Z digest=sha256:a4ac59f1076bc5935e4768c0b36fb4e0ac940b3a165178117259d4e9b6fe7889

Observation 7f3c46e5-a305-4c8f-8f23-2cdc7a32c81f · inbound

Transfer between Modalities with MetaQueries cites this paper.

Transfer between Modalities with MetaQueries Planting a SEED of Vision in Large Language Model

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T22:49:23.162924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-14T22:49:23.074271Z digest=sha256:bb2f78b259323e4e4a54622e3a21510687bd15c290fb4830071b884a6e023750

Observation 3d2fe74f-9855-4096-83cc-2da854771b37 · inbound

Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation cites this paper.

Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation Planting a SEED of Vision in Large Language Model

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T07:24:04.821007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T07:24:04.460276Z digest=sha256:75b52e8ac048545a51162a52ab8bf01e9739873a2e7c60b4d3ead4fc85f86bc3

Observation 9121c834-8d28-4554-bafa-d53507f3dc9c · inbound

FUDOKI: Discrete Flow-based Unified Understanding and Generation via Kinetic-Optimal Velocities cites this paper.

FUDOKI: Discrete Flow-based Unified Understanding and Generation via Kinetic-Optimal Velocities Planting a SEED of Vision in Large Language Model

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:04:54.247781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:04:54.247781Z digest=sha256:72c590fd9e3880d89a14b86d89358859541d6f1a20100bf60dba2c9de04d9aae

Observation a314e9b2-206b-4574-adc3-c2b4999361bd · inbound

From Standalone LLMs to Integrated Intelligence: A Survey of Compound Al Systems cites this paper.

From Standalone LLMs to Integrated Intelligence: A Survey of Compound Al Systems Planting a SEED of Vision in Large Language Model

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:52:16.301701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T11:49:36.574471Z digest=sha256:cfbd390677273a2c06448d036e7224c0a444fbc0d61197c8b95fb3aacce029d3

Observation 5244ed3a-2046-4a49-a4ee-77feb14e17cc · inbound

Instella-T2I: Pushing the Limits of 1D Discrete Latent Space Image Generation cites this paper.

Instella-T2I: Pushing the Limits of 1D Discrete Latent Space Image Generation Planting a SEED of Vision in Large Language Model

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T22:43:02.997268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:43:02.997268Z digest=sha256:130c4e4cbaf0c0a647821d7fd681482f98b1015c4af3ebb0129d6b4cef937bb1

Observation ff4062fc-f51c-4b6a-8b42-aeb20bd31974 · inbound

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs cites this paper.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs Planting a SEED of Vision in Large Language Model

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:15.540771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:15.540771Z digest=sha256:2adf5296dd22d2a79e15485f0e18ca5b37273b9a10b3e81f8b63f379d03dfcc9

Observation 81ff7b25-87d0-4a82-9adf-ca1f202ec84a · inbound

Colon-Bench: An Agentic Workflow for Scalable Dense Lesion Annotation in Full-Procedure Colonoscopy Videos cites this paper.

Colon-Bench: An Agentic Workflow for Scalable Dense Lesion Annotation in Full-Procedure Colonoscopy Videos Planting a SEED of Vision in Large Language Model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-13T18:06:18.381075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T18:06:18.381075Z digest=sha256:3c1a2b9987077afc1a0f0419c22ec611ebd7581e781175a021df25b0ded844f6

Observation 4c277ecd-dbf2-4959-a401-7b8862c8793e · inbound

DeepSight: Long-Horizon World Modeling via Latent States Prediction for End-to-End Autonomous Driving cites this paper.

DeepSight: Long-Horizon World Modeling via Latent States Prediction for End-to-End Autonomous Driving Planting a SEED of Vision in Large Language Model

Reference 206

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:31:26.428162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-12T04:13:37.421188Z digest=sha256:4e8e8fda502cb76712ed0c6232636da826abed5f616cc6388c1ffd9f38648c4e

Observation adfc6a0b-763b-49c7-8317-ce7147a39521 · inbound

Learning to See What You Need: Gaze Attention for Multimodal Large Language Models cites this paper.

Learning to See What You Need: Gaze Attention for Multimodal Large Language Models Planting a SEED of Vision in Large Language Model

Reference 59

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T20:22:56.374615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-14T20:13:18.813131Z digest=sha256:3a44ca1af1bf68f3ba40dfc791fef3ae023f86ca8fad10dd6ecb95c65d161b42

Observation 24696ebf-5e9e-48c5-9253-40d51874cfd6 · inbound

When Recovery Matters: The Blind Spot of Surrogate Privacy in MLLM Editing cites this paper.

When Recovery Matters: The Blind Spot of Surrogate Privacy in MLLM Editing Planting a SEED of Vision in Large Language Model

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T17:07:13.132709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T22:08:57.792229Z digest=sha256:4c91aec8e2c8ca8ae40e04e01eefe7a1e628a51ba646ee263326df875270fee1

Observation 4df426a2-fa2a-467a-9f90-1df208e42540 · inbound

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes cites this paper.

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes Planting a SEED of Vision in Large Language Model

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T11:55:27.864554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:55:27.864554Z digest=sha256:e4d7b5f6f8769924db9cb63475bec2e10235b8b4c70c01b427fd5582d9fd04ef

Observation c015a76d-6c53-43ef-b39f-fca19bfc22c3 · inbound

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes cites this paper.

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes Planting a SEED of Vision in Large Language Model

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T17:08:54.589720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:08:54.589720Z digest=sha256:4ee47293daf8bd4e475a392920c86f26bb07315500464d9d38cb602f15bbb53b