Pith. sign in

Paper Citation Record · LEDGER

DenseFusion-1M: Merging Vision Experts for Comprehensive Multimodal Perception

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2407.08303.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.08303 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T11:41:44.453339Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T00:17:29.025254Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 6f81342e-bd97-4abe-baba-5a01192baf65 · inbound

Emu3: Next-Token Prediction is All You Need cites this paper.

Emu3: Next-Token Prediction is All You Need DenseFusion-1M: Merging Vision Experts for Comprehensive Multimodal Perception

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:56:08.956957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T10:56:06.418360Z digest=sha256:f088d3964d1e0681ed3572674ca43c37b2846769db083df027fe741f0cb1ab7a

Observation a8f2966c-ba4d-41d2-9f6c-40cda4fe35e9 · inbound

Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation cites this paper.

Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation DenseFusion-1M: Merging Vision Experts for Comprehensive Multimodal Perception

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:09:16.163144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T22:09:16.001309Z digest=sha256:61c4692c41c2eaebe675317d0da4908e2c9881a347d8c218cb16c681a7170686

Observation b65a97ec-b51e-406b-b27d-24f6728f9732 · inbound

DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding cites this paper.

DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding DenseFusion-1M: Merging Vision Experts for Comprehensive Multimodal Perception

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:09:24.036364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T10:09:21.542356Z digest=sha256:3950ce1b6dd2b4fb747b0d9076d336384fa091939a080483f021f366a818fa9f

Observation 1c76186e-249a-4591-8610-c7f7bf3feff7 · inbound

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding cites this paper.

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding DenseFusion-1M: Merging Vision Experts for Comprehensive Multimodal Perception

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:19:59.819018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T01:19:59.603343Z digest=sha256:a94fc9f62bdcbcf075d2e7c7d428df5ae7c258a9fe295c72174b2090cd8ade56

Observation 43b09f0f-cfdd-4c8f-9c43-1143b7357f62 · inbound

COCONut-PanCap: Joint Panoptic Segmentation and Grounded Captions for Fine-Grained Understanding and Generation cites this paper.

COCONut-PanCap: Joint Panoptic Segmentation and Grounded Captions for Fine-Grained Understanding and Generation DenseFusion-1M: Merging Vision Experts for Comprehensive Multimodal Perception

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-09T11:41:44.453339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:41:44.453339Z digest=sha256:b8a7e07554d6b8ee3753cde96da5c6af951a2ea8d97d2aeecd4def2583f46d49

Observation e5b52559-e7e7-42ee-8c3b-0be5be797755 · inbound

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models cites this paper.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models DenseFusion-1M: Merging Vision Experts for Comprehensive Multimodal Perception

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.743104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.743104Z digest=sha256:5b5e5c71ff29d687dc3c047700e93963bdb03b61d1183d41657c5085242481db

Observation f93e43ac-949c-4386-9124-a7ab07fdf843 · inbound

Show-o2: Improved Native Unified Multimodal Models cites this paper.

Show-o2: Improved Native Unified Multimodal Models DenseFusion-1M: Merging Vision Experts for Comprehensive Multimodal Perception

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:51:15.991623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T18:51:15.428692Z digest=sha256:b4acda7453bd82dba6d82d5ee10f32d9287a23e0dd68cd05df4d51cb7abe90fd

Observation e3aa9530-bfdd-4844-9640-dd64e7f01dc9 · inbound

GenRecal: Generation after Recalibration from Large to Small Vision-Language Models cites this paper.

GenRecal: Generation after Recalibration from Large to Small Vision-Language Models DenseFusion-1M: Merging Vision Experts for Comprehensive Multimodal Perception

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T23:57:24.067810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:57:24.067810Z digest=sha256:d50d0c80cbb57b511bc57387d6b59fe2f071709860e32b277fbf64f9b71f3a55

Observation ea64e559-00bb-49fa-9315-8e5b1445afbb · inbound

OmniGen2: Towards Instruction-Aligned Multimodal Generation cites this paper.

OmniGen2: Towards Instruction-Aligned Multimodal Generation DenseFusion-1M: Merging Vision Experts for Comprehensive Multimodal Perception

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:52:10.872328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:12b3371c0e3dbdfed9c428b3763c2cedc4b2e3c5b6f8a7d84f8d53aa60a7d82a

Observation 75fc37fc-9c21-4d22-829f-a0592d761a66 · inbound

Efficient Deployment of Vision-Language Models on Mobile Devices: A Case Study on OnePlus 13R cites this paper.

Efficient Deployment of Vision-Language Models on Mobile Devices: A Case Study on OnePlus 13R DenseFusion-1M: Merging Vision Experts for Comprehensive Multimodal Perception

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T18:20:53.520305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:20:53.520305Z digest=sha256:9c040be4bc3e9ce224f49d8bfbbeb3753dbc883007c9b4c0e34d64bc154a66ab

Observation 736aeb7c-7f49-471f-88a5-e1ac314cf19b · inbound

CapRL++: Unified Reinforcement Learning with Verifiable Rewards for Dense Image and Video Captioning cites this paper.

CapRL++: Unified Reinforcement Learning with Verifiable Rewards for Dense Image and Video Captioning DenseFusion-1M: Merging Vision Experts for Comprehensive Multimodal Perception

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:17:29.026830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T17:21:38.543724Z digest=sha256:ec486b5750d74e352c7efd2efc35938c032ff9825a182fac6f4812ab8206346b