Pith. sign in

Paper Citation Record · LEDGER

DenseFusion-1M: Merging Vision Experts for Comprehensive Multimodal Perception

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2407.08303.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.08303 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T11:41:44.453339Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T00:17:29.025254Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 6f81342e-bd97-4abe-baba-5a01192baf65 · inbound

Emu3: Next-Token Prediction is All You Need cites this paper.

Emu3: Next-Token Prediction is All You Need DenseFusion-1M: Merging Vision Experts for Comprehensive Multimodal Perception

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:56:08.956957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T10:56:06.418360Z digest=sha256:ea002871ddc6a4af6f0f132f8090b52ba830f3906176fd332a5605b4d362b838

Observation a8f2966c-ba4d-41d2-9f6c-40cda4fe35e9 · inbound

Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation cites this paper.

Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation DenseFusion-1M: Merging Vision Experts for Comprehensive Multimodal Perception

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:09:16.163144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T22:09:16.001309Z digest=sha256:5a024dc62f4f6b0e4803eeb03a9f26f44799cf8be76688013c30a6bf33448607

Observation b65a97ec-b51e-406b-b27d-24f6728f9732 · inbound

DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding cites this paper.

DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding DenseFusion-1M: Merging Vision Experts for Comprehensive Multimodal Perception

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:09:24.036364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T10:09:21.542356Z digest=sha256:69d5af554b5291377d285f7152763001bd4f3df1be6b2029ae07e7a7377842af

Observation 1c76186e-249a-4591-8610-c7f7bf3feff7 · inbound

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding cites this paper.

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding DenseFusion-1M: Merging Vision Experts for Comprehensive Multimodal Perception

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:19:59.819018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T01:19:59.603343Z digest=sha256:c212c796a1c9d0e0e5d79cafe5d9770a5ac0a74af80b1a65e314fb14c23e312d

Observation 43b09f0f-cfdd-4c8f-9c43-1143b7357f62 · inbound

COCONut-PanCap: Joint Panoptic Segmentation and Grounded Captions for Fine-Grained Understanding and Generation cites this paper.

COCONut-PanCap: Joint Panoptic Segmentation and Grounded Captions for Fine-Grained Understanding and Generation DenseFusion-1M: Merging Vision Experts for Comprehensive Multimodal Perception

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-09T11:41:44.453339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:41:44.453339Z digest=sha256:395b2a8a9926dc534e942634f2e4a834a461ad940a6d7fcbce5a1d76aa1cef42

Observation e5b52559-e7e7-42ee-8c3b-0be5be797755 · inbound

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models cites this paper.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models DenseFusion-1M: Merging Vision Experts for Comprehensive Multimodal Perception

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.743104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.743104Z digest=sha256:2f6c81619ed726f741c97daf60e7255902da304ec8d08d3e2bc28c2a6b08448e

Observation f93e43ac-949c-4386-9124-a7ab07fdf843 · inbound

Show-o2: Improved Native Unified Multimodal Models cites this paper.

Show-o2: Improved Native Unified Multimodal Models DenseFusion-1M: Merging Vision Experts for Comprehensive Multimodal Perception

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:51:15.991623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T18:51:15.428692Z digest=sha256:6dac2c4354763c18c0de927fb0a6922296ebb3b4c95cb15a132af9b1d70fbce5

Observation e3aa9530-bfdd-4844-9640-dd64e7f01dc9 · inbound

GenRecal: Generation after Recalibration from Large to Small Vision-Language Models cites this paper.

GenRecal: Generation after Recalibration from Large to Small Vision-Language Models DenseFusion-1M: Merging Vision Experts for Comprehensive Multimodal Perception

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T23:57:24.067810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:57:24.067810Z digest=sha256:7d4ea3b342c07bab1b9c7f1ccf8b128892ab768d360929c865505e3c7be52be2

Observation ea64e559-00bb-49fa-9315-8e5b1445afbb · inbound

OmniGen2: Towards Instruction-Aligned Multimodal Generation cites this paper.

OmniGen2: Towards Instruction-Aligned Multimodal Generation DenseFusion-1M: Merging Vision Experts for Comprehensive Multimodal Perception

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:52:10.872328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:4f9a64fbb400d6c77b334b1702a9a5e4ac7333c6debd414aa794a6520203eb1d

Observation 75fc37fc-9c21-4d22-829f-a0592d761a66 · inbound

Efficient Deployment of Vision-Language Models on Mobile Devices: A Case Study on OnePlus 13R cites this paper.

Efficient Deployment of Vision-Language Models on Mobile Devices: A Case Study on OnePlus 13R DenseFusion-1M: Merging Vision Experts for Comprehensive Multimodal Perception

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T18:20:53.520305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:20:53.520305Z digest=sha256:4b33dff992000c1a34bf0ccf0be7f2a7af4ce129d28846a7203105a90306bbd9

Observation 736aeb7c-7f49-471f-88a5-e1ac314cf19b · inbound

CapRL++: Unified Reinforcement Learning with Verifiable Rewards for Dense Image and Video Captioning cites this paper.

CapRL++: Unified Reinforcement Learning with Verifiable Rewards for Dense Image and Video Captioning DenseFusion-1M: Merging Vision Experts for Comprehensive Multimodal Perception

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:17:29.026830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T17:21:38.543724Z digest=sha256:5ce40c16cc25b087ef7e6f50745d7110f7737c8bf50e54377531010def509b7c