Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 15 inbound Pith citation observations for arXiv:2406.10328.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T13:09:56.862819Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-06-29T13:23:28.438005Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation d15b7e7e-f44d-4dd6-9422-b149b8c24e0a · inbound
Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation From Pixels to Prose: A Large Dataset of Dense Image Captions
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fdf7a16c-0e59-4fc1-91d5-130f6ccbcf13 · inbound
DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding From Pixels to Prose: A Large Dataset of Dense Image Captions
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3e95b839-3773-4b95-99c6-bf0a01e15116 · inbound
FLARE: Fully Integration of Vision-Language Representations for Deep Cross-Modal Understanding From Pixels to Prose: A Large Dataset of Dense Image Captions
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8f4a2872-484e-4da9-85af-e0d56c56e405 · inbound
Chain-of-Talkers (CoTalk): Fast Human Annotation of Dense Image Captions From Pixels to Prose: A Large Dataset of Dense Image Captions
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2cc224ce-97fb-4208-a204-b9c5ec103c97 · inbound
Document-Level Text Generation with Minimum Bayes Risk Decoding using Optimal Transport From Pixels to Prose: A Large Dataset of Dense Image Captions
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4aaaa547-1faf-4365-b5ba-e27ee147e874 · inbound
MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings From Pixels to Prose: A Large Dataset of Dense Image Captions
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation becd85ce-93b1-4790-9592-8da2c45b33fb · inbound
MS-DPPs: Multi-Source Determinantal Point Processes for Contextual Diversity Refinement of Composite Attributes in Text to Image Retrieval From Pixels to Prose: A Large Dataset of Dense Image Captions
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce87387f-6774-4174-adf5-76bfa3e64914 · inbound
ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation From Pixels to Prose: A Large Dataset of Dense Image Captions
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93220cf9-883b-4f79-8d90-b1cfa87f3d93 · inbound
Transition Models: Rethinking the Generative Learning Objective From Pixels to Prose: A Large Dataset of Dense Image Captions
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a40919a6-6724-4fa7-b093-3f5f9d14a557 · inbound
FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark From Pixels to Prose: A Large Dataset of Dense Image Captions
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e3a1767-f545-4f45-b653-e5e7622b1d96 · inbound
WinTok: A Win-Win Hybrid Tokenizer via Decomposing Visual Understanding and Generation with Transferable Tokens From Pixels to Prose: A Large Dataset of Dense Image Captions
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6fb820e6-d82a-455e-ac0b-fcf50011016b · inbound
BEiTScore: Reference-free Image Captioning Evaluation with an Efficient Cross-Encoder Model From Pixels to Prose: A Large Dataset of Dense Image Captions
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9416758b-e8e1-4ecd-9f1b-7af77abf22f4 · inbound
VCap: Hypergeometric Rewards for Weak-to-Strong Visual Captioning From Pixels to Prose: A Large Dataset of Dense Image Captions
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f2f6443d-f7f4-4dff-bcb1-acd6bcf3d080 · inbound
Boogu-Image-0.1: Boosting Open Agentic Multimodal Generation via Understanding under a Minimal Budget From Pixels to Prose: A Large Dataset of Dense Image Captions
Reference 114
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99b7a9c9-e0ce-4fd2-9a40-4f126a3beb0d · inbound
MIDAL: A Dataset of Math Image Descriptions for Accessible Learning From Pixels to Prose: A Large Dataset of Dense Image Captions
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.