Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T05:59:15.664441Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 0 inbound Pith citation observations for arXiv:2508.01008.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T05:59:15.664441Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
62 of 62 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 77de1c1f-5091-4938-912a-95ea6d7ab884 · outbound
ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f33e43e0-5ac2-4c4c-8846-dc8f5da62df4 · outbound
ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation Multidiffusion: Fusing diffusion paths for controlled image generation, 2023
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ef88c5b7-ba2f-4644-8814-2ded100596c5 · outbound
ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation Coyo-700m: Image-text pair dataset
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3e6b6bcd-aa20-4a8d-b405-71ccf698ba34 · outbound
ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation InternLM2 Technical Report
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a802e98-d67a-45ec-9e3b-86fedc30fa33 · outbound
ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a57f74c1-180f-4dfc-b297-6559e7eb33a7 · outbound
ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation Training-free layout control with cross-attention guidance
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2701598b-bad2-4951-823f-fd574c4888c2 · outbound
ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fff1f503-be48-4288-bdd8-83c895fcb5e5 · outbound
ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation Yolo-world: Real-time open-vocabulary object detection
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7b7cdf4a-83c0-4d7a-b55d-36bb5c69dc3c · outbound
ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation Laion pop: 600,000 high-resolution images with detailed descriptions
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 50095dc7-82b3-46e9-aa8d-9d29971cf9c3 · outbound
ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation The Llama 3 Herd of Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9438a0ce-ce74-4342-978f-57d03d86a504 · outbound
ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation Scaling recti- fied flow transformers for high-resolution image synthesis
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 93e15a82-0e6c-4825-9092-8b0da2d411f8 · outbound
ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation Ranni: Taming text-to-image diffu- sion for accurate instruction following
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e6dc0a98-1f1e-4dd8-af7e-2e8d392853db · outbound
ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation ImageInWords: Unlocking Hyper-Detailed Image Descriptions
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a4d026d-e51f-4c69-a263-2353e07f88ec · outbound
ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec4ef736-a22c-4e54-80d8-2e2fcf0e8f87 · outbound
ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation Generative adversarial networks
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d6ebf94-b546-4cf6-a2f0-048608780046 · outbound
ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a68c98e0-5d59-45c3-bc17-f5f3da9c8edc · outbound
ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation LVIS: A dataset for large vocabulary instance segmentation
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2e451c29-f81e-4700-93a1-111666ef7e24 · outbound
ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation Retrieval-augmented open-vocabulary object detec- tion
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 56937916-9015-48ff-8881-f86b7f20af56 · outbound
ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation Dense text-to-image generation with attention modulation
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9f95e6be-3bc2-4573-a786-0fcd6d081036 · outbound
ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 464d037d-087b-49e0-9697-6f1fcd9db6f0 · outbound
ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation Unresolved cited work
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d3545566-8d32-40b4-aa70-4d3e0ffd41e4 · outbound
ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6613a77-2f33-4490-a7b4-c1de1d298e74 · outbound
ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation Grounded language-image pre-training
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65e9596d-df25-4270-9e01-7f6ddaa0cb37 · outbound
ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation Autoregressive Image Generation without Vector Quantization
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 584874c4-b51e-4df2-be1c-ad3ddeaf8f06 · outbound
ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation Gligen: Open-set grounded text-to-image generation
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0a12d2be-2083-4d49-a18a-e96e4b528ae4 · outbound
ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8cef9ff0-f18e-4e6d-8890-d331dedf1ee3 · outbound
ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation Lawrence Zitnick
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2972c372-c46f-4923-8e19-7c1e3701b694 · outbound
ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation Visual instruction tuning
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 799ba3a8-7c66-47d6-a6db-eaa58a0cc94a · outbound
ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 798db917-1c23-4eec-abe9-48093b64d801 · outbound
ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation Scaling open-vocabulary object detection
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c1b18570-5a77-4e14-9520-86630e70cc47 · outbound
ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation DOCCI: Descriptions of Connected and Contrasting Images
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85db6f9c-64c4-4290-9846-e28f3d23432d · outbound
ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation Gpt-4v(ision) system card
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98b81ee6-cd5e-49ce-9fd9-270e1e435e55 · outbound
ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation Scalable diffusion models with transformers
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1837cca4-5c02-4d3f-8e1d-aaf441661d65 · outbound
ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation Grounded text-to-image synthesis with attention refocusing
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 908b97f6-22de-4fb2-afdc-1d8c298085bc · outbound
ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7024cd76-2a0d-4eac-b776-c3a1ede5c2a1 · outbound
ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation Stanza: A Python Natural Language Processing Toolkit for Many Human Languages
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation afaf257b-0d06-4304-b251-59a663769a15 · outbound
ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation Learning transferable visual models from natural language supervi- sion
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37938214-507b-4431-8e29-4124034599db · outbound
ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation Zero-shot text-to-image generation
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8a4b91e-1cbe-4866-8916-23bf32daaf28 · outbound
ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation Stable diffusion v1.4 checkpoint
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e9e2b47d-7207-4603-92cc-7fafccf75e71 · outbound
ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation High-resolution image synthesis with latent diffusion models
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9d4a3a27-5444-48bf-a621-28759a4034b5 · outbound
ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation Laion-aesthetics
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f4fa66e1-ddd9-4b72-a679-bfe17225964c · outbound
ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation Laion-5b: An open large-scale dataset for training next generation image-text models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d9d324e-7f0f-45f4-b0ec-5ecbacbdd76f · outbound
ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation Objects365: A large-scale, high-quality dataset for object detection
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ce87387f-6774-4174-adf5-76bfa3e64914 · outbound
ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation From Pixels to Prose: A Large Dataset of Dense Image Captions
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2d401de-1157-4434-96b8-28f7dea11fce · outbound
ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38c09d6d-ded2-4e7b-b0ad-e15987f7c9b6 · outbound
ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9742e30f-4e89-4261-add6-acde697eeb25 · outbound
ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation OV-DINO: Unified Open-Vocabulary Detection with Language-Aware Selective Fusion
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f6a7ee7-658f-4415-8b68-9d49ee5de36f · outbound
ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation V3det: Vast vocabulary visual detection dataset
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bfc92a07-7475-461a-83ab-b3b935a2b4bb · outbound
ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33c7eb08-1cfd-4758-a1b4-9805c31cf4a4 · outbound
ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation CogVLM: Visual Expert for Pretrained Language Models
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96fecee6-52b3-4c19-86cc-d243bece0746 · outbound
ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation Instancediffusion: Instance-level control for image generation, 2024
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 20414764-bd0a-4c08-859f-ab2748069f4c · outbound
ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation IFAdapter: Instance Feature Control for Grounded Text-to-Image Generation
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47190e2e-cb6f-487f-bcd1-6790ca4821e5 · outbound
ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation Boxdiff: Text-to-image synthesis with training-free box-constrained diffusion
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 11966a8c-fafa-4e4c-96d9-89124b9e32c0 · outbound
ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation Qwen2 Technical Report
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 022bea05-3bf8-41fd-8391-c2ac7527aef8 · outbound
ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation Detclipv3: To- wards versatile generative open-vocabulary object detection
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 75f515f5-2b13-4031-b42c-f47808b2960d · outbound
ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation Scaling Autoregressive Models for Content-Rich Text-to-Image Generation
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ecd1a846-9a9e-466a-abba-79e3f324ca89 · outbound
ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation Glipv2: Unifying localiza- tion and vision-language understanding
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7288b7b2-a5be-45c4-afd7-53cb58ab7dde · outbound
ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation Recognize anything: A strong image tagging model
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7edd6207-3e5e-40d8-8ede-308d5e5fdb1f · outbound
ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation Taming Self-Training for Open-Vocabulary Object Detection
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a440a126-4c4c-4fcc-bac8-52a6dfdd7968 · outbound
ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation Migc++: Advanced multi-instance generation controller for image synthesis, 2024
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f78d3fea-f5e6-4e64-b566-b5aa1489c56c · outbound
ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation Migc: Multi-instance generation controller for text-to-image synthesis, 2024
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 13b8e0a3-f0bb-44e8-b3f7-b80e4667dd9b · outbound
ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation Unresolved cited work
Reference 2014
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.