Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2303.05657.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-12T16:22:44.297638Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-22T23:02:13.659646Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 5d6d4010-0a8e-4b89-a0fd-e76c14b4cc0b · inbound
VideoChat: Chat-Centric Video Understanding Tag2Text: Guiding Vision-Language Model via Image Tagging
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 7b59a083-4cb4-41c2-a172-3ac2e76c5913 · inbound
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Tag2Text: Guiding Vision-Language Model via Image Tagging
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d868495b-9719-45de-b4e5-5c3ac49776d7 · inbound
SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension Tag2Text: Guiding Vision-Language Model via Image Tagging
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation bd29317d-ac9d-4833-a1d1-e30f80578013 · inbound
VBench++: Comprehensive and Versatile Benchmark Suite for Video Generative Models Tag2Text: Guiding Vision-Language Model via Image Tagging
Reference 150
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5458b101-cb05-4f45-9869-7f477235f072 · inbound
Exploring Aleatoric Uncertainty in Object Detection via Vision Foundation Models Tag2Text: Guiding Vision-Language Model via Image Tagging
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2cea303c-2bea-4d47-b397-696cac09c554 · inbound
Detailed Object Description with Controllable Dimensions Tag2Text: Guiding Vision-Language Model via Image Tagging
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb5d9ed6-7bbe-4ceb-bbc1-ad7da6072e75 · inbound
Hallucination Elimination and Semantic Enhancement Framework for Vision-Language Models in Traffic Scenarios Tag2Text: Guiding Vision-Language Model via Image Tagging
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af232d5a-a921-430e-8390-9b06824f6370 · inbound
Object Style Diffusion for Generalized Object Detection in Urban Scene Tag2Text: Guiding Vision-Language Model via Image Tagging
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d75b84b2-9fa2-40bf-9032-5f11a14aa1e2 · inbound
An Empirical Study of Validating Synthetic Data for Text-Based Person Retrieval Tag2Text: Guiding Vision-Language Model via Image Tagging
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 7463baa3-84c9-4423-8d52-7d1e6ef6921c · inbound
Seamless and Efficient Interactions within a Mixed-Dimensional Information Space Tag2Text: Guiding Vision-Language Model via Image Tagging
Reference 149
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2df76ad5-470a-462b-9220-e003ad1b6221 · inbound
See Different, Think Better: Visual Variations Mitigating Hallucinations in LVLMs Tag2Text: Guiding Vision-Language Model via Image Tagging
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.