Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T17:10:22.228795Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 0 inbound Pith citation observations for arXiv:2412.09283.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T17:10:22.228795Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
60 of 60 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation f098c3e9-9f2f-4ffc-ae90-4dc2769f8607 · outbound
InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Chen and William B
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d57e239f-5270-4cfb-9c6a-ac55c55d4ef2 · outbound
InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption VideoCrafter2: Overcoming data limitations for high-quality video diffusion models
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 4ebfefe2-7ddf-48c5-b74f-a80380af84ad · outbound
InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85085582-55ce-400c-b4df-279fb124da4a · outbound
InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d23350ec-f60d-4ba0-b2ce-ae3afb25ebc2 · outbound
InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption T2I-CompBench++: An Enhanced and Comprehensive Benchmark for Compositional Text-to-image Generation
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ec20df2-4806-43e7-802c-3ede21157525 · outbound
InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption VBench: Com- prehensive benchmark suite for video generative models
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 8edaebb5-2364-49c7-81e8-10cffce4725a · outbound
InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Survey of hallucination in natural language generation
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 4c7fbfbb-a11c-4a9e-9d90-b56edcc76973 · outbound
InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Pyramidal flow matching for efficient video generative modeling
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5af02d9c-2cc1-411a-8e94-3e020a7ddb98 · outbound
InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption MiraData: A Large-Scale Video Dataset with Long Durations and Structured Captions
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6118f070-b677-4b94-bfa0-8843cfd55cb1 · outbound
InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Unresolved cited work
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 77183dd7-47d2-40e3-b7b2-6b85611a20a7 · outbound
InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Open-sora-plan, 2024
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation af50f70e-e17f-423e-ae7f-d685af56199d · outbound
InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption T2v- turbo-v2: Enhancing video generation model post-training through data, reward, and conditional guidance design.arXiv preprint arXiv:2410.05677, 2024
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b43d96d-9d72-4105-866b-a4156803ea0e · outbound
InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Evaluating text-to-visual generation with image-to-text gen- eration, 2024
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 73ec0bcf-72e9-4f34-a786-a13eef6bfc78 · outbound
InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption EvalCrafter: Benchmarking and Evaluating Large Video Generation Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd21542b-0112-4ad1-99bd-b9d3eb2f66ab · outbound
InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Latte: Latent Diffusion Transformer for Video Generation
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c2e2bc9-c3de-4886-87df-a7dc1680c0ce · outbound
InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc576829-b7bc-4388-b783-e820dfee3991 · outbound
InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Animal kingdom: A large and diverse dataset for animal behavior understanding
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 0054ac26-a500-4ce8-910c-64dd8e702de0 · outbound
InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Pika 1.0
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 0ec2a843-92e2-4d53-a88b-bc84f7fd44a4 · outbound
InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Learning transferable visual models from natural language supervision, 2021
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 9bea47ba-645f-4e3f-b1da-5748cc258da8 · outbound
InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption SAM 2: Segment Anything in Images and Videos
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed6640e2-e24f-4d5d-aa8e-584e8ddaf2be · outbound
InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Unresolved cited work
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d9812e93-3ff8-40bc-bfca-468994f453ef · outbound
InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption What does clip know about a red circle? vi- sual prompt engineering for vlms
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation fee51e57-7e0b-4dac-aff5-61ac3bdbaa7a · outbound
InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption ModelScope Text-to-Video Technical Report
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e21e2c1-3a3e-4bb0-bc4c-0e12308fd022 · outbound
InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Vatex: A large-scale, high- quality multilingual dataset for video-and-language research,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation aa616b46-75ad-4d33-9d72-c21d88d7e1c5 · outbound
InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption LAVIE: High-Quality Video Generation with Cascaded Latent Diffusion Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c4a398b-520a-4cc2-b9ba-31dfd64e4ff7 · outbound
InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Internvid: A large-scale video-text dataset for multimodal understanding and generation
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 95c1768c-6195-40ed-9b97-b37311e9020d · outbound
InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Ifadapter: Instance feature control for grounded text-to-image generation, 2024
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 6ac07ef2-95a5-4b98-a847-e3c1166cad5c · outbound
InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Vript: A video is worth thousands of words, 2024
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a450c8be-4544-4f05-b717-c9bd27bf6505 · outbound
InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Mastering text-to-image dif- fusion: Recaptioning, planning, and generating with multi- modal llms
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 68f197f2-abba-4d39-9807-5a81253164fe · outbound
InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 031a6f72-fe6e-42f0-ade5-027b6ab50fe2 · outbound
InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Cpt: Colorful prompt tuning for pre-trained vision-language models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5fa98523-e966-4445-9cde-15042c4c94a0 · outbound
InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Show-1: Marrying pixel and latent diffusion models for text-to-video generation
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d19740f5-efad-4cb3-a74c-17647e40ea17 · outbound
InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Efros, Eli Shecht- man, and Oliver Wang
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation ca537b78-a784-43d9-9adb-6e20807cc484 · outbound
InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Video instruction tuning with synthetic data, 2024
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation f809e469-f5a1-44f6-a958-4e226d7f06e2 · outbound
InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Open-sora: Democratizing efficient video production for all, 2024
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation cc17b6b7-170b-4017-870e-873af743e6aa · outbound
InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Open-Sora: Democratizing efficient video production for all
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 95ddd32b-5419-4f6e-bcbd-1f2343d8fd40 · outbound
InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Unresolved cited work
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 270d6517-3512-4130-80b8-5de4d34bb86d · outbound
InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Detrs with col- laborative hybrid assignments training
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 318e6561-5430-4add-b673-a6b1154ef611 · outbound
InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Conversely, we manually constructed a Negative Lexicon, which was further enriched using the powerful LLM, GPT-4o
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 73d602a7-17e8-4818-a477-2f4a5bf17aef · outbound
InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Please describe the car by its color, make, model, condition, license plate (if visible), and any distinguishing features such as stickers, dents, or modifications
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 231bd2a2-cb13-47c6-85f5-b7a4a6953f5b · outbound
InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Please describe this video in one sentence, no more than 20 words
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 57eef00a-8192-48ac-9c89-8e9d1f2d0758 · outbound
InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption To pro- vide more precise instructions to LLMs, we meticulously designed multiple examples as part of the CoT, which are fed into the LLMs
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 05cacb38-087a-4f84-9b17-3324de6d7a87 · outbound
InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption subject" +
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 88254254-5fef-4b45-a053-24043bc1c34b · outbound
InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Unresolved cited work
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 37a2c3dc-17c1-47f4-ac61-a21a0412bb15 · outbound
InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Unresolved cited work
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 86159046-675d-4300-944e-ae7f004f9195 · outbound
InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Unresolved cited work
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 4b89d5ee-6326-4aa5-ad58-fce3b1848ceb · outbound
InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Unresolved cited work
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 09805f3b-88cf-4c9d-ad31-12813f0b514a · outbound
InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Unresolved cited work
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 68c54bb7-34e2-47d1-9e7b-eb28801a1685 · outbound
InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Unresolved cited work
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 2733460a-5410-4e53-a018-d61641052d41 · outbound
InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Unresolved cited work
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 97ad5e53-6765-42f1-98b7-882556d71f13 · outbound
InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Unresolved cited work
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation ea38dc8d-3864-4df4-9681-ab08e95dbe5e · outbound
InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Global Description
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 6b3ff845-bbdb-4d94-a251-0079f1e3e7a1 · outbound
InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption 3) Ex- trinsic Hallucination: Evaluate whether the text introduces content that is not present in the video
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c46daef7-2472-4254-905a-553e26f6af1c · outbound
InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption counter-intuitive
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 02fd437d-b1f0-4537-bef6-fd4705d237b6 · outbound
InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption ,".join([f
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation f2da4476-454b-4771-bfa6-fb07cec18eee · outbound
InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption subject" +
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 90cdd62c-75c5-439f-9dd4-c3f4f6cc0e8f · outbound
InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption The video shows
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation b486091b-22ac-4c79-92e5-72752dcde0a6 · outbound
InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption A man... and a woman..., and a man
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e2d51b3a-c5fb-4e9b-ba52-394fe430e6eb · outbound
InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption The scene is
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a1f40f2a-eb5d-465d-ae52-8fc043bcb8c4 · outbound
InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Aligning prompt used during alignment with the open source model
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
No inbound Pith citation observations are available.