Pith. sign in

Paper Citation Record · LEDGER

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption

As of 18 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 0 inbound Pith citation observations for arXiv:2412.09283.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.09283 v1

Coverage vector

measured 60 of 60 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T17:10:22.228795Z

measured 60 of 60 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

60 of 60 outbound references displayed

  • verified exact0
  • verified fuzzy35
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f098c3e9-9f2f-4ffc-ae90-4dc2769f8607 · outbound

This paper cites Chen and William B.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Chen and William B

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:10:53.322697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T17:10:21.885108Z digest=sha256:a614c005818c91dce1e505df9f24b3ea63fbea32d6488fb0402744f49ed3d6ea

Observation d57e239f-5270-4cfb-9c6a-ac55c55d4ef2 · outbound

This paper cites VideoCrafter2: Overcoming data limitations for high-quality video diffusion models.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption VideoCrafter2: Overcoming data limitations for high-quality video diffusion models

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:10:53.306471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T17:10:21.891074Z digest=sha256:45a69da85caf07efac38249177efc9ced640e4c2501940417f35aad568b46499

Observation 4ebfefe2-7ddf-48c5-b74f-a80380af84ad · outbound

This paper cites ShareGPT4Video: Improving Video Understanding and Generation with Better Captions.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T17:10:21.897243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:10:21.897243Z digest=sha256:d75d5e7711a2d8fe1004a6283c06af596bf640f80ebc19dfd728ce23bd42569d

Observation 85085582-55ce-400c-b4df-279fb124da4a · outbound

This paper cites Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T17:10:21.903106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:10:21.903106Z digest=sha256:cd60537c4a6eb512bcb63b992e137f6979b68158a4ff2e5ec945abd91944d604

Observation d23350ec-f60d-4ba0-b2ce-ae3afb25ebc2 · outbound

This paper cites T2I-CompBench++: An Enhanced and Comprehensive Benchmark for Compositional Text-to-image Generation.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption T2I-CompBench++: An Enhanced and Comprehensive Benchmark for Compositional Text-to-image Generation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T17:10:21.909455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:10:21.909455Z digest=sha256:cbbdc42db4a1a519224fcd604e8bfccb9f4307368c75d070393cb97a83eea096

Observation 1ec20df2-4806-43e7-802c-3ede21157525 · outbound

This paper cites VBench: Com- prehensive benchmark suite for video generative models.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption VBench: Com- prehensive benchmark suite for video generative models

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:10:53.290914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T17:10:21.915566Z digest=sha256:caa778525b3f275014a026e0da84c44d03a111ca0cacbe6b5436b4e1c80f523b

Observation 8edaebb5-2364-49c7-81e8-10cffce4725a · outbound

This paper cites Survey of hallucination in natural language generation.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Survey of hallucination in natural language generation

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:10:53.274241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T17:10:21.920916Z digest=sha256:e147db2c5abc383a2fef434ecfb659d0d2e7867aef9ac98dcafa5fffb281e77e

Observation 4c7fbfbb-a11c-4a9e-9d90-b56edcc76973 · outbound

This paper cites Pyramidal flow matching for efficient video generative modeling.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Pyramidal flow matching for efficient video generative modeling

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T17:10:21.927418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:10:21.927418Z digest=sha256:2167768caa645e7b4775b6bd00600834a43c2690a2fe0dfcef29ec9845d88957

Observation 5af02d9c-2cc1-411a-8e94-3e020a7ddb98 · outbound

This paper cites MiraData: A Large-Scale Video Dataset with Long Durations and Structured Captions.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption MiraData: A Large-Scale Video Dataset with Long Durations and Structured Captions

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T17:10:21.932432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:10:21.932432Z digest=sha256:f57ba268f8e8e27323cbaec8321cfd7392e7d2c0c2648abefc6527733d623504

Observation 6118f070-b677-4b94-bfa0-8843cfd55cb1 · outbound

This paper cites an unresolved cited work.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-11T17:10:53.257171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T17:10:21.937711Z digest=sha256:1c09c26dad0621964512bd04ac8bbf1b698facfa22b7b304d746c067779c977a

Observation 77183dd7-47d2-40e3-b7b2-6b85611a20a7 · outbound

This paper cites Open-sora-plan, 2024.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Open-sora-plan, 2024

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:10:53.241808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T17:10:21.942609Z digest=sha256:6a54d742440547ff6dbd0c2ba4a27711a51b80ea550fbb1ec257f39c1d2c909a

Observation af50f70e-e17f-423e-ae7f-d685af56199d · outbound

This paper cites T2v- turbo-v2: Enhancing video generation model post-training through data, reward, and conditional guidance design.arXiv preprint arXiv:2410.05677, 2024.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption T2v- turbo-v2: Enhancing video generation model post-training through data, reward, and conditional guidance design.arXiv preprint arXiv:2410.05677, 2024

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T17:10:21.947570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:10:21.947570Z digest=sha256:1b6e5b9169c506a128a2a3090b7fb9183b6b97a59602c751bca265d1796444c7

Observation 4b43d96d-9d72-4105-866b-a4156803ea0e · outbound

This paper cites Evaluating text-to-visual generation with image-to-text gen- eration, 2024.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Evaluating text-to-visual generation with image-to-text gen- eration, 2024

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:10:53.223906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T17:10:21.952218Z digest=sha256:a5089f2e22b7fb55f2aff0551f0379c4fac6dac05a7343fd88c6b0119e74ae4d

Observation 73ec0bcf-72e9-4f34-a786-a13eef6bfc78 · outbound

This paper cites EvalCrafter: Benchmarking and Evaluating Large Video Generation Models.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption EvalCrafter: Benchmarking and Evaluating Large Video Generation Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T17:10:21.957349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:10:21.957349Z digest=sha256:c8cff230d361693c08e1baefe72e9ac33eb6781abaffd6dcb2af8a6bd03b860c

Observation fd21542b-0112-4ad1-99bd-b9d3eb2f66ab · outbound

This paper cites Latte: Latent Diffusion Transformer for Video Generation.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Latte: Latent Diffusion Transformer for Video Generation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T17:10:21.963009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:10:21.963009Z digest=sha256:4b056309843e886162c8cacb87fb0e1786a27dccf997832d5a8ee71f7dd13eb0

Observation 2c2e2bc9-c3de-4886-87df-a7dc1680c0ce · outbound

This paper cites OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T17:10:21.969388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:10:21.969388Z digest=sha256:da74954e05768913c575e294e0671a66e4771f6d20ece7ef04c21d00c929a66d

Observation fc576829-b7bc-4388-b783-e820dfee3991 · outbound

This paper cites Animal kingdom: A large and diverse dataset for animal behavior understanding.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Animal kingdom: A large and diverse dataset for animal behavior understanding

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:10:53.206937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T17:10:21.975014Z digest=sha256:bd9f50e50664d41bd52822409e9e77afae57af5a9b46d87fce7732ac4775f143

Observation 0054ac26-a500-4ce8-910c-64dd8e702de0 · outbound

This paper cites Pika 1.0.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Pika 1.0

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:10:53.189391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T17:10:21.980801Z digest=sha256:fd51c7e69e955455d08e5d1398d7648eab3bcd9210094ba48b5e76d38101a1a4

Observation 0ec2a843-92e2-4d53-a88b-bc84f7fd44a4 · outbound

This paper cites Learning transferable visual models from natural language supervision, 2021.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Learning transferable visual models from natural language supervision, 2021

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:10:53.172796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T17:10:21.986290Z digest=sha256:124b9cdc57b94ca3095dea0a8825dabf9846d8b58260c6dbda0eb8762f0e1586

Observation 9bea47ba-645f-4e3f-b1da-5748cc258da8 · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption SAM 2: Segment Anything in Images and Videos

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T17:10:21.991371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:10:21.991371Z digest=sha256:6b2e49589f5ce52d626b7986c2948ae06ebd6a3c0444f48e38135964a9665c77

Observation ed6640e2-e24f-4d5d-aa8e-584e8ddaf2be · outbound

This paper cites an unresolved cited work.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-11T17:10:53.156086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T17:10:21.996655Z digest=sha256:19bbd4643e30d423bc90322c4f4da426cd9685b1832c1152851584948c358ef3

Observation d9812e93-3ff8-40bc-bfca-468994f453ef · outbound

This paper cites What does clip know about a red circle? vi- sual prompt engineering for vlms.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption What does clip know about a red circle? vi- sual prompt engineering for vlms

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:10:53.137911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T17:10:22.003117Z digest=sha256:344801d269418440fde2d62810ed7a32cfd255272f8dee89cce92ba5abac0b04

Observation fee51e57-7e0b-4dac-aff5-61ac3bdbaa7a · outbound

This paper cites ModelScope Text-to-Video Technical Report.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption ModelScope Text-to-Video Technical Report

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T17:10:22.008174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:10:22.008174Z digest=sha256:0c8f9dbd41b84dc8b9a6b0010036a1bc1abd1a541f21e2610d0972adccde176a

Observation 9e21e2c1-3a3e-4bb0-bc4c-0e12308fd022 · outbound

This paper cites Vatex: A large-scale, high- quality multilingual dataset for video-and-language research,.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Vatex: A large-scale, high- quality multilingual dataset for video-and-language research,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:10:53.114885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T17:10:22.014362Z digest=sha256:b82539f9762fa70e7af65e31dee06b14cfb483bca3675848bf6132e7e5b6152b

Observation aa616b46-75ad-4d33-9d72-c21d88d7e1c5 · outbound

This paper cites LAVIE: High-Quality Video Generation with Cascaded Latent Diffusion Models.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption LAVIE: High-Quality Video Generation with Cascaded Latent Diffusion Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T17:10:22.020323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:10:22.020323Z digest=sha256:57e9298894103531227388c6bd841b83c50dea096560d0b70be101fb968a25bf

Observation 7c4a398b-520a-4cc2-b9ba-31dfd64e4ff7 · outbound

This paper cites Internvid: A large-scale video-text dataset for multimodal understanding and generation.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Internvid: A large-scale video-text dataset for multimodal understanding and generation

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:10:53.097058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T17:10:22.030087Z digest=sha256:a23ff0c8c1ca643fa700c3ac1c4cfd464bb893e3ddfdb0bb3b63ae352c1f8983

Observation 95c1768c-6195-40ed-9b97-b37311e9020d · outbound

This paper cites Ifadapter: Instance feature control for grounded text-to-image generation, 2024.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Ifadapter: Instance feature control for grounded text-to-image generation, 2024

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:10:53.079417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T17:10:22.037006Z digest=sha256:612d699ead87b58ad4d839e0a895882b39062c527d2d7b90b0cb801d95b1679d

Observation 6ac07ef2-95a5-4b98-a847-e3c1166cad5c · outbound

This paper cites Vript: A video is worth thousands of words, 2024.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Vript: A video is worth thousands of words, 2024

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:10:53.061606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T17:10:22.042805Z digest=sha256:270ebca9a18ec9cf81fb222da20576c434eb368c4a7f21c50938c04e9e8e5fd8

Observation a450c8be-4544-4f05-b717-c9bd27bf6505 · outbound

This paper cites Mastering text-to-image dif- fusion: Recaptioning, planning, and generating with multi- modal llms.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Mastering text-to-image dif- fusion: Recaptioning, planning, and generating with multi- modal llms

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:10:53.045703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T17:10:22.055946Z digest=sha256:2f5f560ae1f89359bcdbf19d774e26c39d063df20d4da533bd1148ec52b9631f

Observation 68f197f2-abba-4d39-9807-5a81253164fe · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T17:10:22.062619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:10:22.062619Z digest=sha256:06d826be58cf0643de45507a2733fbb8f81ea8b76676c1f13ccad4838a5683ab

Observation 031a6f72-fe6e-42f0-ade5-027b6ab50fe2 · outbound

This paper cites Cpt: Colorful prompt tuning for pre-trained vision-language models.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Cpt: Colorful prompt tuning for pre-trained vision-language models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T17:10:22.069047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:10:22.069047Z digest=sha256:9ccf46f21f22430169e54c5594cca5b055a90bef08a66d3b038654552cc783a6

Observation 5fa98523-e966-4445-9cde-15042c4c94a0 · outbound

This paper cites Show-1: Marrying pixel and latent diffusion models for text-to-video generation.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Show-1: Marrying pixel and latent diffusion models for text-to-video generation

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:10:53.018846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T17:10:22.074359Z digest=sha256:642b3c1ec634666b2607df9c8e834c4bf5a14bfff8bc0ae2b5445ac5f50f26b6

Observation d19740f5-efad-4cb3-a74c-17647e40ea17 · outbound

This paper cites Efros, Eli Shecht- man, and Oliver Wang.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Efros, Eli Shecht- man, and Oliver Wang

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:10:53.002158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T17:10:22.081845Z digest=sha256:cfe233cbf4e78f398d73b24628dc92684ed87f0100706716f5e3fd96332f3ef7

Observation ca537b78-a784-43d9-9adb-6e20807cc484 · outbound

This paper cites Video instruction tuning with synthetic data, 2024.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Video instruction tuning with synthetic data, 2024

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:10:52.984990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T17:10:22.087507Z digest=sha256:1c5e97f73a199be94ada7cbd1af118391cb5605eb0132ee923d70b47eef57f1e

Observation f809e469-f5a1-44f6-a958-4e226d7f06e2 · outbound

This paper cites Open-sora: Democratizing efficient video production for all, 2024.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Open-sora: Democratizing efficient video production for all, 2024

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:10:52.968053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T17:10:22.095179Z digest=sha256:38673b8dc5ce0cb42d1309c69a7bcd9c88b2eb080031ce4b3f915c0a1540c848

Observation cc17b6b7-170b-4017-870e-873af743e6aa · outbound

This paper cites Open-Sora: Democratizing efficient video production for all.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Open-Sora: Democratizing efficient video production for all

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:10:52.951776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T17:10:22.101440Z digest=sha256:def8c0c7e7d7faf43b561e4ab76555712f55d52668cfc2179aadd61fd8fa756a

Observation 95ddd32b-5419-4f6e-bcbd-1f2343d8fd40 · outbound

This paper cites an unresolved cited work.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-11T17:10:52.935738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T17:10:22.107772Z digest=sha256:780a145b49eb858cd5f99d98edf2c77fa338be47621299e200679d2247d1ab10

Observation 270d6517-3512-4130-80b8-5de4d34bb86d · outbound

This paper cites Detrs with col- laborative hybrid assignments training.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Detrs with col- laborative hybrid assignments training

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:10:52.917077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T17:10:22.113084Z digest=sha256:5e7a86d8d6d789766f55984ec6f55fe63f090e592bf692918568d85b90b07cd7

Observation 318e6561-5430-4add-b673-a6b1154ef611 · outbound

This paper cites Conversely, we manually constructed a Negative Lexicon, which was further enriched using the powerful LLM, GPT-4o.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Conversely, we manually constructed a Negative Lexicon, which was further enriched using the powerful LLM, GPT-4o

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:10:52.899020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T17:10:22.118425Z digest=sha256:49c8be242d230736eca4e9cc01c020704c97dd32a5ff6f29b858f99eae9799ed

Observation 73d602a7-17e8-4818-a477-2f4a5bf17aef · outbound

This paper cites Please describe the car by its color, make, model, condition, license plate (if visible), and any distinguishing features such as stickers, dents, or modifications.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Please describe the car by its color, make, model, condition, license plate (if visible), and any distinguishing features such as stickers, dents, or modifications

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:10:52.881436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T17:10:22.125141Z digest=sha256:cef69ba91fb7c26a765d5560434937e2e9ab342a47147cd939ac6f33b54648c1

Observation 231bd2a2-cb13-47c6-85f5-b7a4a6953f5b · outbound

This paper cites Please describe this video in one sentence, no more than 20 words.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Please describe this video in one sentence, no more than 20 words

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:10:52.866461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T17:10:22.130340Z digest=sha256:11b616cd546b654106a7c6e53901c7b3ce491063841fab7235580cac173e8304

Observation 57eef00a-8192-48ac-9c89-8e9d1f2d0758 · outbound

This paper cites To pro- vide more precise instructions to LLMs, we meticulously designed multiple examples as part of the CoT, which are fed into the LLMs.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption To pro- vide more precise instructions to LLMs, we meticulously designed multiple examples as part of the CoT, which are fed into the LLMs

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:10:52.851374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T17:10:22.135176Z digest=sha256:10a728b268804267e81bdeaf77d503cdbb33280bdaa01551dc2be9318dc56bef

Observation 05cacb38-087a-4f84-9b17-3324de6d7a87 · outbound

This paper cites subject" +.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption subject" +

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:10:52.835383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T17:10:22.141222Z digest=sha256:ef6cdfc7542bbd952ffc219116fe916ef76ab97cf0f1ea660ffff06f5355fd46

Observation 88254254-5fef-4b45-a053-24043bc1c34b · outbound

This paper cites an unresolved cited work.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-11T17:10:52.818327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T17:10:22.146262Z digest=sha256:0a81ab46bf446cd00d527d50cd11d801d904133fefdfaec212d90c843bedd56d

Observation 37a2c3dc-17c1-47f4-ac61-a21a0412bb15 · outbound

This paper cites an unresolved cited work.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-11T17:10:52.802974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T17:10:22.151056Z digest=sha256:24ce3722e8b0af10a32241c66b7de19ac3d623b2a709977d3a6343fc91eda493

Observation 86159046-675d-4300-944e-ae7f004f9195 · outbound

This paper cites an unresolved cited work.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-11T17:10:52.786731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T17:10:22.156056Z digest=sha256:a23a24d518cdb45fc13d41cd1782cd49d6595282bdad5ac3adebd6355aa73ffa

Observation 4b89d5ee-6326-4aa5-ad58-fce3b1848ceb · outbound

This paper cites an unresolved cited work.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-11T17:10:52.767594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T17:10:22.160851Z digest=sha256:5c7b2a6791d2841511238d3815f1ce9bc353adb1656619b5b92821670b670952

Observation 09805f3b-88cf-4c9d-ad31-12813f0b514a · outbound

This paper cites an unresolved cited work.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-11T17:10:52.753427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T17:10:22.167400Z digest=sha256:5629b9ea803b6f51c7b15b6293afc8f5930f2e3cf5ca3a3d5f2a6336f551414d

Observation 68c54bb7-34e2-47d1-9e7b-eb28801a1685 · outbound

This paper cites an unresolved cited work.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-11T17:10:52.738613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T17:10:22.172740Z digest=sha256:ecad88e10db603ff98276d1057708adbbd5a84a95ff352b50e7dc8e24d30b30a

Observation 2733460a-5410-4e53-a018-d61641052d41 · outbound

This paper cites an unresolved cited work.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-11T17:10:52.723950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T17:10:22.177978Z digest=sha256:38160884d3be95227740b5032bd7d338e1e69e41a96375014da7b7d69734f813

Observation 97ad5e53-6765-42f1-98b7-882556d71f13 · outbound

This paper cites an unresolved cited work.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-11T17:10:52.709918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T17:10:22.183120Z digest=sha256:a41ab9c68ee247e14f1a821d5e8ca03f906abdaa43260b1f8e0e5490e45058e8

Observation ea38dc8d-3864-4df4-9681-ab08e95dbe5e · outbound

This paper cites Global Description.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Global Description

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:10:52.695354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T17:10:22.188382Z digest=sha256:3d648def39d8c19bbb7c57fb6276d9eb953757689cc76b3f09de1f0165730064

Observation 6b3ff845-bbdb-4d94-a251-0079f1e3e7a1 · outbound

This paper cites 3) Ex- trinsic Hallucination: Evaluate whether the text introduces content that is not present in the video.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption 3) Ex- trinsic Hallucination: Evaluate whether the text introduces content that is not present in the video

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:10:52.680131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T17:10:22.193343Z digest=sha256:30869e614e917bfe04170745132ba93703278acdaacd1bb74ecfe9a15809127f

Observation c46daef7-2472-4254-905a-553e26f6af1c · outbound

This paper cites counter-intuitive.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption counter-intuitive

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:10:52.665365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T17:10:22.198272Z digest=sha256:f146deff414f4efd28b7938c3854d35cc59fea885feba287f193697315a3e77f

Observation 02fd437d-b1f0-4537-bef6-fd4705d237b6 · outbound

This paper cites ,".join([f.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption ,".join([f

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:10:52.648749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T17:10:22.202993Z digest=sha256:259a0e6fe40121a90f769885d8f71fcb95ad5d81a5f31f48a634adf8c9762d65

Observation f2da4476-454b-4771-bfa6-fb07cec18eee · outbound

This paper cites subject" +.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption subject" +

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:10:52.631507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T17:10:22.208917Z digest=sha256:1b03667f3f6c4a55ff85a6d01a72f2273219bda96d42a16b29be85729426074d

Observation 90cdd62c-75c5-439f-9dd4-c3f4f6cc0e8f · outbound

This paper cites The video shows.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption The video shows

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:10:52.614918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T17:10:22.213643Z digest=sha256:936363972e212bdcb012191a39d551e107edf104fd194259a563566044c6505f

Observation b486091b-22ac-4c79-92e5-72752dcde0a6 · outbound

This paper cites A man... and a woman..., and a man.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption A man... and a woman..., and a man

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:10:52.596761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T17:10:22.218481Z digest=sha256:6038baf76ba84103bab177f4fa8138d291ad28bf252609413cf388f0bcf5706e

Observation e2d51b3a-c5fb-4e9b-ba52-394fe430e6eb · outbound

This paper cites The scene is.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption The scene is

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:10:52.581473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T17:10:22.223004Z digest=sha256:466b481c50241ac36b2cd11d0d3258d4b8be020fe49a1066dd666300fbc87f9d

Observation a1f40f2a-eb5d-465d-ae52-8fc043bcb8c4 · outbound

This paper cites Aligning prompt used during alignment with the open source model.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Aligning prompt used during alignment with the open source model

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:10:52.564812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T17:10:22.228795Z digest=sha256:93f3eac45da8cc9ba8e3933bf93986829c81f496d9e3ef598188059a605a7f26

Pith citing papers

No inbound Pith citation observations are available.