Pith. sign in

Paper Citation Record · LEDGER

Can video generation replace cinematographers? Research on the cinematic language of generated video

As of 21 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 3 inbound Pith citation observations for arXiv:2412.12223.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.12223 v2

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T14:52:04.456970Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:31:56.426125Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-12T14:06:17.109079Z

Reference resolution

46 of 46 outbound references displayed

  • verified exact0
  • verified fuzzy20
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 375f553b-a259-44bb-9d47-037546a7f72b · outbound

This paper cites Genetic Algorithm: Reviews, Implementations, and Applications.

Can video generation replace cinematographers? Research on the cinematic language of generated video Genetic Algorithm: Reviews, Implementations, and Applications

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T14:52:04.268282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:52:04.268282Z digest=sha256:63a28c8a7509c93d605618a6ff27ba48c504f9e7bf12307178e200c1607612ce

Observation 959e2757-a431-47c2-9afb-7fc8a9a34b6b · outbound

This paper cites Frozen in time: A joint video and image encoder for end-to-end retrieval.

Can video generation replace cinematographers? Research on the cinematic language of generated video Frozen in time: A joint video and image encoder for end-to-end retrieval

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T14:52:04.272775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:52:04.272775Z digest=sha256:99f90e6ce6c656d0219bd85ef54c35ab7c836288aab34fba00e0cbba10515ea4

Observation 14bbcf77-fa86-45ce-a1b5-f9f38415fdbc · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

Can video generation replace cinematographers? Research on the cinematic language of generated video Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T14:52:04.276344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:52:04.276344Z digest=sha256:594eaee5ceefec495904f26f3aff91a1ecac4db7b65ece894d68f28ddebb49d4

Observation 31b30414-febf-45de-9ade-323fb3e8f4e4 · outbound

This paper cites Align your latents: High-resolution video synthesis with la- tent diffusion models.

Can video generation replace cinematographers? Research on the cinematic language of generated video Align your latents: High-resolution video synthesis with la- tent diffusion models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T14:52:04.280401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:52:04.280401Z digest=sha256:5206ef6b6aca31d9888727936f692656294614d5d0baec657ca6ebe92c860cdc

Observation d2e5ba77-fe6e-4cb0-90ad-49704a170433 · outbound

This paper cites Collecting highly paral- lel data for paraphrase evaluation.

Can video generation replace cinematographers? Research on the cinematic language of generated video Collecting highly paral- lel data for paraphrase evaluation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T14:52:04.284320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:52:04.284320Z digest=sha256:11708e3909abdfdb05d8e79c86d72bed458868351bd3e48cd032d57ecd5cbe2a

Observation 0ee5c73a-211b-4372-8d96-788bca3dddde · outbound

This paper cites VideoCrafter1: Open Diffusion Models for High-Quality Video Generation.

Can video generation replace cinematographers? Research on the cinematic language of generated video VideoCrafter1: Open Diffusion Models for High-Quality Video Generation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T14:52:04.288493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:52:04.288493Z digest=sha256:8c5f459270f727d3a7d543079f1a5e02adf43e77502f87e0ef1ed0dcf725e92c

Observation 17a07dd3-0067-4586-a54f-f922941b0114 · outbound

This paper cites ShareGPT4Video: Improving Video Understanding and Generation with Better Captions.

Can video generation replace cinematographers? Research on the cinematic language of generated video ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T14:52:04.292949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:52:04.292949Z digest=sha256:5fb60ad5bd67eac28cb986147a84b0baa5b02f8106aa52410f5322c17c7f1341

Observation ef0ff874-56c8-4b24-afb8-8a6bb97b263f · outbound

This paper cites Vqgan-clip: Open domain image generation and editing with natural language guidance.

Can video generation replace cinematographers? Research on the cinematic language of generated video Vqgan-clip: Open domain image generation and editing with natural language guidance

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:52:04.984997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:52:04.297145Z digest=sha256:e7f139bcbf3f1163eed71aed271cd82a45c951fc5dbaee3b745e41965a01914c

Observation 63aa037b-4a69-4310-b07b-29a87e1f2656 · outbound

This paper cites Clipdraw: Exploring text-to-drawing synthesis through language-image encoders.

Can video generation replace cinematographers? Research on the cinematic language of generated video Clipdraw: Exploring text-to-drawing synthesis through language-image encoders

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:52:04.974673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:52:04.301240Z digest=sha256:e45c4b2b3f2f32b5fd4878672162ae89efb52de20028005a4fec0988c5a5a335

Observation aaeb91af-d23c-4643-8bfe-7252def77f5d · outbound

This paper cites Genetic algorithms in search, optimiza- tion, and machine learning.

Can video generation replace cinematographers? Research on the cinematic language of generated video Genetic algorithms in search, optimiza- tion, and machine learning

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:52:04.964337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:52:04.304835Z digest=sha256:aeb3cd5ede38a65d049f6455f415df2b5b3307b56aa4edad7ae5c62d0e3222c9

Observation 8941ef0b-cdae-4c48-abe0-e0e0fa26874b · outbound

This paper cites Open-vocabulary object detection via vision and language knowledge distillation.

Can video generation replace cinematographers? Research on the cinematic language of generated video Open-vocabulary object detection via vision and language knowledge distillation

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:52:04.952005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:52:04.308891Z digest=sha256:d29ae0b2dbe4f98c37bf9bb2f48f6cfa4a2fde935b0748eccc64af3e8cbd9ea8

Observation 75bfbd1d-0555-409e-8621-1a54827bcb45 · outbound

This paper cites Animatediff: Animate your personalized text- to-image diffusion models without specific tuning.

Can video generation replace cinematographers? Research on the cinematic language of generated video Animatediff: Animate your personalized text- to-image diffusion models without specific tuning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T14:52:04.312360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:52:04.312360Z digest=sha256:bac6397acccf9117aeda3fec58fd9c811e134d92f47499567c63425fa8d1a4b2

Observation aae3ae33-3bff-4ec0-b6d5-c3563893a060 · outbound

This paper cites CameraCtrl: Enabling Camera Control for Text-to-Video Generation.

Can video generation replace cinematographers? Research on the cinematic language of generated video CameraCtrl: Enabling Camera Control for Text-to-Video Generation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T14:52:04.315759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:52:04.315759Z digest=sha256:779c41eb881d7441937512d98a8dadc175e7b1a67ed5a4b775a77b0225800107

Observation 129abff1-a78b-4dae-82b8-d3dc4df0ee76 · outbound

This paper cites Denoising dif- fusion probabilistic models.

Can video generation replace cinematographers? Research on the cinematic language of generated video Denoising dif- fusion probabilistic models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T14:52:04.319886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:52:04.319886Z digest=sha256:05229e3949a34b4210305f700b0e498e986bd238ac43955f117b409e3704c652

Observation 8e3aa7a1-95e4-4cc7-9e79-3f6b98bb1e5e · outbound

This paper cites Imagen Video: High Definition Video Generation with Diffusion Models.

Can video generation replace cinematographers? Research on the cinematic language of generated video Imagen Video: High Definition Video Generation with Diffusion Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T14:52:04.323354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:52:04.323354Z digest=sha256:826c904d8b698f569f9a09e7e9a44be73381828915819a29bb011887d570685c

Observation f9d1b718-66db-4cd5-89e9-023d8afd0947 · outbound

This paper cites Lora: Low- rank adaptation of large language models.

Can video generation replace cinematographers? Research on the cinematic language of generated video Lora: Low- rank adaptation of large language models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T14:52:04.327298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:52:04.327298Z digest=sha256:34a730419d06b8a8ecc8e228da54bfd6e23eb718cb79601401dc876c88fc75df

Observation 3aee8cf5-323d-4cc5-9052-8489189d668b · outbound

This paper cites Language-driven semantic seg- mentation.

Can video generation replace cinematographers? Research on the cinematic language of generated video Language-driven semantic seg- mentation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T14:52:04.330518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:52:04.330518Z digest=sha256:b49cf178df283a1e2cfdd86852ae7d6e01ba6ba8957dfc4abb43cb5b1ca66fdc

Observation 87c2dbb2-1dd9-42d1-8725-6819dd82e18a · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Can video generation replace cinematographers? Research on the cinematic language of generated video LLaVA-OneVision: Easy Visual Task Transfer

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T14:52:04.333805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:52:04.333805Z digest=sha256:4e9340007d43bf5a05c3389208063429adb759198ec2ee802ead1752917f0f18

Observation 8d1806eb-b47b-4d55-a369-008f5faa7af8 · outbound

This paper cites Grounded language-image pre-training.

Can video generation replace cinematographers? Research on the cinematic language of generated video Grounded language-image pre-training

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T14:52:04.337506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:52:04.337506Z digest=sha256:c641133a02a031e21deb8c3185585d8fcd5658f9150b4cf1205795e0184885c5

Observation 7a93eeef-ab2e-478f-8457-7d093b3a809e · outbound

This paper cites Clip4clip: An empirical study of clip for end to end video clip retrieval and captioning.

Can video generation replace cinematographers? Research on the cinematic language of generated video Clip4clip: An empirical study of clip for end to end video clip retrieval and captioning

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:52:04.904761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:52:04.340939Z digest=sha256:ada4f7536f8f145fdd61b2ffe2176ff8fa77d6605e66a609aa64ce1cdb88374d

Observation 527fd4cb-1cc0-44a2-977f-e261c85129d4 · outbound

This paper cites OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation.

Can video generation replace cinematographers? Research on the cinematic language of generated video OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T14:52:04.344269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:52:04.344269Z digest=sha256:82f5f7c8d09d89f3d013ed1ca3d3156ca37eb6bfba986d72bae568ddb4609019

Observation 9a9d807c-3634-4dc7-9029-fafc9cc4cea8 · outbound

This paper cites Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever.

Can video generation replace cinematographers? Research on the cinematic language of generated video Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:52:04.893730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:52:04.348319Z digest=sha256:01fe13e0674e9bd413f13500d9909f0184311388c0b518da795363b43f3fd0f6

Observation f9bc02aa-9d28-466d-8401-f29774ebfc17 · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

Can video generation replace cinematographers? Research on the cinematic language of generated video Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T14:52:04.351789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:52:04.351789Z digest=sha256:dd167e96febf3be7f76e8d4e41c839a6ab8fb792585f9bf3739a31174dc7b74a

Observation 105c4f95-1fa1-4882-8c20-078cf3e041df · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Can video generation replace cinematographers? Research on the cinematic language of generated video High-resolution image synthesis with latent diffusion models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T14:52:04.356725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:52:04.356725Z digest=sha256:719c524b22a52747e2e25a6c319ad1f47dc694f8f9b3e40cb4b09cc65abbc9b7

Observation 85e6d462-c248-4215-be66-d2d105201021 · outbound

This paper cites an unresolved cited work.

Can video generation replace cinematographers? Research on the cinematic language of generated video Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:52:04.876175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:52:04.360114Z digest=sha256:5d05b7e3ed81c6a84d68041085588f7a95392e65010c727dffef0cafc0e2ceda

Observation 0b3caa55-963c-4c50-9421-8a316f6a87d9 · outbound

This paper cites Photorealistic text-to-image diffusion models with deep language understanding.

Can video generation replace cinematographers? Research on the cinematic language of generated video Photorealistic text-to-image diffusion models with deep language understanding

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:52:04.865528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:52:04.363428Z digest=sha256:4c485b097a4d08a57cfb7e45cdc9d8403a432d614d85a8e1a90abd580cbe2edb

Observation 827a4352-3bb0-4b90-a251-fcb1bc7381bb · outbound

This paper cites Tempo- ral generative adversarial nets with singular value clipping.

Can video generation replace cinematographers? Research on the cinematic language of generated video Tempo- ral generative adversarial nets with singular value clipping

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T14:52:04.367087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:52:04.367087Z digest=sha256:951d3a53ca9f02872ec11a0beaa3097ea11cf73627d8b5f0eff4bb78fb44f753

Observation 28e1a3b2-0c8f-4f15-a32c-277ba58c38b6 · outbound

This paper cites Cinescale: A dataset of cinematic shot scale in movies.

Can video generation replace cinematographers? Research on the cinematic language of generated video Cinescale: A dataset of cinematic shot scale in movies

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:52:04.845584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:52:04.371117Z digest=sha256:e738e7f68b3196d98c05cedba74474856049b5e2583d65843365a64a7ba0742b

Observation 3720e106-351c-43bd-ac50-d979466ff682 · outbound

This paper cites Cinescale2: a dataset of cinematic camera features in movies.

Can video generation replace cinematographers? Research on the cinematic language of generated video Cinescale2: a dataset of cinematic camera features in movies

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:52:04.833584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:52:04.375215Z digest=sha256:89dc5bd9e098c6fc4a09b0a8f0ea904ffda439ad83273040d7664e181ec1ab0c

Observation a544eaf5-1c54-4c3c-bee1-4a301d0081d7 · outbound

This paper cites Laion-5b: An open large-scale dataset for training next generation image-text models.

Can video generation replace cinematographers? Research on the cinematic language of generated video Laion-5b: An open large-scale dataset for training next generation image-text models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T14:52:04.379548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:52:04.379548Z digest=sha256:64a1e2beb47e119f56ae3e8f40428d80001553c3224169fcf47fcad8014565e2

Observation bf293bf8-f741-4037-a771-6b0713b2c3f3 · outbound

This paper cites Stylegan-v: A continuous video generator with the price, image quality and perks of stylegan2.

Can video generation replace cinematographers? Research on the cinematic language of generated video Stylegan-v: A continuous video generator with the price, image quality and perks of stylegan2

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:52:04.814403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:52:04.384245Z digest=sha256:60d5255261d78d57f24e97ca4e4d585b4f463852b4a619e11a6ca1d7caef3657

Observation 8edb603d-e2b5-49cf-91c7-a9c25da554f5 · outbound

This paper cites Mocogan: Decomposing motion and content for video generation.

Can video generation replace cinematographers? Research on the cinematic language of generated video Mocogan: Decomposing motion and content for video generation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T14:52:04.388280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:52:04.388280Z digest=sha256:02a403fd2c88336bed7ecd2c158e28e85f26c82be5411c843e813283bab217b0

Observation d4614245-0881-4484-a1f4-50a4edcc97e0 · outbound

This paper cites Towards Accurate Generative Models of Video: A New Metric & Challenges.

Can video generation replace cinematographers? Research on the cinematic language of generated video Towards Accurate Generative Models of Video: A New Metric & Challenges

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T14:52:04.392788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:52:04.392788Z digest=sha256:3be1d50a137e28acdb9c39110880f887115bd0946a13e696dc7215ee7a28fd39

Observation 221e54e0-23b1-4f57-ace5-bdb43787a163 · outbound

This paper cites Clipasso: Semantically-aware object sketching.

Can video generation replace cinematographers? Research on the cinematic language of generated video Clipasso: Semantically-aware object sketching

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T14:52:04.399561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:52:04.399561Z digest=sha256:e270078d3738e3cf21b9b235e3f16b2ce33a13bbc284c897ea13491daebe702c

Observation f9673331-481e-4a84-93d7-c12370abcc61 · outbound

This paper cites Generating videos with scene dynamics.

Can video generation replace cinematographers? Research on the cinematic language of generated video Generating videos with scene dynamics

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:52:04.786848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:52:04.403491Z digest=sha256:7c18b7de1979fa216f2b849f3ddf1bf5d3678de9850159e5d9f4abb4efcb2525

Observation 01eae950-b823-4b0a-8309-939db9388289 · outbound

This paper cites Cocaclip: Exploring distillation of fully- connected knowledge interaction graph for lightweight text- image retrieval.

Can video generation replace cinematographers? Research on the cinematic language of generated video Cocaclip: Exploring distillation of fully- connected knowledge interaction graph for lightweight text- image retrieval

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:52:04.774618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:52:04.410138Z digest=sha256:1f8249dc4a0e67aea453d7c467d3fdc30ab148768c078fbe8b2abbf83104d196

Observation f1e1a69a-c6a8-406d-a4a1-9b528f994fa6 · outbound

This paper cites Videoclip-xl: Advancing long descrip- tion understanding for video clip models.

Can video generation replace cinematographers? Research on the cinematic language of generated video Videoclip-xl: Advancing long descrip- tion understanding for video clip models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:52:04.761169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:52:04.414924Z digest=sha256:3b39decd7f42926688850be3ba93a5d3c07034530d83c75839536f47c7e05671

Observation 3b7a3a6c-6e3d-4998-90ea-e40b57284cc0 · outbound

This paper cites Videocomposer: Compositional video synthesis with motion controllability.

Can video generation replace cinematographers? Research on the cinematic language of generated video Videocomposer: Compositional video synthesis with motion controllability

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:52:04.749565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:52:04.420380Z digest=sha256:034b9f8613beb1d074f5ec12a8e44efac3f3a919d7e660e92f4c9769648547b0

Observation 4d7eafa1-72e3-448a-9e89-79500ab9becc · outbound

This paper cites Internvid: A large-scale video-text dataset for multimodal understanding and generation.

Can video generation replace cinematographers? Research on the cinematic language of generated video Internvid: A large-scale video-text dataset for multimodal understanding and generation

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:52:04.739134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:52:04.424691Z digest=sha256:f5ef9a20b8fd2b7af1286eff9c8c3762986ff5d7c054d1550fcdb0f7e1576328

Observation 349c49d7-2360-4e99-91d4-3954c72df6a0 · outbound

This paper cites Motionctrl: A unified and flexible motion controller for video generation.

Can video generation replace cinematographers? Research on the cinematic language of generated video Motionctrl: A unified and flexible motion controller for video generation

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:52:04.727900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:52:04.429014Z digest=sha256:6305a0dc647d0a6a3fbaf2bdea9654f98087ae323e7d2c3abf256ab6bb3640d3

Observation 34559d55-451e-4522-b182-e07537dd5f04 · outbound

This paper cites Videoclip: Contrastive pre-training for zero-shot video-text understanding.

Can video generation replace cinematographers? Research on the cinematic language of generated video Videoclip: Contrastive pre-training for zero-shot video-text understanding

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:52:04.715605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:52:04.433727Z digest=sha256:37be52c840a2aaab01c52d82b14bba0e7c949891031886c23cb42f21964cf5ab

Observation 040774fe-f20e-4acd-a470-052d8ef5a949 · outbound

This paper cites Groupvit: Semantic segmentation emerges from text supervision.

Can video generation replace cinematographers? Research on the cinematic language of generated video Groupvit: Semantic segmentation emerges from text supervision

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T14:52:04.438399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:52:04.438399Z digest=sha256:2f91da141bf8fffabe82afe92ddbb23edaecc9ec5dafb31388b5ff9028c1308f

Observation 0ff325ef-4319-4dd7-87a8-83048304ceef · outbound

This paper cites Direct-a-video: Customized video generation with user- directed camera movement and object motion.

Can video generation replace cinematographers? Research on the cinematic language of generated video Direct-a-video: Customized video generation with user- directed camera movement and object motion

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:52:04.695723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:52:04.443170Z digest=sha256:8146043bc51f4bce8d4f32dd87c60c20da5776d43d805e59aeaa43ed1bf3c32f

Observation 30bfe89b-fcac-463a-a20e-59773296072e · outbound

This paper cites Long-clip: Unlocking the long-text capability of clip.

Can video generation replace cinematographers? Research on the cinematic language of generated video Long-clip: Unlocking the long-text capability of clip

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:52:04.682926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:52:04.448122Z digest=sha256:8621cde949a94001dcca18ed9a7a5e0fa80199d204ef120de19f9ef7db802e8c

Observation 38b99fee-15b0-499a-99bc-18db78878446 · outbound

This paper cites Multi-LoRA Composition for Image Generation.

Can video generation replace cinematographers? Research on the cinematic language of generated video Multi-LoRA Composition for Image Generation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T14:52:04.452484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:52:04.452484Z digest=sha256:c49945a200dc6d8174088410a0b5fbc6a87fd52187b5391179328e639d5c9175

Observation 4546e254-d153-4496-8f75-4e064a1458de · outbound

This paper cites Stereo magnification: learning view synthesis using multiplane images.

Can video generation replace cinematographers? Research on the cinematic language of generated video Stereo magnification: learning view synthesis using multiplane images

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:52:04.670235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:52:04.456970Z digest=sha256:1abb574092522da00aa2564f90e91bb12864f667981e4e5b068987940a2523a1

Pith citing papers

Observation da734167-211c-4d1c-a33b-8ed90ab449ab · inbound

Towards Understanding Camera Motions in Any Video cites this paper.

Towards Understanding Camera Motions in Any Video Can video generation replace cinematographers? Research on the cinematic language of generated video

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-16T11:31:56.426125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:31:56.426125Z digest=sha256:b05373afc4d7bec46444e0dafe1d956ae6a796ee8200176d1ec2f19069464afa

Observation 57df0b89-4c9c-4c6a-8a6a-1aac18ca925e · inbound

Natural Language Camera Movement Understanding cites this paper.

Natural Language Camera Movement Understanding Can video generation replace cinematographers? Research on the cinematic language of generated video

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-12T05:16:19.750976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T05:16:19.750976Z digest=sha256:4ffcaf04cc0808ce32cb9c7faf11a22f0b9a5d637871a14e6099144edae9eb9b

Observation ef6adfc5-8d40-421e-8d1c-141d5467603e · inbound

Temporally Grounded Compositional Camera Motion Understanding via Geometric Knowledge Distillation cites this paper.

Temporally Grounded Compositional Camera Motion Understanding via Geometric Knowledge Distillation Can video generation replace cinematographers? Research on the cinematic language of generated video

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-12T14:06:17.112505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T14:06:16.740457Z digest=sha256:70fe6c216ba876f9ce255c983900ad99b65456873e3a027d8a06257854ddf946