Pith. sign in

Paper Citation Record · LEDGER

From Image to Video: An Empirical Study of Diffusion Representations

As of 20 August 2026, this Paper Citation Record lists 69 of 69 outbound references and 2 inbound Pith citation observations for arXiv:2502.07001.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.07001 v2

Coverage vector

measured 69 of 69 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T14:10:55.455485Z

measured 71 of 71 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-19T03:42:54.620069Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-19T03:42:57.302590Z

Reference resolution

69 of 69 outbound references displayed

  • verified exact1
  • verified fuzzy42
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 603821a0-9355-48c5-9aeb-a0481d66b0c7 · outbound

This paper cites Self-supervised learning from images with a joint-embedding predictive architecture.

From Image to Video: An Empirical Study of Diffusion Representations Self-supervised learning from images with a joint-embedding predictive architecture

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:56.543774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:10:55.116949Z digest=sha256:b88b4a4ae04eff8f648dcdd2555120705352060e5f45a834f3601d0934d7f3ab

Observation cf7f64f8-25c1-4c34-910c-da56a3c62b98 · outbound

This paper cites Video diffusion models learn the struc- ture of the dynamic world.

From Image to Video: An Empirical Study of Diffusion Representations Video diffusion models learn the struc- ture of the dynamic world

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:56.529087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:10:55.122254Z digest=sha256:4565ed5e12897aa8f79eec1ef1056ce5c70f8099370df7f9ff5dd7770547ceb1

Observation 84998fb5-9c41-4862-ad36-25ae68f7595f · outbound

This paper cites Learning by Reconstruction Produces Uninformative Features For Perception.

From Image to Video: An Empirical Study of Diffusion Representations Learning by Reconstruction Produces Uninformative Features For Perception

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T14:10:55.127624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:10:55.127624Z digest=sha256:5ba2a5fd704aeec93706385d242470ed8927490c2353412d2cc6a9fd230221d8

Observation e8f2fc10-1067-4eb7-b442-9c176f7863b6 · outbound

This paper cites Label-efficient se- mantic segmentation with diffusion models.

From Image to Video: An Empirical Study of Diffusion Representations Label-efficient se- mantic segmentation with diffusion models

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:56.513716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:10:55.133861Z digest=sha256:9aa3b97ce93ba67922ee3e8ac78f200cd0dba0201a17a4cb487050fb1a31f02e

Observation a21fbcaf-e472-4e30-a2ab-a570cd1ceb2a · outbound

This paper cites Revisiting Feature Prediction for Learning Visual Representations from Video.

From Image to Video: An Empirical Study of Diffusion Representations Revisiting Feature Prediction for Learning Visual Representations from Video

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T14:10:55.138930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:10:55.138930Z digest=sha256:7ed187ec264d832bae40983940173054a3ea53349a60adbc18f5af81a34be50e

Observation 65321793-a5fe-4987-8f30-8828dc9fc8d7 · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

From Image to Video: An Empirical Study of Diffusion Representations Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T14:10:55.144179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:10:55.144179Z digest=sha256:157249591f00eb08b339354f37da444be2b62592d99e8d3c2cb4634c0b54d8ac

Observation ae5814d3-b1c7-4258-a241-1b50a0120411 · outbound

This paper cites Deep regression on manifolds: a 3D rota- tion case study.

From Image to Video: An Empirical Study of Diffusion Representations Deep regression on manifolds: a 3D rota- tion case study

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:56.499007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:10:55.149974Z digest=sha256:cc0c535b38b15a50fb875639b33ecd02f9c54b97bfc5a7b3132b2ca1df6ab81e

Observation 2b88fef2-7f80-42e0-b5a4-fc00d09576d3 · outbound

This paper cites Emerg- ing properties in self-supervised vision transformers.

From Image to Video: An Empirical Study of Diffusion Representations Emerg- ing properties in self-supervised vision transformers

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:56.482929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:10:55.154906Z digest=sha256:0a0f9587cf401fc36419bc9d5fbe44cce0299b121a8fac59fe64647ee264167e

Observation 2c952667-f4b4-437d-946e-a2045f6b52b9 · outbound

This paper cites Quo vadis, action recognition? a new model and the kinetics dataset.

From Image to Video: An Empirical Study of Diffusion Representations Quo vadis, action recognition? a new model and the kinetics dataset

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T14:10:55.159745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:10:55.159745Z digest=sha256:9ddd13c34c02d8843134ddd9e59ec42f9731021efc78b08de9cf2dc9f9f3aad6

Observation ddf9d714-8dfe-4364-b1e4-2a330064446e · outbound

This paper cites A Short Note on the Kinetics-700 Human Action Dataset.

From Image to Video: An Empirical Study of Diffusion Representations A Short Note on the Kinetics-700 Human Action Dataset

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T14:10:55.165909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:10:55.165909Z digest=sha256:2c150842f55592d1b79d9ede28de06ba78cce9cae13c2277e5d368dabbf41fc5

Observation 647ea506-fb21-412f-8709-191a86556497 · outbound

This paper cites Scaling 4D Representations.

From Image to Video: An Empirical Study of Diffusion Representations Scaling 4D Representations

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T14:10:55.171296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:10:55.171296Z digest=sha256:67884906cba4c8b8be87d2606c641c72cab7fdba65255316cfa64a23e30d5a45

Observation fd97b2ea-81ad-4ab1-84fb-bf8e4fc288ed · outbound

This paper cites an unresolved cited work.

From Image to Video: An Empirical Study of Diffusion Representations Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:10:56.454616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:10:55.176297Z digest=sha256:357c16b71efc47820c9238360c7fbe8be48cb243cef2e693eb7a904fd22a0c51

Observation 62146370-3bbb-4834-85ed-c1be60865c89 · outbound

This paper cites PaLI-3 Vision Language Models: Smaller, Faster, Stronger.

From Image to Video: An Empirical Study of Diffusion Representations PaLI-3 Vision Language Models: Smaller, Faster, Stronger

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T14:10:55.182071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:10:55.182071Z digest=sha256:380572e7151686153c1c61f6d9e7e94a9f8dbc0b800e8937fe0cdb0566fe413d

Observation c7fff663-46f1-4cee-9c79-59e461514fc6 · outbound

This paper cites Text-to-image diffusion mod- els are zero shot classifiers.

From Image to Video: An Empirical Study of Diffusion Representations Text-to-image diffusion mod- els are zero shot classifiers

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:56.437657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:10:55.187263Z digest=sha256:de17ecc030ab63f4f2c59a3e5022840ca9f613e4e2f105855236666db54babb6

Observation 6babfe3c-9302-4e56-9aee-1c71d713bea9 · outbound

This paper cites Chang, Manolis Savva, Maciej Hal- ber, Thomas Funkhouser, and Matthias Nießner.

From Image to Video: An Empirical Study of Diffusion Representations Chang, Manolis Savva, Maciej Hal- ber, Thomas Funkhouser, and Matthias Nießner

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:56.421615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:10:55.191848Z digest=sha256:3d9720d0f65de5dc84089bf8255b7218b24ecbcbdd87d56a13177dc199440f24

Observation ceb7d32a-da1d-453d-bd88-0f0e747d5aa5 · outbound

This paper cites Depth map prediction from a single image using a multi-scale deep net- work.

From Image to Video: An Empirical Study of Diffusion Representations Depth map prediction from a single image using a multi-scale deep net- work

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:56.404034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:10:55.196339Z digest=sha256:435d36ae9e7677f24b0d942dbf646b6a61c47beda8e7e7a34cc6d8a626fd51aa

Observation a71f2f6c-c156-4b51-a343-bf48d93d0515 · outbound

This paper cites Taming transformers for high-resolution image synthesis.

From Image to Video: An Empirical Study of Diffusion Representations Taming transformers for high-resolution image synthesis

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T14:10:55.200902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:10:55.200902Z digest=sha256:6aa9840769fb5627ef9ded5a1717e477e30a81564223288c2e3b0cb4836a1d71

Observation 48a5b6a2-1bea-47c0-9fda-cc34bcdc70fc · outbound

This paper cites Masked autoencoders as spatiotemporal learners.

From Image to Video: An Empirical Study of Diffusion Representations Masked autoencoders as spatiotemporal learners

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:56.374861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:10:55.206228Z digest=sha256:300fdc8c7e40012e9ffb6ef8f42fd80ede28a81b8e7b5d837db40834feb8b45b

Observation ade7fd2f-118a-4d8b-a295-b245f66e7678 · outbound

This paper cites Diffusion Models and Representation Learning: A Survey.

From Image to Video: An Empirical Study of Diffusion Representations Diffusion Models and Representation Learning: A Survey

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T14:10:55.211931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:10:55.211931Z digest=sha256:4426727180aabc01946be8a34cf6995590196e567c0cbc87318a21078c92c349

Observation 16a378de-c69e-41e8-8dcc-8909361ab035 · outbound

This paper cites Something Something.

From Image to Video: An Empirical Study of Diffusion Representations Something Something

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:56.354536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:10:55.217130Z digest=sha256:20701e02f452a1e9938364a4e0accf40fe09c8aa278c90c2ca726d381a57ed74

Observation 20c8d60a-5b6a-427b-b8fd-cfcf907d1c47 · outbound

This paper cites Kubric: A scalable dataset generator.

From Image to Video: An Empirical Study of Diffusion Representations Kubric: A scalable dataset generator

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:56.338487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:10:55.222921Z digest=sha256:fb34570d68cb5a7d6ad41b73f02473bf218c8fa7feff25c47ff2cacdd6d4436a

Observation 8d92fef7-1769-4742-b1e0-aeed694afa2b · outbound

This paper cites Photorealistic video generation with diffusion models.

From Image to Video: An Empirical Study of Diffusion Representations Photorealistic video generation with diffusion models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:56.321232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:10:55.227765Z digest=sha256:961955a20138fc4b122568d9e41b8dca6fa5162964abb2901adda83950a95638

Observation 7302ca0f-4509-4216-a41e-7b1fd9d63e17 · outbound

This paper cites Masked autoencoders are scalable vision learners.

From Image to Video: An Empirical Study of Diffusion Representations Masked autoencoders are scalable vision learners

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:56.303949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:10:55.232445Z digest=sha256:5ad85bee2f03323bca47c922b7af5ea2f5c07baa083b63a3f8bb25283fa2dd2f

Observation 737ae461-0659-4691-a059-b1c40a504c27 · outbound

This paper cites Unsupervised keypoints from pretrained diffusion models.

From Image to Video: An Empirical Study of Diffusion Representations Unsupervised keypoints from pretrained diffusion models

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:56.288117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:10:55.237362Z digest=sha256:f83ba46754fd16060a18f8e2451c43f8fc04b8f53b163c065b69aa4d4f620b2b

Observation a2be131a-35d6-40b5-873b-dd347fdbb184 · outbound

This paper cites Denoising diffu- sion probabilistic models.

From Image to Video: An Empirical Study of Diffusion Representations Denoising diffu- sion probabilistic models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T14:10:55.242828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:10:55.242828Z digest=sha256:f28cd7501dad0f48a83c8545cb989780ae226e988a13edc444efdd90601da97f

Observation a32fe247-85c8-4dc7-8b9e-4ab376ec1002 · outbound

This paper cites DepthCrafter: Generating Consistent Long Depth Sequences for Open-world Videos.

From Image to Video: An Empirical Study of Diffusion Representations DepthCrafter: Generating Consistent Long Depth Sequences for Open-world Videos

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T14:10:55.247542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:10:55.247542Z digest=sha256:996ef68e1059a4b2c2589931e2301c8dd35ffa2f03853fa3e36805a8df1359a6

Observation 68812302-5910-4492-ada0-0d5cf32723e1 · outbound

This paper cites Elucidating the design space of diffusion-based generative models.

From Image to Video: An Empirical Study of Diffusion Representations Elucidating the design space of diffusion-based generative models

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:56.261203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:10:55.252875Z digest=sha256:6dfe127dea316a56cd8316aee79e946b4a8cf1e1e5028b2b7e8b36f6a59bfe11

Observation 87a222f0-25fc-41b2-bad7-3d32451b08f8 · outbound

This paper cites Your diffusion model is secretly a zero-shot classifier.

From Image to Video: An Empirical Study of Diffusion Representations Your diffusion model is secretly a zero-shot classifier

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:56.243722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:10:55.257709Z digest=sha256:e27efd55cf7bbdbaad2d34f96c647bf37ffcb331f937d9a3b4c2a417a38b97d8

Observation c3a21647-dd1e-4432-8077-36b0cb0d8030 · outbound

This paper cites Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models.

From Image to Video: An Empirical Study of Diffusion Representations Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T14:10:55.263104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:10:55.263104Z digest=sha256:9db58b3bbe14dc7d55e430ead70f7562ba4d739a42e1e29fc01ed71028e01d66

Observation fe5d3173-800e-4721-9961-1d22c15a3088 · outbound

This paper cites Diffusion hyperfeatures: Search- ing through time and space for semantic correspondence.

From Image to Video: An Empirical Study of Diffusion Representations Diffusion hyperfeatures: Search- ing through time and space for semantic correspondence

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:56.226488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:10:55.268287Z digest=sha256:166a77836127dd70581d5468cd53ed0f9703c73a7740969eda142765cf81f678

Observation db54a12f-71f6-4882-b412-13304013a4d5 · outbound

This paper cites Understanding deep image representations by inverting them.

From Image to Video: An Empirical Study of Diffusion Representations Understanding deep image representations by inverting them

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:56.211266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:10:55.272944Z digest=sha256:487dd474476edda12295c628fcaa4fc30889e14a9885b97bc2bb9d8751271078

Observation f6c3539d-ca5e-4a6c-b193-e525ed96fc62 · outbound

This paper cites Lexicon3d: Probing vi- sual foundation models for complex 3d scene understanding.

From Image to Video: An Empirical Study of Diffusion Representations Lexicon3d: Probing vi- sual foundation models for complex 3d scene understanding

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:56.194835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:10:55.277612Z digest=sha256:4288fd906e8f8b3942c092b6b0b4d08c863cf5c18756ecaeede7b87f7a681de5

Observation 24f582d3-87f7-4921-a4c4-d2b6ba373ad6 · outbound

This paper cites NeRF: Representing scenes as neural radiance fields for view syn- thesis.

From Image to Video: An Empirical Study of Diffusion Representations NeRF: Representing scenes as neural radiance fields for view syn- thesis

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:56.178589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:10:55.282206Z digest=sha256:055a4d9674e169073acf31e9fdb7c44d579e50e9b693d868bea43b6e6baeae80

Observation 7cc0493e-586f-4637-b833-9ad2376161ba · outbound

This paper cites Diffusion Models Beat GANs on Image Classification.

From Image to Video: An Empirical Study of Diffusion Representations Diffusion Models Beat GANs on Image Classification

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T14:10:55.286821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:10:55.286821Z digest=sha256:8865cf2dcd9d63e74eee413405e23cde8f5ff1dcc398336c3619f2d8f05185ae

Observation 41f6959e-fb2f-4959-a742-3f5306b03908 · outbound

This paper cites DiffTAD: Temporal action detection with proposal denoising diffusion.

From Image to Video: An Empirical Study of Diffusion Representations DiffTAD: Temporal action detection with proposal denoising diffusion

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:56.162756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:10:55.291762Z digest=sha256:f9ca2b2239c16286bb78a981efab9846e96b4a4adc3668fc712bf14e079c288f

Observation 297db8f9-bf23-4d29-abec-819a1b70cac8 · outbound

This paper cites an unresolved cited work.

From Image to Video: An Empirical Study of Diffusion Representations Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-08T14:10:55.296307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:10:55.296307Z digest=sha256:c2cf9dd0b372429cff4d053af5e0e6afc0ce286f2e8d9543fc0f5a9f34e80e57

Observation 1bcfb7f4-2bf5-4b4b-98e8-0b52de14c06d · outbound

This paper cites Self-supervised video pretraining yields robust and more human-aligned visual representations.

From Image to Video: An Empirical Study of Diffusion Representations Self-supervised video pretraining yields robust and more human-aligned visual representations

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:56.136122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:10:55.301353Z digest=sha256:6bcf7a2e617aba5a5570a322e8072d6c0c143293674e254560a3d77f7034203f

Observation 08856bea-5579-4067-91d6-e9ddf695b466 · outbound

This paper cites Per- ception Test: A diagnostic benchmark for multimodal video models.

From Image to Video: An Empirical Study of Diffusion Representations Per- ception Test: A diagnostic benchmark for multimodal video models

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:56.120448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:10:55.306794Z digest=sha256:052e9900b1a12ab2dab7a6d698d439eab24ba54d734655a38e9717774ff84c11

Observation babb581f-0cb3-4ed5-a5c6-ed0d335fa9bc · outbound

This paper cites Scalable diffusion models with transformers.

From Image to Video: An Empirical Study of Diffusion Representations Scalable diffusion models with transformers

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-08T14:10:55.311336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:10:55.311336Z digest=sha256:cc35d6a8abede7a79ffdd0861a977e4f9635c068858cbdb05fa9fe66021cf661

Observation 08672b8d-6e27-4aa5-a28f-9001a1def01e · outbound

This paper cites The 2017 DAVIS Challenge on Video Object Segmentation.

From Image to Video: An Empirical Study of Diffusion Representations The 2017 DAVIS Challenge on Video Object Segmentation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-08T14:10:55.315744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:10:55.315744Z digest=sha256:6a4e9fa552e5c68225a9a9f24f7b35c50a3501d6852d9dea3e1a4207784fdf57

Observation a81a3c84-5fe0-4f51-a6bc-e6a6fb922909 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

From Image to Video: An Empirical Study of Diffusion Representations Learn- ing transferable visual models from natural language super- vision

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-08T14:10:55.320638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:10:55.320638Z digest=sha256:13d6020423128044067c6249eac9e5f9f263f51efe719f28b8877436ccde62a6

Observation a5c0f855-667d-4bca-8468-481b32e18231 · outbound

This paper cites High-resolution image syn- thesis with latent diffusion models.

From Image to Video: An Empirical Study of Diffusion Representations High-resolution image syn- thesis with latent diffusion models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-08T14:10:55.326083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:10:55.326083Z digest=sha256:64842526b0a73d2dc9930e34ac450507c92f67be519b6ca6c64bdd0e34802765

Observation 116c1a14-2760-4986-a65c-d87aa487f408 · outbound

This paper cites U- Net: Convolutional networks for biomedical image segmen- tation.

From Image to Video: An Empirical Study of Diffusion Representations U- Net: Convolutional networks for biomedical image segmen- tation

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:56.074226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:10:55.331019Z digest=sha256:df272c73808ee4c9651f14be4149dd4df7c66bed180f4a2fd807a69085200c86

Observation dbfba91e-e117-43b3-a687-098d0627e39b · outbound

This paper cites Berg, and Li Fei-Fei.

From Image to Video: An Empirical Study of Diffusion Representations Berg, and Li Fei-Fei

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:56.056491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:10:55.335679Z digest=sha256:4b3466c850de11ab9f78f1e3bbe188634d58f31904ab1ce8fa52b2da8dcebbde

Observation c6ab9844-88b4-41a7-b0c7-50e4d4d98c4c · outbound

This paper cites Scene representation transformer: Geometry-free novel view syn- thesis through set-latent scene representations.

From Image to Video: An Empirical Study of Diffusion Representations Scene representation transformer: Geometry-free novel view syn- thesis through set-latent scene representations

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:56.041323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:10:55.340453Z digest=sha256:cbeb383c64b788ea61fce3d67fb7c207b90d766ae7855a2bb55d022221523f5a

Observation 3b8c6641-290a-4a7d-8e5b-11179bae4712 · outbound

This paper cites Only time can tell: Discovering temporal data for temporal modeling.

From Image to Video: An Empirical Study of Diffusion Representations Only time can tell: Discovering temporal data for temporal modeling

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:56.025424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:10:55.345836Z digest=sha256:23d09ed83b8062984d2ce70deb3c414ad34165245c9eaf8bb9d155fb7b2b1ba7

Observation 6fb7003e-13f5-43c8-bb60-fff423a83da4 · outbound

This paper cites MonoDiffusion: Self-Supervised Monocular Depth Estimation Using Diffusion Model.

From Image to Video: An Empirical Study of Diffusion Representations MonoDiffusion: Self-Supervised Monocular Depth Estimation Using Diffusion Model

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-08-08T14:10:55.563436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:10:55.350854Z digest=sha256:e1bfd6b6ccab384187da5491ad46f3d48e9bc3a4b6d4ca42c9901f5895b4a1ca

Observation f154804d-1108-4532-90cb-1f8bf6e6e25a · outbound

This paper cites Deep unsupervised learning using nonequilibrium thermodynamics.

From Image to Video: An Empirical Study of Diffusion Representations Deep unsupervised learning using nonequilibrium thermodynamics

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-08T14:10:55.355636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:10:55.355636Z digest=sha256:6ecc492b6ae54675c05377fd662753d9d55e9261563275da5c2bd3a058bdf5d0

Observation 7b302f72-d9ab-45a3-bb7e-53a11f481eb2 · outbound

This paper cites Scalability in perception for autonomous driving: Waymo Open Dataset.

From Image to Video: An Empirical Study of Diffusion Representations Scalability in perception for autonomous driving: Waymo Open Dataset

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:55.998842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:10:55.360284Z digest=sha256:80bcc919ce93613bac2622ca62d3b1e0dd1c025a82a47151a28a251d6a57d24a

Observation 23f0ce83-05b5-4ff3-bb97-8c14bd76ba9e · outbound

This paper cites Emergent correspondence from image diffusion.

From Image to Video: An Empirical Study of Diffusion Representations Emergent correspondence from image diffusion

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:55.984019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:10:55.364835Z digest=sha256:8c359cfde66eeadb04c3ce1960d298aefc6ad3657b93073ab3fe008fc4598255

Observation ea064e57-1183-4773-92f6-55510f4c043b · outbound

This paper cites VideoMAE: Masked autoencoders are data-efficient learners for self-supervised video pre-training.

From Image to Video: An Empirical Study of Diffusion Representations VideoMAE: Masked autoencoders are data-efficient learners for self-supervised video pre-training

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:55.968978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:10:55.369731Z digest=sha256:0c3d3a93088d206d9a852e2fd09cf1c937450fd7039ad7350e003fb703ce1aaa

Observation 2ffa4df1-3728-42d1-89f9-fbbfc6e56103 · outbound

This paper cites Towards Accurate Generative Models of Video: A New Metric & Challenges.

From Image to Video: An Empirical Study of Diffusion Representations Towards Accurate Generative Models of Video: A New Metric & Challenges

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-08T14:10:55.374419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:10:55.374419Z digest=sha256:770b3284b4148f8ef252b749c26493eb6c0773c281ae4c459dfcfe6da70ec0ef

Observation 94445917-140d-4ed1-83c5-d54ac64332f1 · outbound

This paper cites Neural discrete representation learning.

From Image to Video: An Empirical Study of Diffusion Representations Neural discrete representation learning

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:55.953590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:10:55.379343Z digest=sha256:12911ec25563c10cee8ed041d36f591909d9f0cab920654e042dd6630342e6ee

Observation 958500c9-b3bb-4d5f-9c02-e736c5693d7a · outbound

This paper cites The iNaturalist species classification and detection dataset.

From Image to Video: An Empirical Study of Diffusion Representations The iNaturalist species classification and detection dataset

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:55.938315Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:10:55.383979Z digest=sha256:cf1c930f0add481271004a968f2b96c35602ab0705b79985fed6235910012875

Observation af177f6d-7537-46b7-923d-127918f55d41 · outbound

This paper cites Hudson, Thomas Albert Keck, Joao Carreira, Alexey Doso- vitskiy, Mehdi S.

From Image to Video: An Empirical Study of Diffusion Representations Hudson, Thomas Albert Keck, Joao Carreira, Alexey Doso- vitskiy, Mehdi S

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:55.923033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:10:55.389085Z digest=sha256:5314585175d2b45c0a34934a7999ed0b0728198e2d181f03bf2ecf40acb0d1ba

Observation 3c76d85f-a3a4-41f4-b74a-885ad9d09450 · outbound

This paper cites Attention is all you need.

From Image to Video: An Empirical Study of Diffusion Representations Attention is all you need

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:55.907171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:10:55.393683Z digest=sha256:4973c8d7129504230ab1427a8a92fc066ddae1fe38cb3081bc8c66ff74312c58

Observation 7502c86a-a9a9-438b-82d0-d8af2e1c7a85 · outbound

This paper cites VideoMAE v2: Scaling video masked autoencoders with dual masking.

From Image to Video: An Empirical Study of Diffusion Representations VideoMAE v2: Scaling video masked autoencoders with dual masking

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:55.891728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:10:55.398243Z digest=sha256:8af355a96c326b32a511a48a87a0af08fad163444aa9789fdff8e772b6aacbd1

Observation 2f9ede65-af59-4efc-aee6-1dbd688e7143 · outbound

This paper cites Controlling Space and Time with Diffusion Models.

From Image to Video: An Empirical Study of Diffusion Representations Controlling Space and Time with Diffusion Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-08T14:10:55.402609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:10:55.402609Z digest=sha256:af8e3046832700d752f3e0748c4683af543ff71a6027126b4c3286d89a9b151d

Observation 0fe87e0a-d124-4519-87a1-eba6c744e2ac · outbound

This paper cites Denoising diffusion autoencoders are unified self-supervised learners.

From Image to Video: An Empirical Study of Diffusion Representations Denoising diffusion autoencoders are unified self-supervised learners

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:55.874562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:10:55.407888Z digest=sha256:83e7a4e5b9e988a5931065f9b3506aee6a107fa4d446e9fd045283b6b728c5b4

Observation a6a4e390-b173-415c-9063-ac78ec4bac71 · outbound

This paper cites Open-vocabulary panop- tic segmentation with text-to-image diffusion models.

From Image to Video: An Empirical Study of Diffusion Representations Open-vocabulary panop- tic segmentation with text-to-image diffusion models

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:55.858262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:10:55.412593Z digest=sha256:82397e5034ce402aec3cf41117e459a5ad462aa4262f3e02ab51ae2127bc0ba3

Observation 27edcb1a-eb8b-4d72-8352-ff598a435d7b · outbound

This paper cites Diffusion Model as Rep- resentation Learner.

From Image to Video: An Empirical Study of Diffusion Representations Diffusion Model as Rep- resentation Learner

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:55.842384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:10:55.417404Z digest=sha256:c76632e8b2f099f5564f059f338e41541ec91352da90c176c3112d5ced085c44

Observation b0ac0337-c4a8-47d4-ba5a-16d07178cb95 · outbound

This paper cites Gundavarapu, Luca Ver- sari, Kihyuk Sohn, David Minnen, Yong Cheng, Vigh- nesh Birodkar, Agrim Gupta, Xiuye Gu, Alexander G.

From Image to Video: An Empirical Study of Diffusion Representations Gundavarapu, Luca Ver- sari, Kihyuk Sohn, David Minnen, Yong Cheng, Vigh- nesh Birodkar, Agrim Gupta, Xiuye Gu, Alexander G

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:55.826303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:10:55.421970Z digest=sha256:6bb09f64f1d26a9cde28d401bb6dbae9982422767fa7e963b66131ed525be02b

Observation af5a3197-f2c1-4c5a-901f-b84be0514da7 · outbound

This paper cites A tale of two features: Stable diffusion complements DINO for zero-shot semantic correspondence.

From Image to Video: An Empirical Study of Diffusion Representations A tale of two features: Stable diffusion complements DINO for zero-shot semantic correspondence

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:55.809804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:10:55.426588Z digest=sha256:6a4f10997ecfa7f1b70a66488d0edc7fcfdc938f9b62167f1587ad3178ab1a8b

Observation adc973b1-8a9f-4c9f-810d-c9d4f59e998a · outbound

This paper cites A Survey of Diffusion Based Image Generation Models: Issues and Their Solutions.

From Image to Video: An Empirical Study of Diffusion Representations A Survey of Diffusion Based Image Generation Models: Issues and Their Solutions

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-08T14:10:55.431225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:10:55.431225Z digest=sha256:5980bb9d4c8ec1edf62156fd255bf22a381331523203a136774452583139249c

Observation 54c34bff-d2ce-477f-8bfd-a6ac1a5c169a · outbound

This paper cites an unresolved cited work.

From Image to Video: An Empirical Study of Diffusion Representations Unresolved cited work

Reference 65

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:10:55.793939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:10:55.437068Z digest=sha256:109a4cdba5bee848ddbf2b7214b245bfc10153c71d8d0fa4a6c27a26040d41a4

Observation 1462df58-ed28-4474-a511-dcab881afb14 · outbound

This paper cites Unleashing text-to-image diffusion models for visual perception.

From Image to Video: An Empirical Study of Diffusion Representations Unleashing text-to-image diffusion models for visual perception

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-08T14:10:55.441538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:10:55.441538Z digest=sha256:69fa42d413ed4694006b7ba21b1e410482c357a2debc24a140bf2a8dabe6f168

Observation 0f76226c-d04b-4335-801b-81acca099df6 · outbound

This paper cites Places: A 10 million image database for scene recognition.

From Image to Video: An Empirical Study of Diffusion Representations Places: A 10 million image database for scene recognition

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:55.767105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:10:55.446497Z digest=sha256:466c1224fffa302e80dc75b60314b7b132e8f7971b6d99ec5a8bbd7604b9671e

Observation 3b551034-2701-48cf-a3d5-164bf2b1f3d6 · outbound

This paper cites Stereo magnification: Learning view syn- thesis using multiplane images.

From Image to Video: An Empirical Study of Diffusion Representations Stereo magnification: Learning view syn- thesis using multiplane images

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:55.750105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:10:55.450982Z digest=sha256:511f3d7f854994f02faa684af470d161cf5a13c1b9b24ba86229a8d2f4fec7c6

Observation 6e333442-b790-49a6-871e-4f4bcf9c06e7 · outbound

This paper cites Exploring Pre-trained Text-to-Video Diffusion Models for Referring Video Object Segmentation.

From Image to Video: An Empirical Study of Diffusion Representations Exploring Pre-trained Text-to-Video Diffusion Models for Referring Video Object Segmentation

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-08T14:10:55.455485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:10:55.455485Z digest=sha256:690a62a8d348d4485425b12b694b50b0edd5a24167b6fcb1aae656e1f5ca7620

Pith citing papers

Observation a2d5ddbd-90f3-4119-9328-a330dd5eaf9c · inbound

Frozen Forecasting: A Unified Evaluation cites this paper.

Frozen Forecasting: A Unified Evaluation From Image to Video: An Empirical Study of Diffusion Representations

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-19T03:42:57.304495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-19T03:42:54.620069Z digest=sha256:838f7ee0b3363650c09d3369e054475633e3c565676a51111400e95e838fb94b

Observation 0b36fdbc-3921-4f57-8d81-c211e5d371ff · inbound

Video Generation with Predictive Latents cites this paper.

Video Generation with Predictive Latents From Image to Video: An Empirical Study of Diffusion Representations

Reference 48

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T16:26:09.069938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-09T16:55:54.709581Z digest=sha256:f95c75e9e2159bff1bca306b912ae466429de2223c1e1c08f2cbc0749d6a9264