Pith. sign in

Paper Citation Record · LEDGER

From Image to Video: An Empirical Study of Diffusion Representations

As of 9 August 2026, this Paper Citation Record lists 69 of 69 outbound references and 2 inbound Pith citation observations for arXiv:2502.07001.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.07001 v2

Coverage vector

measured 69 of 69 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T14:10:55.455485Z

measured 71 of 71 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-19T03:42:54.620069Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-19T03:42:57.302590Z

Reference resolution

69 of 69 outbound references displayed

  • verified exact1
  • verified fuzzy42
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 603821a0-9355-48c5-9aeb-a0481d66b0c7 · outbound

This paper cites Self-supervised learning from images with a joint-embedding predictive architecture.

From Image to Video: An Empirical Study of Diffusion Representations Self-supervised learning from images with a joint-embedding predictive architecture

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:56.543774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T14:10:55.116949Z digest=sha256:ba4df4e5332440add053258e5c6c2a7b705a4fde86db4d8271735a9502138653

Observation cf7f64f8-25c1-4c34-910c-da56a3c62b98 · outbound

This paper cites Video diffusion models learn the struc- ture of the dynamic world.

From Image to Video: An Empirical Study of Diffusion Representations Video diffusion models learn the struc- ture of the dynamic world

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:56.529087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T14:10:55.122254Z digest=sha256:2c731b9de3e539b1fdfa1619a847e55466435e1760e68997e180a8b865f70f75

Observation 84998fb5-9c41-4862-ad36-25ae68f7595f · outbound

This paper cites Learning by Reconstruction Produces Uninformative Features For Perception.

From Image to Video: An Empirical Study of Diffusion Representations Learning by Reconstruction Produces Uninformative Features For Perception

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T14:10:55.127624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:10:55.127624Z digest=sha256:29b9d5e7a3f8e499e53fda7c11a6937fe19bf588c552f888b22f2c0803a5432d

Observation e8f2fc10-1067-4eb7-b442-9c176f7863b6 · outbound

This paper cites Label-efficient se- mantic segmentation with diffusion models.

From Image to Video: An Empirical Study of Diffusion Representations Label-efficient se- mantic segmentation with diffusion models

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:56.513716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T14:10:55.133861Z digest=sha256:8ad074581384ead9b36f3c9814ac2a25f9ab944170b037865b12a4c4c480d2a3

Observation a21fbcaf-e472-4e30-a2ab-a570cd1ceb2a · outbound

This paper cites Revisiting Feature Prediction for Learning Visual Representations from Video.

From Image to Video: An Empirical Study of Diffusion Representations Revisiting Feature Prediction for Learning Visual Representations from Video

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T14:10:55.138930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:10:55.138930Z digest=sha256:240a1b95a6b4893c440371f85f34380ed3f47291f0ac29be61a22a2508cb8119

Observation 65321793-a5fe-4987-8f30-8828dc9fc8d7 · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

From Image to Video: An Empirical Study of Diffusion Representations Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T14:10:55.144179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:10:55.144179Z digest=sha256:0dcefd1cd84709a75c227626c1d7377c29891e3dd41a24103a73b5ff95d9b4c3

Observation ae5814d3-b1c7-4258-a241-1b50a0120411 · outbound

This paper cites Deep regression on manifolds: a 3D rota- tion case study.

From Image to Video: An Empirical Study of Diffusion Representations Deep regression on manifolds: a 3D rota- tion case study

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:56.499007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T14:10:55.149974Z digest=sha256:079289088ad59bbe648c945de890e29e74f2ed77c5f9e98e8b693526958be5ff

Observation 2b88fef2-7f80-42e0-b5a4-fc00d09576d3 · outbound

This paper cites Emerg- ing properties in self-supervised vision transformers.

From Image to Video: An Empirical Study of Diffusion Representations Emerg- ing properties in self-supervised vision transformers

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:56.482929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T14:10:55.154906Z digest=sha256:033d4191d26bf6e049c8053154d389779ab5a9babed755622d4a1eef134a249a

Observation 2c952667-f4b4-437d-946e-a2045f6b52b9 · outbound

This paper cites Quo vadis, action recognition? a new model and the kinetics dataset.

From Image to Video: An Empirical Study of Diffusion Representations Quo vadis, action recognition? a new model and the kinetics dataset

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T14:10:55.159745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:10:55.159745Z digest=sha256:d700e2df2b0e4769620f669bbf54a96710e8f3e1aed05d707242d405d9de2c6b

Observation ddf9d714-8dfe-4364-b1e4-2a330064446e · outbound

This paper cites A Short Note on the Kinetics-700 Human Action Dataset.

From Image to Video: An Empirical Study of Diffusion Representations A Short Note on the Kinetics-700 Human Action Dataset

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T14:10:55.165909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:10:55.165909Z digest=sha256:de7cdcbea34d803a9615c433f69db3dca7c7452c54d018f690a535aa85b6eed1

Observation 647ea506-fb21-412f-8709-191a86556497 · outbound

This paper cites Scaling 4D Representations.

From Image to Video: An Empirical Study of Diffusion Representations Scaling 4D Representations

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T14:10:55.171296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:10:55.171296Z digest=sha256:0f937f40bc2b48df7611ac65398084a617770509132f28ecfc9b8ac794390b57

Observation fd97b2ea-81ad-4ab1-84fb-bf8e4fc288ed · outbound

This paper cites an unresolved cited work.

From Image to Video: An Empirical Study of Diffusion Representations Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:10:56.454616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T14:10:55.176297Z digest=sha256:8367494622dbb4a3e601f1721ddabed975e70ee8575d513998c07042a1f15b69

Observation 62146370-3bbb-4834-85ed-c1be60865c89 · outbound

This paper cites PaLI-3 Vision Language Models: Smaller, Faster, Stronger.

From Image to Video: An Empirical Study of Diffusion Representations PaLI-3 Vision Language Models: Smaller, Faster, Stronger

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T14:10:55.182071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:10:55.182071Z digest=sha256:c776701586fe1e0f7cf2ac5de106d359dbb9c3045efc0a5b00b1ed0e80a628f2

Observation c7fff663-46f1-4cee-9c79-59e461514fc6 · outbound

This paper cites Text-to-image diffusion mod- els are zero shot classifiers.

From Image to Video: An Empirical Study of Diffusion Representations Text-to-image diffusion mod- els are zero shot classifiers

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:56.437657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T14:10:55.187263Z digest=sha256:4a29d9dd246ca73045f684b1a0927fd78f96ca4f40b7f9bd1d6530853c52ac92

Observation 6babfe3c-9302-4e56-9aee-1c71d713bea9 · outbound

This paper cites Chang, Manolis Savva, Maciej Hal- ber, Thomas Funkhouser, and Matthias Nießner.

From Image to Video: An Empirical Study of Diffusion Representations Chang, Manolis Savva, Maciej Hal- ber, Thomas Funkhouser, and Matthias Nießner

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:56.421615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T14:10:55.191848Z digest=sha256:0a9385892cbcaee03563c2bdef6f8bf04fb399443e015bf575fce7eb149e8179

Observation ceb7d32a-da1d-453d-bd88-0f0e747d5aa5 · outbound

This paper cites Depth map prediction from a single image using a multi-scale deep net- work.

From Image to Video: An Empirical Study of Diffusion Representations Depth map prediction from a single image using a multi-scale deep net- work

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:56.404034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T14:10:55.196339Z digest=sha256:ba21471c797cb71fa8b87c7915cd1fa96ada14a43ef381172cffa9317177b5d3

Observation a71f2f6c-c156-4b51-a343-bf48d93d0515 · outbound

This paper cites Taming transformers for high-resolution image synthesis.

From Image to Video: An Empirical Study of Diffusion Representations Taming transformers for high-resolution image synthesis

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T14:10:55.200902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:10:55.200902Z digest=sha256:4e65bb7610286e26524b0f01503c98271380aed2a1922d36113249e324f53693

Observation 48a5b6a2-1bea-47c0-9fda-cc34bcdc70fc · outbound

This paper cites Masked autoencoders as spatiotemporal learners.

From Image to Video: An Empirical Study of Diffusion Representations Masked autoencoders as spatiotemporal learners

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:56.374861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T14:10:55.206228Z digest=sha256:62f597d52362c784b2058f6b5a850814c686bd6c570287121488c1ce2ce6a685

Observation ade7fd2f-118a-4d8b-a295-b245f66e7678 · outbound

This paper cites Diffusion Models and Representation Learning: A Survey.

From Image to Video: An Empirical Study of Diffusion Representations Diffusion Models and Representation Learning: A Survey

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T14:10:55.211931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:10:55.211931Z digest=sha256:d2ea526ee973c80844d81b83424075683969e5c415d00e2972f314cbd1229391

Observation 16a378de-c69e-41e8-8dcc-8909361ab035 · outbound

This paper cites Something Something.

From Image to Video: An Empirical Study of Diffusion Representations Something Something

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:56.354536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T14:10:55.217130Z digest=sha256:641dcb2a4f9d8c69b0f60e56b93c95d338ab1fcfcf56c9c9e7fd05a9ec8c3762

Observation 20c8d60a-5b6a-427b-b8fd-cfcf907d1c47 · outbound

This paper cites Kubric: A scalable dataset generator.

From Image to Video: An Empirical Study of Diffusion Representations Kubric: A scalable dataset generator

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:56.338487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T14:10:55.222921Z digest=sha256:88bee9d5923d999eaaf25e0edaf0b45a2e21e2d870faea72f50c6a5a3e4e1a27

Observation 8d92fef7-1769-4742-b1e0-aeed694afa2b · outbound

This paper cites Photorealistic video generation with diffusion models.

From Image to Video: An Empirical Study of Diffusion Representations Photorealistic video generation with diffusion models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:56.321232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T14:10:55.227765Z digest=sha256:c02a346dd10c38171930ada61c2635a39fd7c00307e67ce47b5a4e8b802f02c0

Observation 7302ca0f-4509-4216-a41e-7b1fd9d63e17 · outbound

This paper cites Masked autoencoders are scalable vision learners.

From Image to Video: An Empirical Study of Diffusion Representations Masked autoencoders are scalable vision learners

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:56.303949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T14:10:55.232445Z digest=sha256:afcd602c731c98bf3f4d9379cdbb1719b2df7f1f32060392053463aa9789b568

Observation 737ae461-0659-4691-a059-b1c40a504c27 · outbound

This paper cites Unsupervised keypoints from pretrained diffusion models.

From Image to Video: An Empirical Study of Diffusion Representations Unsupervised keypoints from pretrained diffusion models

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:56.288117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T14:10:55.237362Z digest=sha256:cb635b40293bf65a34911f87b20cd1c0940a648775d6533f62aee98fd39d9d3c

Observation a2be131a-35d6-40b5-873b-dd347fdbb184 · outbound

This paper cites Denoising diffu- sion probabilistic models.

From Image to Video: An Empirical Study of Diffusion Representations Denoising diffu- sion probabilistic models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T14:10:55.242828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:10:55.242828Z digest=sha256:bb532bd625e318bdc6adee7c744e38d793cc6225f67b52c4e2cedb1ddb034050

Observation a32fe247-85c8-4dc7-8b9e-4ab376ec1002 · outbound

This paper cites DepthCrafter: Generating Consistent Long Depth Sequences for Open-world Videos.

From Image to Video: An Empirical Study of Diffusion Representations DepthCrafter: Generating Consistent Long Depth Sequences for Open-world Videos

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T14:10:55.247542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:10:55.247542Z digest=sha256:c82fccb753adbe58636fd2c022c48c36435bf3bb3ed2d66ca96412de12731d7a

Observation 68812302-5910-4492-ada0-0d5cf32723e1 · outbound

This paper cites Elucidating the design space of diffusion-based generative models.

From Image to Video: An Empirical Study of Diffusion Representations Elucidating the design space of diffusion-based generative models

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:56.261203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T14:10:55.252875Z digest=sha256:853bed0f395726f2efa128a43304736cb466e4322b50a75fa935864c8b081d25

Observation 87a222f0-25fc-41b2-bad7-3d32451b08f8 · outbound

This paper cites Your diffusion model is secretly a zero-shot classifier.

From Image to Video: An Empirical Study of Diffusion Representations Your diffusion model is secretly a zero-shot classifier

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:56.243722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T14:10:55.257709Z digest=sha256:2eaa4d1b9af50525859040a852026c250820f7b281fa3f4f06c97462cb5f25a6

Observation c3a21647-dd1e-4432-8077-36b0cb0d8030 · outbound

This paper cites Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models.

From Image to Video: An Empirical Study of Diffusion Representations Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T14:10:55.263104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:10:55.263104Z digest=sha256:9a1712b964d59c79fa551cd8e89773f3e6b2f6a031cd08c5187ea79b5e99d3d9

Observation fe5d3173-800e-4721-9961-1d22c15a3088 · outbound

This paper cites Diffusion hyperfeatures: Search- ing through time and space for semantic correspondence.

From Image to Video: An Empirical Study of Diffusion Representations Diffusion hyperfeatures: Search- ing through time and space for semantic correspondence

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:56.226488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T14:10:55.268287Z digest=sha256:1641e27e2a130b5e38d2249dcf2312cbbeee89cdf940286ba14bb62fc65ecd4d

Observation db54a12f-71f6-4882-b412-13304013a4d5 · outbound

This paper cites Understanding deep image representations by inverting them.

From Image to Video: An Empirical Study of Diffusion Representations Understanding deep image representations by inverting them

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:56.211266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T14:10:55.272944Z digest=sha256:e24aea6d73c3642212aaf4ca82634043ef508c51f8011e55c0006109c8d28dd0

Observation f6c3539d-ca5e-4a6c-b193-e525ed96fc62 · outbound

This paper cites Lexicon3d: Probing vi- sual foundation models for complex 3d scene understanding.

From Image to Video: An Empirical Study of Diffusion Representations Lexicon3d: Probing vi- sual foundation models for complex 3d scene understanding

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:56.194835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T14:10:55.277612Z digest=sha256:073b572fcbeb2b3d3cd05b728366e5f4fbb5f95015806df0d66ff8cd46b17b05

Observation 24f582d3-87f7-4921-a4c4-d2b6ba373ad6 · outbound

This paper cites NeRF: Representing scenes as neural radiance fields for view syn- thesis.

From Image to Video: An Empirical Study of Diffusion Representations NeRF: Representing scenes as neural radiance fields for view syn- thesis

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:56.178589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T14:10:55.282206Z digest=sha256:df6387c6b9fa206fcedd104784a28e7d33d4f91a47840416b4fb7ee82bc33a8e

Observation 7cc0493e-586f-4637-b833-9ad2376161ba · outbound

This paper cites Diffusion Models Beat GANs on Image Classification.

From Image to Video: An Empirical Study of Diffusion Representations Diffusion Models Beat GANs on Image Classification

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T14:10:55.286821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:10:55.286821Z digest=sha256:d0de7fe315f256dd2a25c2db63b2540045ba4c9fbd05403585a2d90b04b36d16

Observation 41f6959e-fb2f-4959-a742-3f5306b03908 · outbound

This paper cites DiffTAD: Temporal action detection with proposal denoising diffusion.

From Image to Video: An Empirical Study of Diffusion Representations DiffTAD: Temporal action detection with proposal denoising diffusion

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:56.162756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T14:10:55.291762Z digest=sha256:2199baf8e1f5d6bdcd74d658fff4e2734f83d5a370331309d24cca37291bbc17

Observation 297db8f9-bf23-4d29-abec-819a1b70cac8 · outbound

This paper cites an unresolved cited work.

From Image to Video: An Empirical Study of Diffusion Representations Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-08T14:10:55.296307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:10:55.296307Z digest=sha256:ea0a7c6ec87b128716b320c82bfabc11f3d002343e025973d039f3aeda8c983a

Observation 1bcfb7f4-2bf5-4b4b-98e8-0b52de14c06d · outbound

This paper cites Self-supervised video pretraining yields robust and more human-aligned visual representations.

From Image to Video: An Empirical Study of Diffusion Representations Self-supervised video pretraining yields robust and more human-aligned visual representations

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:56.136122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T14:10:55.301353Z digest=sha256:bac6d872bb098ed90074815288c4243fa33d5dde5003126f348a842105828ba6

Observation 08856bea-5579-4067-91d6-e9ddf695b466 · outbound

This paper cites Per- ception Test: A diagnostic benchmark for multimodal video models.

From Image to Video: An Empirical Study of Diffusion Representations Per- ception Test: A diagnostic benchmark for multimodal video models

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:56.120448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T14:10:55.306794Z digest=sha256:538941e36a9f52cbf64ce4376b6aabff533cb9c061b135286b94f43bf44b2aaa

Observation babb581f-0cb3-4ed5-a5c6-ed0d335fa9bc · outbound

This paper cites Scalable diffusion models with transformers.

From Image to Video: An Empirical Study of Diffusion Representations Scalable diffusion models with transformers

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-08T14:10:55.311336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:10:55.311336Z digest=sha256:a51c53551c0a9378165fb87ec230aa58f6f8cc35923bad34dc725be2da35f1bd

Observation 08672b8d-6e27-4aa5-a28f-9001a1def01e · outbound

This paper cites The 2017 DAVIS Challenge on Video Object Segmentation.

From Image to Video: An Empirical Study of Diffusion Representations The 2017 DAVIS Challenge on Video Object Segmentation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-08T14:10:55.315744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:10:55.315744Z digest=sha256:a0841bf5361d4bf7b162ee34ab22466d4bb7b4a191072d979b9c301729b01954

Observation a81a3c84-5fe0-4f51-a6bc-e6a6fb922909 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

From Image to Video: An Empirical Study of Diffusion Representations Learn- ing transferable visual models from natural language super- vision

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-08T14:10:55.320638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:10:55.320638Z digest=sha256:e922842491dad20fa508b292eec708d4cfa920da0f92d6ff99090d626d0cd640

Observation a5c0f855-667d-4bca-8468-481b32e18231 · outbound

This paper cites High-resolution image syn- thesis with latent diffusion models.

From Image to Video: An Empirical Study of Diffusion Representations High-resolution image syn- thesis with latent diffusion models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-08T14:10:55.326083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:10:55.326083Z digest=sha256:3c2db964d93a48636aeb4740f80b060e89edada37f286dee39d350c51d22a548

Observation 116c1a14-2760-4986-a65c-d87aa487f408 · outbound

This paper cites U- Net: Convolutional networks for biomedical image segmen- tation.

From Image to Video: An Empirical Study of Diffusion Representations U- Net: Convolutional networks for biomedical image segmen- tation

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:56.074226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T14:10:55.331019Z digest=sha256:c99dfe4fce73c5f50a30772fdedb11c295641be02ef93b0cb85f36d254883d9f

Observation dbfba91e-e117-43b3-a687-098d0627e39b · outbound

This paper cites Berg, and Li Fei-Fei.

From Image to Video: An Empirical Study of Diffusion Representations Berg, and Li Fei-Fei

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:56.056491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T14:10:55.335679Z digest=sha256:9f8a3bdd7bf226264678756fde591e44a6434e52d71ada1a8e22bd7fce87cbf7

Observation c6ab9844-88b4-41a7-b0c7-50e4d4d98c4c · outbound

This paper cites Scene representation transformer: Geometry-free novel view syn- thesis through set-latent scene representations.

From Image to Video: An Empirical Study of Diffusion Representations Scene representation transformer: Geometry-free novel view syn- thesis through set-latent scene representations

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:56.041323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T14:10:55.340453Z digest=sha256:2dd14e9be54112917f6417df92cad840d57557a4090a83bc3337e1a7ac35baa6

Observation 3b8c6641-290a-4a7d-8e5b-11179bae4712 · outbound

This paper cites Only time can tell: Discovering temporal data for temporal modeling.

From Image to Video: An Empirical Study of Diffusion Representations Only time can tell: Discovering temporal data for temporal modeling

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:56.025424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T14:10:55.345836Z digest=sha256:844e2079f0efdfd334e57ed1a07934b2fef0b0794c3d3fa7e55ca6d386c09fdd

Observation 6fb7003e-13f5-43c8-bb60-fff423a83da4 · outbound

This paper cites MonoDiffusion: Self-Supervised Monocular Depth Estimation Using Diffusion Model.

From Image to Video: An Empirical Study of Diffusion Representations MonoDiffusion: Self-Supervised Monocular Depth Estimation Using Diffusion Model

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-08-08T14:10:55.563436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T14:10:55.350854Z digest=sha256:166b420c1ab6eb8719f8a25a26abc5f5ae1eebff4aefc7006f8ad5d513011a5e

Observation f154804d-1108-4532-90cb-1f8bf6e6e25a · outbound

This paper cites Deep unsupervised learning using nonequilibrium thermodynamics.

From Image to Video: An Empirical Study of Diffusion Representations Deep unsupervised learning using nonequilibrium thermodynamics

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-08T14:10:55.355636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:10:55.355636Z digest=sha256:48e7342f7a34b0506689d5329099935564997740b884e7e1611fb9e982a35939

Observation 7b302f72-d9ab-45a3-bb7e-53a11f481eb2 · outbound

This paper cites Scalability in perception for autonomous driving: Waymo Open Dataset.

From Image to Video: An Empirical Study of Diffusion Representations Scalability in perception for autonomous driving: Waymo Open Dataset

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:55.998842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T14:10:55.360284Z digest=sha256:4a31e4220642a0d398329b3bdb35d9cd1ad1cc5f08aaba489206a16053c658f2

Observation 23f0ce83-05b5-4ff3-bb97-8c14bd76ba9e · outbound

This paper cites Emergent correspondence from image diffusion.

From Image to Video: An Empirical Study of Diffusion Representations Emergent correspondence from image diffusion

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:55.984019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T14:10:55.364835Z digest=sha256:75780a96a90513cf890fef2aa6b9c63b50452eb8f49cc65db034baf07712ce4e

Observation ea064e57-1183-4773-92f6-55510f4c043b · outbound

This paper cites VideoMAE: Masked autoencoders are data-efficient learners for self-supervised video pre-training.

From Image to Video: An Empirical Study of Diffusion Representations VideoMAE: Masked autoencoders are data-efficient learners for self-supervised video pre-training

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:55.968978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T14:10:55.369731Z digest=sha256:13aea8d3dd139ffb276ed94354e292d28b55389ac46112b6417c9d03089005d4

Observation 2ffa4df1-3728-42d1-89f9-fbbfc6e56103 · outbound

This paper cites Towards Accurate Generative Models of Video: A New Metric & Challenges.

From Image to Video: An Empirical Study of Diffusion Representations Towards Accurate Generative Models of Video: A New Metric & Challenges

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-08T14:10:55.374419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:10:55.374419Z digest=sha256:c63207b7b39aee3dfe6849bdc92f62edd209ea1c70676746e41c46790d6611b2

Observation 94445917-140d-4ed1-83c5-d54ac64332f1 · outbound

This paper cites Neural discrete representation learning.

From Image to Video: An Empirical Study of Diffusion Representations Neural discrete representation learning

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:55.953590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T14:10:55.379343Z digest=sha256:699a139f0a1e1cfd79da113206ef32a53b1a769bbeec5a8a07a77b3af519de8d

Observation 958500c9-b3bb-4d5f-9c02-e736c5693d7a · outbound

This paper cites The iNaturalist species classification and detection dataset.

From Image to Video: An Empirical Study of Diffusion Representations The iNaturalist species classification and detection dataset

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:55.938315Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T14:10:55.383979Z digest=sha256:ce8828a1ecd13df78258187520f3eb39375d42f1878a7b9c6aa10d0fa5dacdc4

Observation af177f6d-7537-46b7-923d-127918f55d41 · outbound

This paper cites Hudson, Thomas Albert Keck, Joao Carreira, Alexey Doso- vitskiy, Mehdi S.

From Image to Video: An Empirical Study of Diffusion Representations Hudson, Thomas Albert Keck, Joao Carreira, Alexey Doso- vitskiy, Mehdi S

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:55.923033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T14:10:55.389085Z digest=sha256:9fd259b076e476b0277b14b4e4a030d34d19aff093b7ee157ccfe2093825f5d2

Observation 3c76d85f-a3a4-41f4-b74a-885ad9d09450 · outbound

This paper cites Attention is all you need.

From Image to Video: An Empirical Study of Diffusion Representations Attention is all you need

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:55.907171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T14:10:55.393683Z digest=sha256:47779f539a863145a7c94ea165f887a588043a994c55616a4065492e50246199

Observation 7502c86a-a9a9-438b-82d0-d8af2e1c7a85 · outbound

This paper cites VideoMAE v2: Scaling video masked autoencoders with dual masking.

From Image to Video: An Empirical Study of Diffusion Representations VideoMAE v2: Scaling video masked autoencoders with dual masking

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:55.891728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T14:10:55.398243Z digest=sha256:2fb92ad7a2cbe411f0889000ba0dd29bebb8b297ed4fd207919a9f314344ac6e

Observation 2f9ede65-af59-4efc-aee6-1dbd688e7143 · outbound

This paper cites Controlling Space and Time with Diffusion Models.

From Image to Video: An Empirical Study of Diffusion Representations Controlling Space and Time with Diffusion Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-08T14:10:55.402609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:10:55.402609Z digest=sha256:e967ccb753316134c6abd343b549a2c6e9f0f53e44f59bd45bcfc584db13f594

Observation 0fe87e0a-d124-4519-87a1-eba6c744e2ac · outbound

This paper cites Denoising diffusion autoencoders are unified self-supervised learners.

From Image to Video: An Empirical Study of Diffusion Representations Denoising diffusion autoencoders are unified self-supervised learners

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:55.874562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T14:10:55.407888Z digest=sha256:793782b6a20f9a8a24e1f5587271863307347c605889e6310abb9912987d8860

Observation a6a4e390-b173-415c-9063-ac78ec4bac71 · outbound

This paper cites Open-vocabulary panop- tic segmentation with text-to-image diffusion models.

From Image to Video: An Empirical Study of Diffusion Representations Open-vocabulary panop- tic segmentation with text-to-image diffusion models

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:55.858262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T14:10:55.412593Z digest=sha256:7d99e5a101914140fdb3430ff41304acd6d1c6a2f683d9a1312656ed3530b95f

Observation 27edcb1a-eb8b-4d72-8352-ff598a435d7b · outbound

This paper cites Diffusion Model as Rep- resentation Learner.

From Image to Video: An Empirical Study of Diffusion Representations Diffusion Model as Rep- resentation Learner

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:55.842384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T14:10:55.417404Z digest=sha256:20d284f0f78c10a6d9f022e742d9792d401d958ded0ed3b2614b554ea9800c01

Observation b0ac0337-c4a8-47d4-ba5a-16d07178cb95 · outbound

This paper cites Gundavarapu, Luca Ver- sari, Kihyuk Sohn, David Minnen, Yong Cheng, Vigh- nesh Birodkar, Agrim Gupta, Xiuye Gu, Alexander G.

From Image to Video: An Empirical Study of Diffusion Representations Gundavarapu, Luca Ver- sari, Kihyuk Sohn, David Minnen, Yong Cheng, Vigh- nesh Birodkar, Agrim Gupta, Xiuye Gu, Alexander G

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:55.826303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T14:10:55.421970Z digest=sha256:65a5390fe5fea0ea610cfb1a094e9180b8422819309019f8f92bac186cb0d6e2

Observation af5a3197-f2c1-4c5a-901f-b84be0514da7 · outbound

This paper cites A tale of two features: Stable diffusion complements DINO for zero-shot semantic correspondence.

From Image to Video: An Empirical Study of Diffusion Representations A tale of two features: Stable diffusion complements DINO for zero-shot semantic correspondence

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:55.809804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T14:10:55.426588Z digest=sha256:b53ae598066c5054e021a537f1b7e018fb16d0e2f442b6e1f7d9205729816c79

Observation adc973b1-8a9f-4c9f-810d-c9d4f59e998a · outbound

This paper cites A Survey of Diffusion Based Image Generation Models: Issues and Their Solutions.

From Image to Video: An Empirical Study of Diffusion Representations A Survey of Diffusion Based Image Generation Models: Issues and Their Solutions

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-08T14:10:55.431225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:10:55.431225Z digest=sha256:a97b30027c3dd91e828ad0bb6f5ceddf181639cae3d208890ec204ab7bd7a49f

Observation 54c34bff-d2ce-477f-8bfd-a6ac1a5c169a · outbound

This paper cites an unresolved cited work.

From Image to Video: An Empirical Study of Diffusion Representations Unresolved cited work

Reference 65

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:10:55.793939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T14:10:55.437068Z digest=sha256:a2efe776f267b781b027a380d9ccde8aab5b796387d93bbed9a8734fcb2d6f14

Observation 1462df58-ed28-4474-a511-dcab881afb14 · outbound

This paper cites Unleashing text-to-image diffusion models for visual perception.

From Image to Video: An Empirical Study of Diffusion Representations Unleashing text-to-image diffusion models for visual perception

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-08T14:10:55.441538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:10:55.441538Z digest=sha256:b5a3ee697b4d4c09a2d92156f99716b5cc48d3e521c541543dfcc1d664fe6650

Observation 0f76226c-d04b-4335-801b-81acca099df6 · outbound

This paper cites Places: A 10 million image database for scene recognition.

From Image to Video: An Empirical Study of Diffusion Representations Places: A 10 million image database for scene recognition

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:55.767105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T14:10:55.446497Z digest=sha256:b33ca6c0bc9634a021f9672d46980e7bc3eb2abb5df8c4486f44292953352198

Observation 3b551034-2701-48cf-a3d5-164bf2b1f3d6 · outbound

This paper cites Stereo magnification: Learning view syn- thesis using multiplane images.

From Image to Video: An Empirical Study of Diffusion Representations Stereo magnification: Learning view syn- thesis using multiplane images

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:10:55.750105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T14:10:55.450982Z digest=sha256:517dc045dd579b8579f6be1b006ee6307101d11a96367adf4e86facc609a91ef

Observation 6e333442-b790-49a6-871e-4f4bcf9c06e7 · outbound

This paper cites Exploring Pre-trained Text-to-Video Diffusion Models for Referring Video Object Segmentation.

From Image to Video: An Empirical Study of Diffusion Representations Exploring Pre-trained Text-to-Video Diffusion Models for Referring Video Object Segmentation

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-08T14:10:55.455485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:10:55.455485Z digest=sha256:364e19cecd9a69792c63ffcceb2edc1cbba03f2937745be8d46f9da3532f61be

Pith citing papers

Observation a2d5ddbd-90f3-4119-9328-a330dd5eaf9c · inbound

Frozen Forecasting: A Unified Evaluation cites this paper.

Frozen Forecasting: A Unified Evaluation From Image to Video: An Empirical Study of Diffusion Representations

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-19T03:42:57.304495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T03:42:54.620069Z digest=sha256:1dbcf4de6cf0f14ffc8d1a4973f8a30f101d15644c19e8492c77c9583e5fea80

Observation 0b36fdbc-3921-4f57-8d81-c211e5d371ff · inbound

Video Generation with Predictive Latents cites this paper.

Video Generation with Predictive Latents From Image to Video: An Empirical Study of Diffusion Representations

Reference 48

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T16:26:09.069938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T16:55:54.709581Z digest=sha256:52c5e830583d1cb6f7d0c21911261852a74526b85621b98867234a859825a95e