Pith. sign in

Paper Citation Record · LEDGER

Video Generation Models are General-Purpose Vision Learners

As of 14 August 2026, this Paper Citation Record lists 87 of 87 outbound references and 0 inbound Pith citation observations for arXiv:2607.09024.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.09024 v1

Coverage vector

measured 87 of 87 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-13T00:56:18.867382Z

measured 87 of 87 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

87 of 87 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved87
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 46f2b52c-3de1-473c-9ffb-ab0dfdc56485 · outbound

This paper cites From Generation to Generalization: Emergent Few-Shot Learning in Video Diffusion Models.

Video Generation Models are General-Purpose Vision Learners From Generation to Generalization: Emergent Few-Shot Learning in Video Diffusion Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-13T00:56:18.867382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:56:18.867382Z digest=sha256:d80630d56c4b41d767078f079ed124ab2ca04d835d967fcbcdb7d874eb0b649b

Observation df198e50-7374-4026-9ee1-ca56760e38f7 · outbound

This paper cites Advances in Neural Information Processing Systems37, 61872–61911 (2024) 3.

Video Generation Models are General-Purpose Vision Learners Advances in Neural Information Processing Systems37, 61872–61911 (2024) 3

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-13T00:56:18.867382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:56:18.867382Z digest=sha256:f618124aa27f91526488b2e02cedb6d245fc3f035a37a6d86626104b9b4f725b

Observation e2649b28-43be-48d2-ba22-e0d0bf1302f5 · outbound

This paper cites In: Proceedings of the IEEE/CVF International Conference on Computer Vision.

Video Generation Models are General-Purpose Vision Learners In: Proceedings of the IEEE/CVF International Conference on Computer Vision

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-13T00:56:18.867382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:56:18.867382Z digest=sha256:1bbe31156177ecd7c7643cb4018d0a740d1051da7152453a612f7e531f7688cc

Observation 3abcb894-b659-47dc-9789-76ed66baade5 · outbound

This paper cites Revisiting Feature Prediction for Learning Visual Representations from Video.

Video Generation Models are General-Purpose Vision Learners Revisiting Feature Prediction for Learning Visual Representations from Video

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-13T00:56:18.867382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:56:18.867382Z digest=sha256:425ce8ff16d3feb42a059c4162abd922a44f39b80c30b2ad2c1c6c6c7a046b31

Observation 05d88bf4-f476-493b-b396-ae9b1a7e80a7 · outbound

This paper cites HSPACE: Synthetic Parametric Humans Animated in Complex Environments.

Video Generation Models are General-Purpose Vision Learners HSPACE: Synthetic Parametric Humans Animated in Complex Environments

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-13T00:56:18.867382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:56:18.867382Z digest=sha256:ebb403df5d0cd2ef85098e191c0869b27606257fc08d74a0bc41e6a3d7bafca7

Observation 24e2e1f6-3289-4856-8f65-717a446a7b73 · outbound

This paper cites NormalCrafter: Learning Temporally Consistent Normals from Video Diffusion Priors.

Video Generation Models are General-Purpose Vision Learners NormalCrafter: Learning Temporally Consistent Normals from Video Diffusion Priors

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-13T00:56:18.867382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:56:18.867382Z digest=sha256:12d96ea379c8c7f2743928ce98efb207d470086a19880c420fd985a2eea45865

Observation 59e80ac5-d19a-4121-90d3-2830f7cee1f3 · outbound

This paper cites an unresolved cited work.

Video Generation Models are General-Purpose Vision Learners Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-13T00:56:18.867382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:56:18.867382Z digest=sha256:95906c77a9a26be9c0fa97a7c3f429ad99204e06e00b93a97d0c1e4b66893672

Observation ce2c13f8-f005-43d5-976b-93cc77e321af · outbound

This paper cites Advances in neural information processing systems33, 1877–1901 (2020) 3.

Video Generation Models are General-Purpose Vision Learners Advances in neural information processing systems33, 1877–1901 (2020) 3

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-13T00:56:18.867382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:56:18.867382Z digest=sha256:050c32ec63c0a6faa0e4e9f2c457d8b4d9897b84a35f725ba1b07d7f34cea0ae

Observation dd3bb213-52db-44be-be70-8cb564e074c7 · outbound

This paper cites In: European conference on computer vision.

Video Generation Models are General-Purpose Vision Learners In: European conference on computer vision

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-13T00:56:18.867382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:56:18.867382Z digest=sha256:cbdf4a58cd777ace40cb4b9a9ac72d1f02cf1c873bdc2a58e353a95247506960

Observation 63678d3d-0f6e-4ca9-b489-6c30de51d3e5 · outbound

This paper cites SAM 3: Segment Anything with Concepts.

Video Generation Models are General-Purpose Vision Learners SAM 3: Segment Anything with Concepts

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-13T00:56:18.867382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:56:18.867382Z digest=sha256:7e26471561af80d210523f094c090bd7626bc06d0345f9cf0f5a64801dbc12de

Observation 1d63bf22-c443-4e76-af7f-063eb2e40780 · outbound

This paper cites an unresolved cited work.

Video Generation Models are General-Purpose Vision Learners Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-13T00:56:18.867382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:56:18.867382Z digest=sha256:bf93afd1950f6b427934f6838d1907e46fd6a3fc12e10d66cd0e3278c5c08096

Observation faf4feaf-07f4-4faf-857f-6c84300d950c · outbound

This paper cites In: Proceedings of the IEEE/CVF international conference on computer vision.

Video Generation Models are General-Purpose Vision Learners In: Proceedings of the IEEE/CVF international conference on computer vision

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-13T00:56:18.867382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:56:18.867382Z digest=sha256:c2d3936c7f34ed7ea093dd637e8a0e93217acd32d1908da98b20a0076f15593c

Observation b4896612-c9d1-4705-a45b-eeda5a5ed296 · outbound

This paper cites Extending Context Window of Large Language Models via Positional Interpolation.

Video Generation Models are General-Purpose Vision Learners Extending Context Window of Large Language Models via Positional Interpolation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-13T00:56:18.867382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:56:18.867382Z digest=sha256:477c579280f142041ccc3aff08a5a3ec07e30f6e1934268a19c7dd23d13ce758

Observation fa4fc8d4-8e4e-4f43-89d9-9145be9f5c94 · outbound

This paper cites In: Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies, volume 1 (long and short papers).

Video Generation Models are General-Purpose Vision Learners In: Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies, volume 1 (long and short papers)

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-13T00:56:18.867382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:56:18.867382Z digest=sha256:2bc7f0f1b441fa3bf9831dd1311a8ad2558f3c652a2931e4ab45d6a7fc96643a

Observation 23a90ece-fdcb-498b-a98d-4a403ef68d8e · outbound

This paper cites an unresolved cited work.

Video Generation Models are General-Purpose Vision Learners Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-13T00:56:18.867382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:56:18.867382Z digest=sha256:204bf25afff0918caa74f067ed1eb8e928450a77383e77716b87829a2c41a85a

Observation 3b29b38d-0bcf-43aa-a845-370ff61a356d · outbound

This paper cites In: Forty-first international conference on machine learning (2024) 6.

Video Generation Models are General-Purpose Vision Learners In: Forty-first international conference on machine learning (2024) 6

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-13T00:56:18.867382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:56:18.867382Z digest=sha256:85c3b8e67fe84814a283fa16e86cb8fdc68d27f30a6b8fef4642cc4e65c9eca0

Observation d3866ee8-a48f-4cb4-b5b7-d8676c3f3c2b · outbound

This paper cites In: European Conference on Computer Vision (ECCV) (2024) 4.

Video Generation Models are General-Purpose Vision Learners In: European Conference on Computer Vision (ECCV) (2024) 4

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-13T00:56:18.867382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:56:18.867382Z digest=sha256:ae981d29ab7d3f4848f55c3df2d09766db8a573d643b0cf3fe62ea861669ad5d

Observation 5631ba55-81a8-4676-a702-c623c2fb9003 · outbound

This paper cites Image Generators are Generalist Vision Learners.

Video Generation Models are General-Purpose Vision Learners Image Generators are Generalist Vision Learners

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-13T00:56:18.867382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:56:18.867382Z digest=sha256:6993ffed9507ee301c67f148c6adac1dacfc8df71043a679dedbc02e341c6d1b

Observation 09cd76cd-3399-4940-a23a-4e7d26103c76 · outbound

This paper cites In: Proceedings of the IEEE conference on computer vision and pattern recognition.

Video Generation Models are General-Purpose Vision Learners In: Proceedings of the IEEE conference on computer vision and pattern recognition

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-13T00:56:18.867382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:56:18.867382Z digest=sha256:53f5c22192744d80f26321e121647fb363eea48b776e8f4aabccc504f7795886

Observation 0fc70e05-2f21-4466-9ba1-4b76e2ea3728 · outbound

This paper cites In: Proceedings of the Winter Conference on Applications of Computer Vision.

Video Generation Models are General-Purpose Vision Learners In: Proceedings of the Winter Conference on Applications of Computer Vision

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-13T00:56:18.867382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:56:18.867382Z digest=sha256:ef0262ed95278ca80096b0e013a0250487e40b6f341b60490ae5b1c9fc66b9d6

Observation 6e3893fd-ab16-4dc5-a32c-48921804461a · outbound

This paper cites The international journal of robotics research32(11), 1231–1237 (2013) 11.

Video Generation Models are General-Purpose Vision Learners The international journal of robotics research32(11), 1231–1237 (2013) 11

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-13T00:56:18.867382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:56:18.867382Z digest=sha256:5af232b6322af88c5dfc31061970c47811e3815faa975176e458655ce51f6af7

Observation 8753e497-42e3-44ba-ac64-de87a2a702a6 · outbound

This paper cites Lotus-2: Advancing Geometric Dense Prediction with Powerful Image Generative Model.

Video Generation Models are General-Purpose Vision Learners Lotus-2: Advancing Geometric Dense Prediction with Powerful Image Generative Model

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-13T00:56:18.867382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:56:18.867382Z digest=sha256:1a5a50df79edcbe72f8858ff1a79533751cb158393d85839894da01bad3966d8

Observation 41d2635f-c791-44ca-a813-26a28f2f421e · outbound

This paper cites In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition.

Video Generation Models are General-Purpose Vision Learners In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-13T00:56:18.867382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:56:18.867382Z digest=sha256:56336ba7815546cbd6cb6c5498d666e989c5e4382a8184baaabde6970ab509c4

Observation 53012cfd-3327-4c05-98b1-87ac8434e598 · outbound

This paper cites Denoising Diffusion Probabilistic Models.

Video Generation Models are General-Purpose Vision Learners Denoising Diffusion Probabilistic Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-13T00:56:18.867382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:56:18.867382Z digest=sha256:1d92598d80088d80e2cdf5b9eeadeccb7ad949a14496e6160f8da8c739677ee7

Observation a087af19-e580-421e-a627-7d7fe27d2299 · outbound

This paper cites DepthCrafter: Generating Consistent Long Depth Sequences for Open-world Videos.

Video Generation Models are General-Purpose Vision Learners DepthCrafter: Generating Consistent Long Depth Sequences for Open-world Videos

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-13T00:56:18.867382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:56:18.867382Z digest=sha256:bbb6d047d92d4e7fafaa16d635434fb8b1c1e3cc9ae3899f938536b7d94e51d2

Observation 057785f3-a81d-468f-a587-d6bb1def0b18 · outbound

This paper cites arXiv preprint arXiv:2512.07831 (2025) 3.

Video Generation Models are General-Purpose Vision Learners arXiv preprint arXiv:2512.07831 (2025) 3

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-13T00:56:18.867382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:56:18.867382Z digest=sha256:6f1f796cadc34c43ea6d6b2150436a5b81023bd8f54afd8242dff1fc52b39f53

Observation e3f469a3-544d-4b9b-8e1b-8268693b4c01 · outbound

This paper cites In: Proceedings of the IEEE conference on computer vision and pattern recognition.

Video Generation Models are General-Purpose Vision Learners In: Proceedings of the IEEE conference on computer vision and pattern recognition

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-13T00:56:18.867382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:56:18.867382Z digest=sha256:dc63421bf2dbbada2ad0d525012e2e5870027ffd39b08c1604aea1fd0f73c625

Observation 333f737a-c649-4dac-ac89-68a3b759856a · outbound

This paper cites Geo4D: Leveraging Video Generators for Geometric 4D Scene Reconstruction.

Video Generation Models are General-Purpose Vision Learners Geo4D: Leveraging Video Generators for Geometric 4D Scene Reconstruction

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-13T00:56:18.867382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:56:18.867382Z digest=sha256:2432730e8c590a25344437b0989bb1cb21248d538d3102826a07796ffef6fb89

Observation d480ced8-caa3-4ec8-93d6-4721849f8d7d · outbound

This paper cites In: International Conference on Computer Vision (ICCV) (2023) 11.

Video Generation Models are General-Purpose Vision Learners In: International Conference on Computer Vision (ICCV) (2023) 11

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-13T00:56:18.867382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:56:18.867382Z digest=sha256:e93a4545f5e80a5114621bee3bc0bdcd7e4167f144eac591a5971355ee36b164

Observation 3c5efaf4-3f78-42ed-bc77-55623eb30252 · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).

Video Generation Models are General-Purpose Vision Learners In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-13T00:56:18.867382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:56:18.867382Z digest=sha256:9eb04efbd5b21fd6222bcf3182f97085228713f1b57953e6521fd20b6f15c6eb

Observation 7b0d277b-89cb-4cb5-aafa-fd6f3964737c · outbound

This paper cites In: Proceedings of the AAAI Conference on Artificial Intelligence.

Video Generation Models are General-Purpose Vision Learners In: Proceedings of the AAAI Conference on Artificial Intelligence

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-13T00:56:18.867382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:56:18.867382Z digest=sha256:61163860e10a6bf8f5e025699c89a28bf639009f0a0286140b9a688eaada36f9

Observation 523f5417-e211-42ce-aa12-b4f69a0876f6 · outbound

This paper cites MapAnything: Universal Feed-Forward Metric 3D Reconstruction.

Video Generation Models are General-Purpose Vision Learners MapAnything: Universal Feed-Forward Metric 3D Reconstruction

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-13T00:56:18.867382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:56:18.867382Z digest=sha256:59b964fbc109ec76fc9f8bc01f513d8dabee14d5b8433957270693a9b3ba7ae0

Observation b439d091-43b4-4b1c-a8ab-145f3b2767a5 · outbound

This paper cites In: Leonardis, A., Ricci, E., Roth, S., Russakovsky, O., Sattler, T., Varol, G.

Video Generation Models are General-Purpose Vision Learners In: Leonardis, A., Ricci, E., Roth, S., Russakovsky, O., Sattler, T., Varol, G

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-13T00:56:18.867382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:56:18.867382Z digest=sha256:42f2a411c7819e3628c6b45a5bc967483ad7c1ecde104e7fbc1bd43908897643

Observation fb070ba0-8b7a-410f-bad7-51554138b383 · outbound

This paper cites In: Asian conference on computer vision.

Video Generation Models are General-Purpose Vision Learners In: Asian conference on computer vision

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-13T00:56:18.867382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:56:18.867382Z digest=sha256:a64ec528028185c7d382f584641d75370f8f1c4e00e24e1c4bf50d1616ca3673

Observation 385cbbf2-a152-45b0-8942-d9051267dace · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Video Generation Models are General-Purpose Vision Learners Adam: A Method for Stochastic Optimization

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-13T00:56:18.867382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:56:18.867382Z digest=sha256:d83aca5c2269c2c3187deaf07ce8c381bf579a557ba25ce15b414637a8aa25f6

Observation 95294dfb-f786-412a-be1d-53a08a19dbfc · outbound

This paper cites In: Proceedings of the IEEE/CVF international conference on computer vision.

Video Generation Models are General-Purpose Vision Learners In: Proceedings of the IEEE/CVF international conference on computer vision

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-13T00:56:18.867382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:56:18.867382Z digest=sha256:f638e9ee2dab12c7e35f2ebc94548dd8a4da091bc2bb89d699893484fab6a6b4

Observation c46a1f6e-6f13-4b8f-9a6c-98450bc7130f · outbound

This paper cites Buffer Anytime: Zero-Shot Video Depth and Normal from Image Priors.

Video Generation Models are General-Purpose Vision Learners Buffer Anytime: Zero-Shot Video Depth and Normal from Image Priors

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-13T00:56:18.867382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:56:18.867382Z digest=sha256:27de88cc5d976b5da21625708995a2e6676384e1a9c886468f156acefe8e4af9

Observation bde40d8d-3abc-4c88-8dc5-378f0921d422 · outbound

This paper cites GENMO: A GENeralist Model for Human MOtion.

Video Generation Models are General-Purpose Vision Learners GENMO: A GENeralist Model for Human MOtion

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-13T00:56:18.867382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:56:18.867382Z digest=sha256:61b44af194cea0597741c9e4e7ab9fc2476c6e61109273cff3ffbf53b1d2dcc6

Observation bf4c7601-7aa5-4eff-8a20-ac03b233c71f · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.

Video Generation Models are General-Purpose Vision Learners In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-13T00:56:18.867382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:56:18.867382Z digest=sha256:0c3e661cbddbf296077b95bdee02adac60f2e894727e874ce1b694d7e2133b7e

Observation 32a7def1-ae01-4436-8b8f-dde33814c562 · outbound

This paper cites DiffusionRenderer: Neural Inverse and Forward Rendering with Video Diffusion Models.

Video Generation Models are General-Purpose Vision Learners DiffusionRenderer: Neural Inverse and Forward Rendering with Video Diffusion Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-13T00:56:18.867382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:56:18.867382Z digest=sha256:6678b01613a60d316be4078d6d44fe0b70cb8fcde5b5cb7fb2a3b36bb6474d85

Observation 7bbd9e2d-be04-44f0-aebf-89e977fd075e · outbound

This paper cites Depth Anything 3: Recovering the Visual Space from Any Views.

Video Generation Models are General-Purpose Vision Learners Depth Anything 3: Recovering the Visual Space from Any Views

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-13T00:56:18.867382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:56:18.867382Z digest=sha256:135b0acc3631027209973ba6e52dd3761835d8fba0924a5a447aa1b597b7d8e5

Observation eebfe8e9-4eac-4bfe-a9c0-6f119e4799bb · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.

Video Generation Models are General-Purpose Vision Learners In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

Reference 42

Resolution
unresolved
no resolver link, observed 2026-07-13T00:56:18.867382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:56:18.867382Z digest=sha256:29344fb923fbd040fc9d57f329a9d65766464805d1f2ccae1128fcafa9b312f1

Observation 8d62c2fe-2c85-4bb8-9871-77beda147deb · outbound

This paper cites In: Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision.

Video Generation Models are General-Purpose Vision Learners In: Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-13T00:56:18.867382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:56:18.867382Z digest=sha256:4b779db479562c1a820423fb39ca24a25e3d23a1469d8e452307652611735695

Observation e7bad297-4e7c-40d2-9efb-eadc425480b5 · outbound

This paper cites Flow Matching for Generative Modeling.

Video Generation Models are General-Purpose Vision Learners Flow Matching for Generative Modeling

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-13T00:56:18.867382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:56:18.867382Z digest=sha256:c3daea2b23b773a9e469df895095bed55f2c874cc85aea9c6d6e71c13eaeae15

Observation 456fcaeb-d3d6-45ff-874c-88c0179ffcec · outbound

This paper cites Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow.

Video Generation Models are General-Purpose Vision Learners Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow

Reference 45

Resolution
unresolved
no resolver link, observed 2026-07-13T00:56:18.867382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:56:18.867382Z digest=sha256:3a4ac6cce4e3be99c80e688737f425c00d072eafca29f84bcbc46483a1eb7cf4

Observation fb842aeb-cb1a-45a2-9a84-e4a825220287 · outbound

This paper cites an unresolved cited work.

Video Generation Models are General-Purpose Vision Learners Unresolved cited work

Reference 46

Resolution
unresolved
no resolver link, observed 2026-07-13T00:56:18.867382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:56:18.867382Z digest=sha256:35a07e0754f91cc8d3fa47753af70b58d41620ee07084d86065680ab9872ad81

Observation 3b364577-90b7-426f-a329-1dc9f45023b9 · outbound

This paper cites Advances in Neural Information Processing Systems36, 58363– 58408 (2023) 3.

Video Generation Models are General-Purpose Vision Learners Advances in Neural Information Processing Systems36, 58363– 58408 (2023) 3

Reference 47

Resolution
unresolved
no resolver link, observed 2026-07-13T00:56:18.867382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:56:18.867382Z digest=sha256:adf141dad123bc01b624dc1904aa7b669d37bfc12fdd359de8dba81b549161c3

Observation 78301d01-2d02-4a8c-9c85-2d6d4f80a667 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

Video Generation Models are General-Purpose Vision Learners DINOv2: Learning Robust Visual Features without Supervision

Reference 48

Resolution
unresolved
no resolver link, observed 2026-07-13T00:56:18.867382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:56:18.867382Z digest=sha256:bf62afa1180d69e1398668e44ef26c1838862e626e8790c1239fe076bb9a6867

Observation 5e700313-f427-4409-9a04-acf6c6c8c4bc · outbound

This paper cites In: Proceedings of the IEEE conference on computer vision and pattern recognition.

Video Generation Models are General-Purpose Vision Learners In: Proceedings of the IEEE conference on computer vision and pattern recognition

Reference 49

Resolution
unresolved
no resolver link, observed 2026-07-13T00:56:18.867382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:56:18.867382Z digest=sha256:234193aaa641144e8bcf342ec8020a7bcad9e3297a4423bce220f782940c2127

Observation 03288838-d1df-43ef-97a6-0e6d77c840c6 · outbound

This paper cites In: Proceedings of the IEEE conference on computer vision and pattern recognition.

Video Generation Models are General-Purpose Vision Learners In: Proceedings of the IEEE conference on computer vision and pattern recognition

Reference 50

Resolution
unresolved
no resolver link, observed 2026-07-13T00:56:18.867382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:56:18.867382Z digest=sha256:c1f731b050a8f2e64226ef5c0f85136ee05afa0c55618503f2245697636728d1

Observation 232f96fb-ab3a-4361-9c51-d3745d91844a · outbound

This paper cites In: International conference on machine learning.

Video Generation Models are General-Purpose Vision Learners In: International conference on machine learning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-07-13T00:56:18.867382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:56:18.867382Z digest=sha256:cdf5a90cfaa4f14fec0dde4717018ca22150b0fa3e5aff71da6c0c059b6823f2

Observation ea02bec8-aca7-491b-8944-a2b2fc7da32c · outbound

This paper cites IEEE transactions on pattern analysis and machine intelligence44(3), 1623–1637 (2020) 11.

Video Generation Models are General-Purpose Vision Learners IEEE transactions on pattern analysis and machine intelligence44(3), 1623–1637 (2020) 11

Reference 52

Resolution
unresolved
no resolver link, observed 2026-07-13T00:56:18.867382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:56:18.867382Z digest=sha256:e3d69b82ff63b2503cd6b730e57055763e478e2f0fdc6c78e9f1c3c6fe2315eb

Observation 67a694d8-c7bf-4a49-8f59-071f671d7504 · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

Video Generation Models are General-Purpose Vision Learners SAM 2: Segment Anything in Images and Videos

Reference 53

Resolution
unresolved
no resolver link, observed 2026-07-13T00:56:18.867382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:56:18.867382Z digest=sha256:28d86ed6d223cd1fb0c3f57ba4cbaba20af06ba17ef95a727c6a3d1baeb6d2f3

Observation 89f9f7dc-568f-4858-97fe-5d42148eb97f · outbound

This paper cites an unresolved cited work.

Video Generation Models are General-Purpose Vision Learners Unresolved cited work

Reference 54

Resolution
unresolved
no resolver link, observed 2026-07-13T00:56:18.867382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:56:18.867382Z digest=sha256:4388dfe09b87664afde35b9c22a169c42d16601fa27764188f93c809bc7ab068

Observation 15e48454-ed40-417f-8ffd-07c8be949f4a · outbound

This paper cites In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition (CVPR).

Video Generation Models are General-Purpose Vision Learners In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition (CVPR)

Reference 55

Resolution
unresolved
no resolver link, observed 2026-07-13T00:56:18.867382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:56:18.867382Z digest=sha256:8f64546c9d306823feb8741edeba578a672cebd0f6f6a88b2e6949c75840777c

Observation df15b44c-68d9-4506-b5ad-15320431b80f · outbound

This paper cites DAViD: Data-efficient and Accurate Vision Models from Synthetic Data.

Video Generation Models are General-Purpose Vision Learners DAViD: Data-efficient and Accurate Vision Models from Synthetic Data

Reference 56

Resolution
unresolved
no resolver link, observed 2026-07-13T00:56:18.867382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:56:18.867382Z digest=sha256:589e42009a1f596d48820fcc55595f323d373f2f598df68c2372620f413adcc5

Observation e25e0bce-44bf-4120-ad16-0be6408ef6ed · outbound

This paper cites In: Proceedings of the IEEE conference on computer vision and pattern recognition.

Video Generation Models are General-Purpose Vision Learners In: Proceedings of the IEEE conference on computer vision and pattern recognition

Reference 57

Resolution
unresolved
no resolver link, observed 2026-07-13T00:56:18.867382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:56:18.867382Z digest=sha256:a64718d1d9e6d24a54b695dc29528de8a9f61c57bbb5870189fb6b820613104e

Observation e678c719-2113-4e23-b32b-80e5536b1b98 · outbound

This paper cites In: SIGGRAPH Asia Conference Proceedings (2024) 11 17 Video Generation Models are General-Purpose Vision Learners.

Video Generation Models are General-Purpose Vision Learners In: SIGGRAPH Asia Conference Proceedings (2024) 11 17 Video Generation Models are General-Purpose Vision Learners

Reference 58

Resolution
unresolved
no resolver link, observed 2026-07-13T00:56:18.867382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:56:18.867382Z digest=sha256:dcce50f721f121d3c652c4ca0b33befb116118aa1a05080b051cfac6771849a7

Observation 8cd14e0c-6a0a-4cc1-9aab-c08e05334314 · outbound

This paper cites DINOv3.

Video Generation Models are General-Purpose Vision Learners DINOv3

Reference 59

Resolution
unresolved
no resolver link, observed 2026-07-13T00:56:18.867382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:56:18.867382Z digest=sha256:85cc006f74aa2eb48cebffa600a2a56748215fcb7fee29bdfb091c42d0d3f4f1

Observation f8b18b41-b1cc-42e2-ba7b-451f21b156ae · outbound

This paper cites In: International Conference on Learning Representations (2021) 6.

Video Generation Models are General-Purpose Vision Learners In: International Conference on Learning Representations (2021) 6

Reference 60

Resolution
unresolved
no resolver link, observed 2026-07-13T00:56:18.867382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:56:18.867382Z digest=sha256:9408bd5e213ba715a1d9c40f9b8aa94e7ca64d45e88d4fd28201b992d2a6e231

Observation b669bf86-1a50-441e-87a1-ad75bf5d629b · outbound

This paper cites Advances in neural information processing systems35, 10078–10093 (2022) 2, 4, 13.

Video Generation Models are General-Purpose Vision Learners Advances in neural information processing systems35, 10078–10093 (2022) 2, 4, 13

Reference 61

Resolution
unresolved
no resolver link, observed 2026-07-13T00:56:18.867382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:56:18.867382Z digest=sha256:e02da1a645786030ba2f4ebf1683d0b577229e3c1f25a1fa658573d651d4d61c

Observation c1f47cb2-a1bc-4129-a261-406d0b614598 · outbound

This paper cites VoCap: Video Object Captioning and Segmentation from Any Prompt.

Video Generation Models are General-Purpose Vision Learners VoCap: Video Object Captioning and Segmentation from Any Prompt

Reference 62

Resolution
unresolved
no resolver link, observed 2026-07-13T00:56:18.867382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:56:18.867382Z digest=sha256:ef3f6d166fe999c6159e272004dfa688fde1801bdebab18987eeae48fe6b9daf

Observation eb54ce5b-fe02-46f3-8dd0-4c3aa3b1d3d0 · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

Video Generation Models are General-Purpose Vision Learners Wan: Open and Advanced Large-Scale Video Generative Models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-07-13T00:56:18.867382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:56:18.867382Z digest=sha256:1212319c59ad5c776b1fd41905132e295602584962c1dd0817fdbfdc6640a117

Observation bd76484d-79ae-41ca-947e-aa48fdfb73c4 · outbound

This paper cites In: Proceedings of the Computer Vision and Pattern Recognition Conference.

Video Generation Models are General-Purpose Vision Learners In: Proceedings of the Computer Vision and Pattern Recognition Conference

Reference 64

Resolution
unresolved
no resolver link, observed 2026-07-13T00:56:18.867382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:56:18.867382Z digest=sha256:ab35b0da98119e89762035b56ad4a10610b2826aac056b96dd01b67f47b1f653

Observation 5794ff4b-4032-4b20-b929-b5586f1a5e22 · outbound

This paper cites VGGT-$\Omega$.

Video Generation Models are General-Purpose Vision Learners VGGT-$\Omega$

Reference 65

Resolution
unresolved
no resolver link, observed 2026-07-13T00:56:18.867382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:56:18.867382Z digest=sha256:76dbc3f27937d1dad0d138380db0908243bc7bfe94f0ff7ae51c6904731de19e

Observation a31fc39c-0084-46dd-a4a3-6385f685aa8f · outbound

This paper cites arXiv preprintarXiv:2603.25892(2026) 4.

Video Generation Models are General-Purpose Vision Learners arXiv preprintarXiv:2603.25892(2026) 4

Reference 66

Resolution
unresolved
no resolver link, observed 2026-07-13T00:56:18.867382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:56:18.867382Z digest=sha256:68082c4719a61bbd13c16f0908a62f90585d3f3ec52e90f719a7976bcf6e9bc5

Observation 762ef6d4-e269-438d-9be9-4bdce722da14 · outbound

This paper cites In: 2020 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS).

Video Generation Models are General-Purpose Vision Learners In: 2020 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)

Reference 67

Resolution
unresolved
no resolver link, observed 2026-07-13T00:56:18.867382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:56:18.867382Z digest=sha256:6d5f2a03d2d19b4242a7c648ced021e9bf8f9a34094bc4afdb96f3861851c1db

Observation f0d3bbaf-37ca-47cc-8cf3-73c5a1897eff · outbound

This paper cites In: European Conference on Computer Vision.

Video Generation Models are General-Purpose Vision Learners In: European Conference on Computer Vision

Reference 68

Resolution
unresolved
no resolver link, observed 2026-07-13T00:56:18.867382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:56:18.867382Z digest=sha256:bb155483d6b4acf2944790b0f0cbb1ee4dd3e173c924d9d03ce42ac9cf440d49

Observation 8e37c9d8-47b1-4b2e-aff0-3e4f726cd7f5 · outbound

This paper cites Video models are zero-shot learners and reasoners.

Video Generation Models are General-Purpose Vision Learners Video models are zero-shot learners and reasoners

Reference 69

Resolution
unresolved
no resolver link, observed 2026-07-13T00:56:18.867382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:56:18.867382Z digest=sha256:2e1a474fe55a19b983669fbcc369f664de065a6d5e600037eaedb513efa857e0

Observation 18590aad-f3d9-4a8f-a80d-ac8302e3f676 · outbound

This paper cites In: CVPR (2022) 11.

Video Generation Models are General-Purpose Vision Learners In: CVPR (2022) 11

Reference 70

Resolution
unresolved
no resolver link, observed 2026-07-13T00:56:18.867382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:56:18.867382Z digest=sha256:62c9b78567eeda34c9cbaad740c8e049ccfc2bbe424b62d20afcf218048840a5

Observation 1c9d9c9a-dc14-4c5b-ba19-d24d263be397 · outbound

This paper cites What Matters When Repurposing Diffusion Models for General Dense Perception Tasks?.

Video Generation Models are General-Purpose Vision Learners What Matters When Repurposing Diffusion Models for General Dense Perception Tasks?

Reference 71

Resolution
unresolved
no resolver link, observed 2026-07-13T00:56:18.867382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:56:18.867382Z digest=sha256:947e19f2178e46f31e24f5c39d397c21eb0befa91db6b3c13ddb37f1706e0370

Observation e11c055e-67c3-43c1-a361-996caa4aa99a · outbound

This paper cites In: Proceedings of the European conference on computer vision (ECCV).

Video Generation Models are General-Purpose Vision Learners In: Proceedings of the European conference on computer vision (ECCV)

Reference 72

Resolution
unresolved
no resolver link, observed 2026-07-13T00:56:18.867382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:56:18.867382Z digest=sha256:58afbd1d8071ac96d1f49ec36715623d4b9f51711087c41702b53f28090e916a

Observation d9d5893f-6750-47de-b070-aedeb160f5a5 · outbound

This paper cites In: European Conference on Computer Vision.

Video Generation Models are General-Purpose Vision Learners In: European Conference on Computer Vision

Reference 73

Resolution
unresolved
no resolver link, observed 2026-07-13T00:56:18.867382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:56:18.867382Z digest=sha256:06d4441037af98a97a6600c58bf55a1c0c59e26135d841ffd4979280df0a79db

Observation 1139947d-1420-4fa3-be16-9bd4d00c3b98 · outbound

This paper cites Depth Any Video with Scalable Synthetic Data.

Video Generation Models are General-Purpose Vision Learners Depth Any Video with Scalable Synthetic Data

Reference 74

Resolution
unresolved
no resolver link, observed 2026-07-13T00:56:18.867382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:56:18.867382Z digest=sha256:2e8df9a10d9088d217c381180d9633baa93f5ba7dea3a71784310043ee6bc656

Observation 91d0549b-5d32-4294-b52c-e9195afdbe1a · outbound

This paper cites In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition.

Video Generation Models are General-Purpose Vision Learners In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition

Reference 75

Resolution
unresolved
no resolver link, observed 2026-07-13T00:56:18.867382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:56:18.867382Z digest=sha256:1624a403f70c6b8892dc97026314ae4401d5a79f9cacb3140f2cd51b5e27934f

Observation 1082eb65-55ef-45ef-8347-c9b1cd780b8c · outbound

This paper cites Advances in Neural Information Processing Systems37, 21875–21911 (2024) 1, 3, 11.

Video Generation Models are General-Purpose Vision Learners Advances in Neural Information Processing Systems37, 21875–21911 (2024) 1, 3, 11

Reference 76

Resolution
unresolved
no resolver link, observed 2026-07-13T00:56:18.867382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:56:18.867382Z digest=sha256:1c9c356060a7cc5aeba86d13d88b2cfa3a9bc629a355df52586d4072b7c89014

Observation bb32f305-3347-466c-9fab-babe1e83292a · outbound

This paper cites Advances in Neural Information Processing Systems37, 21875–21911 (2025) 1.

Video Generation Models are General-Purpose Vision Learners Advances in Neural Information Processing Systems37, 21875–21911 (2025) 1

Reference 77

Resolution
unresolved
no resolver link, observed 2026-07-13T00:56:18.867382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:56:18.867382Z digest=sha256:8bcbac1f0f2a8b66e2abd12bae56cb225b2f37f13a317ed32864cfea7879c4af

Observation 006aa9f8-098c-44f4-9e24-fd8a1058790f · outbound

This paper cites In: Proceedings of the Computer Vision and Pattern Recognition Conference.

Video Generation Models are General-Purpose Vision Learners In: Proceedings of the Computer Vision and Pattern Recognition Conference

Reference 78

Resolution
unresolved
no resolver link, observed 2026-07-13T00:56:18.867382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:56:18.867382Z digest=sha256:9b75ef4c229bcb02b494cfdb23a11d668ff934cff9bd0248474ae54d2cd8eca6

Observation 87293276-b681-4502-a85a-d5d4f51dd077 · outbound

This paper cites In: Computer Vision and Pattern Recognition (CVPR) (2023) 11.

Video Generation Models are General-Purpose Vision Learners In: Computer Vision and Pattern Recognition (CVPR) (2023) 11

Reference 79

Resolution
unresolved
no resolver link, observed 2026-07-13T00:56:18.867382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:56:18.867382Z digest=sha256:007bd11a92a85a672c1f5039918d66c202c3ff96d018d1eb376ae2c93f6749c8

Observation c2735f94-4133-444f-b54d-defe1915a39e · outbound

This paper cites In: European conference on computer vision.

Video Generation Models are General-Purpose Vision Learners In: European conference on computer vision

Reference 80

Resolution
unresolved
no resolver link, observed 2026-07-13T00:56:18.867382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:56:18.867382Z digest=sha256:9351d9272820c858097194636d0b0191786cca7bf44c618c48e601ed0d757f19

Observation c6162dab-e909-428d-a761-e3e0a8f803fc · outbound

This paper cites In: Proceedings of the IEEE/CVF international conference on computer vision.

Video Generation Models are General-Purpose Vision Learners In: Proceedings of the IEEE/CVF international conference on computer vision

Reference 81

Resolution
unresolved
no resolver link, observed 2026-07-13T00:56:18.867382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:56:18.867382Z digest=sha256:7b2d546cd3ec6d5ce000abb9a90f0a356e379a1473932b30dae248130d8aaf6a

Observation c1ca92e8-6a77-4db2-80a1-46700d7385dd · outbound

This paper cites arXiv preprint arXiv:2512.08924 (2025) 2, 3, 11, 12.

Video Generation Models are General-Purpose Vision Learners arXiv preprint arXiv:2512.08924 (2025) 2, 3, 11, 12

Reference 82

Resolution
unresolved
no resolver link, observed 2026-07-13T00:56:18.867382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:56:18.867382Z digest=sha256:4ee8bc662dc61de2369fa17ef7b2a0677dc451c4daab2dc3c0e8c23bb3cc0829

Observation f9d16dc7-4545-4b4b-9fbe-64a242094f62 · outbound

This paper cites 23345–23366 (2024) 7.

Video Generation Models are General-Purpose Vision Learners 23345–23366 (2024) 7

Reference 83

Resolution
unresolved
no resolver link, observed 2026-07-13T00:56:18.867382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:56:18.867382Z digest=sha256:3e2ed04326d568941f0c4cd296f00ab067944ae40a2b8a01f891c72bbec4b5f8

Observation 0975dc7b-8ee3-47f5-8802-b598c8b84384 · outbound

This paper cites arXiv preprint arXiv:2502.17157 (2025) 4, 11.

Video Generation Models are General-Purpose Vision Learners arXiv preprint arXiv:2502.17157 (2025) 4, 11

Reference 84

Resolution
unresolved
no resolver link, observed 2026-07-13T00:56:18.867382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:56:18.867382Z digest=sha256:244df4cf2f572ed8a9c458ee2b1a6a460a1100a7a44992ab8a7c7c734869e8f6

Observation 9ade3979-a485-4960-abbc-f9f222f3be1d · outbound

This paper cites In: Proceedings of the Computer Vision and Pattern Recognition Conference.

Video Generation Models are General-Purpose Vision Learners In: Proceedings of the Computer Vision and Pattern Recognition Conference

Reference 85

Resolution
unresolved
no resolver link, observed 2026-07-13T00:56:18.867382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:56:18.867382Z digest=sha256:b2314469985b97b507b81be5d388da462cbdbbd3940bc070c08803da48ae1f8c

Observation e6b6b15f-52d8-4796-8ae2-63049ff1a72e · outbound

This paper cites In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition.

Video Generation Models are General-Purpose Vision Learners In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition

Reference 86

Resolution
unresolved
no resolver link, observed 2026-07-13T00:56:18.867382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:56:18.867382Z digest=sha256:d1d17026987b4102895c123a65e3a78a735631917b23cd9f56c9d42a27196c61

Observation fee7102b-d4d0-49e2-80ae-591a188e3128 · outbound

This paper cites Recurrent Video Masked Autoencoders.

Video Generation Models are General-Purpose Vision Learners Recurrent Video Masked Autoencoders

Reference 87

Resolution
unresolved
no resolver link, observed 2026-07-13T00:56:18.867382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:56:18.867382Z digest=sha256:c562daeb439a794e3916392932e6537cf0a3c5542db5194d758a4f80ddd9681f

Pith citing papers

No inbound Pith citation observations are available.