Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T11:36:44.784379Z
Paper Citation Record · LEDGER
As of 23 August 2026, this Paper Citation Record lists 83 of 83 outbound references and 13 inbound Pith citation observations for arXiv:2412.15212.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T11:36:44.784379Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T14:58:05.471418Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-10T00:26:39.628379Z
83 of 83 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 0b339532-bb5e-4af6-a620-76975105b77b · outbound
Scaling 4D Representations Learning to see by moving
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf470e0f-5277-4e63-8bb8-17d39113ebec · outbound
Scaling 4D Representations Flamingo: a visual language model for few-shot learning.NeurIPS, 2022
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 662649da-014c-44dc-9414-82ada87ce246 · outbound
Scaling 4D Representations Self-supervised learning by cross-modal audio-video clustering
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c10bbcd5-9c71-4960-8f22-2645104fbdd0 · outbound
Scaling 4D Representations Look, listen and learn
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c658eb3e-2ea4-46b3-a254-1705fe88234c · outbound
Scaling 4D Representations Objects that sound
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 16179761-9a9e-4210-a727-73d0c9be023b · outbound
Scaling 4D Representations Vivit: A video vi- sion transformer
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 7eddce4d-481e-42b1-8e33-4682642feebd · outbound
Scaling 4D Representations Yuille, Trevor Darrell, Jitendra Malik, and Alexei A
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation a0c4a0cc-5f7d-425e-a74c-82815ff07527 · outbound
Scaling 4D Representations Revisiting feature prediction for learning visual rep- resentations from video
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 2ca458b1-eaac-4e82-8075-37dd30dac47c · outbound
Scaling 4D Representations Physion: Evaluating physical prediction from vision in humans and machines
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 952ee2be-95a3-42b3-a3b1-2f1536bec186 · outbound
Scaling 4D Representations Zoedepth: Zero-shot transfer by com- bining relative and metric depth
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 470119a5-a728-47d2-b01a-33f65d853cf0 · outbound
Scaling 4D Representations Deep regression on manifolds: a 3d rotation case study
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation e7ab0d43-18e1-48cd-a0a1-431000193af7 · outbound
Scaling 4D Representations Activitynet: A large-scale video benchmark for human activity understanding
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e27a8fe5-b26e-4900-81b2-85849eaa52e1 · outbound
Scaling 4D Representations Generative pre- training from pixels
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 78290ec3-c7e5-4978-a6ff-79e2078fb82e · outbound
Scaling 4D Representations A simple framework for contrastive learning of visual representations
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 1b2718de-c323-4f1a-954f-b0760812c3d6 · outbound
Scaling 4D Representations PaLI-3 vision language models: Smaller, faster, stronger
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation f7725220-fe7d-4be8-96c1-e8f4e56e71f3 · outbound
Scaling 4D Representations A unified architecture for natural language processing: Deep neural networks with multitask learning
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation ebcf9b9b-d7fa-4559-a9a8-482ceb41f14d · outbound
Scaling 4D Representations Scannet: Richly-annotated 3d reconstructions of indoor scenes
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c7a9af97-80d3-422b-9df6-2facdef4b3c2 · outbound
Scaling 4D Representations Scaling egocentric vision: The epic-kitchens dataset
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 50e07444-1252-4e4e-9661-0d83d8d4a623 · outbound
Scaling 4D Representations Scaling vision transformers to 22 billion pa- rameters
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 4b68ec81-0a88-4fc0-8921-ee4363d57d59 · outbound
Scaling 4D Representations Unsuper- vised visual representation learning by context prediction
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f9c645e-3faf-4d17-a87a-a946352dbe79 · outbound
Scaling 4D Representations TAP-vid: A bench- mark for tracking any point in a video
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 01d6782e-17ff-437c-840f-d051952c183f · outbound
Scaling 4D Representations TAPIR: Tracking any point with per-frame initialization and temporal refinement
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 980b072f-a4a6-4b6e-85d1-10c0b558138a · outbound
Scaling 4D Representations Boot- sTAP: Bootstrapped training for tracking any point
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation d85b95bb-494b-4c9b-a766-2e1751f1d87b · outbound
Scaling 4D Representations An image is worth 16x16 words: Transformers for image recognition at scale
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 8149e4d6-62b2-46ba-90aa-2aee56ab9d2c · outbound
Scaling 4D Representations Prob- ing the 3d awareness of visual foundation models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 11b57ec8-2b55-496c-81c8-ad1ab73851a7 · outbound
Scaling 4D Representations Scalable pre- training of large autoregressive image models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation cbdcdae6-b054-4dca-82d1-2f5c2517b874 · outbound
Scaling 4D Representations SA Vi++: Towards end-to-end object-centric learning from real-world videos
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation a6f3db81-a5af-404c-9abb-4bffa133347b · outbound
Scaling 4D Representations Spatiotemporal residual networks for video action recogni- tion
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 1cc776fa-20ed-484d-a36b-7c45bbcad05f · outbound
Scaling 4D Representations Slowfast networks for video recognition
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b951d335-bee8-4660-b415-b4ad64cfb981 · outbound
Scaling 4D Representations A large-scale study on unsupervised spatiotemporal representation learning
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 4ac01222-34ab-43e4-96ec-48e8fc11e549 · outbound
Scaling 4D Representations Distributed hier- archical processing in the primate cerebral cortex
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 8529e2cd-b473-4fe8-a8ec-f95a577ef7a0 · outbound
Scaling 4D Representations Learning invariance from transformation se- quences
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 5f749498-3321-4791-a6b3-8fc2c4a6662c · outbound
Scaling 4D Representations Unresolved cited work
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation de41f93e-0a26-4c5d-88fe-6dc04e181cce · outbound
Scaling 4D Representations Learn- ing to linearize under uncertainty
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 6a837ddb-613d-4c10-981b-8d4a1751f540 · outbound
Scaling 4D Representations The” something something” video database for learning and evaluating visual common sense
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 944b4d25-d168-4266-9f56-f4e0face1b11 · outbound
Scaling 4D Representations Unresolved cited work
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 8b952943-1a6a-43ad-b817-66005dd3982e · outbound
Scaling 4D Representations Memory- augmented dense predictive coding for video representation learning
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 264c93e1-577b-42c3-9099-229716f64b24 · outbound
Scaling 4D Representations Self- supervised co-training for video representation learning
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation cd76d221-6c15-49cd-88f3-65c09b1f0f85 · outbound
Scaling 4D Representations Masked autoencoders are scalable vision learners
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 43151a63-3d22-4585-bb0a-0e4fba8352f4 · outbound
Scaling 4D Representations Data-efficient image recognition with contrastive predictive coding
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 1600f371-b1eb-4762-a851-64da1dfddb98 · outbound
Scaling 4D Representations Representation learn- ing with video deep infomax
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 9a4e9834-cdee-46a2-af31-099304396f39 · outbound
Scaling 4D Representations The kinetics human action video dataset, 2017
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 336990ec-141c-4e48-8f93-b07cf87d402f · outbound
Scaling 4D Representations Condi- tional object-centric learning from video
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation ffd50336-2e97-4534-9526-22ee3a190c0e · outbound
Scaling 4D Representations Coopera- tive learning of audio and video models from self-supervised synchronization
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 9b3a51b1-a16d-4ff7-9196-c70c1e81d226 · outbound
Scaling 4D Representations Decoupled weight decay regularization
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7118bcf7-3c7e-4615-bfc4-265bf3cd8f56 · outbound
Scaling 4D Representations Object vision and spatial vision: two cortical path- ways
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 043fbb96-2f40-4c0c-b452-658bcc1ea143 · outbound
Scaling 4D Representations Deep learning from temporal coherence in video
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation fae716d6-0ba1-49bf-9880-919e8337cdfc · outbound
Scaling 4D Representations Audio- visual instance discrimination with cross-modal agreement
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation b74c613d-5ea4-49af-b0d5-fb8aa63ba36c · outbound
Scaling 4D Representations Atlas: End- to-end 3d scene reconstruction from posed images
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1a85c6a-e7e0-4a0b-8d3d-ac99c50bfb22 · outbound
Scaling 4D Representations Unresolved cited work
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c3d11847-36cb-4cf0-b3ba-cc83e4ae29a5 · outbound
Scaling 4D Representations Fully sharded data parallel: faster ai training with fewer gpus
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 6094bae1-83ff-4d21-91e8-ce3e0944b628 · outbound
Scaling 4D Representations Audio-visual scene analysis with self-supervised multisensory features
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 73e05ce9-cdde-4a70-9aff-550ffe691b5a · outbound
Scaling 4D Representations Context encoders: Feature learning by inpainting
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 92e1b67d-3531-45b7-af13-56a231503c45 · outbound
Scaling 4D Representations Learning features by watching ob- jects move
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 65957157-1fdd-4a39-a0fd-c759560f712c · outbound
Scaling 4D Representations Perception test: A diagnostic benchmark for mul- timodal video models
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation e6b5101f-2a7a-464c-b2c9-8edb3d41da86 · outbound
Scaling 4D Representations Asano, Ruth Fong, Jo ˜ao F
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 96ae4680-7f6b-4366-9b2d-4559f66f0d26 · outbound
Scaling 4D Representations Spatiotempo- ral contrastive video representation learning
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 28b8b729-232a-4b07-8567-cf22db7bea2b · outbound
Scaling 4D Representations Improving language understanding by generative pre-training
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation a69fee24-1b84-4747-af8c-6f155bb275cc · outbound
Scaling 4D Representations Learning transferable visual models from natural language supervision, 2021
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 514204a2-b968-4065-a126-d4023f49adf3 · outbound
Scaling 4D Representations Vi- sion transformers for dense prediction
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ead10f5-7ab3-4fe3-86ed-aeb7edd1b75f · outbound
Scaling 4D Representations Video (lan- guage) modeling: a baseline for generative models of natural videos
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 486ee18c-3a64-4bcc-89aa-9fb47a615e7f · outbound
Scaling 4D Representations Broaden your views for self-supervised video learning
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 26dd5823-a6e0-4fa5-87b1-43a83b349148 · outbound
Scaling 4D Representations Learning to localize sound source in visual scenes
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 4fcba358-e2cb-44ac-9b09-72b27f264076 · outbound
Scaling 4D Representations Two-stream con- volutional networks for action recognition in videos
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 8fe5f959-0955-4d4d-90ad-2085c89c9540 · outbound
Scaling 4D Representations A Short Note on the Kinetics-700-2020 Human Action Dataset
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac78d9a9-ec84-435d-9575-4d8af7d2e6de · outbound
Scaling 4D Representations Scalability in perception for autonomous driving: Waymo open dataset
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 2616d095-999d-4c51-a0ef-5ea8d62de70e · outbound
Scaling 4D Representations Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 39d41e55-08ef-43df-9ac7-34f1a761a08f · outbound
Scaling 4D Representations RealEstate10K
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation fd0f2790-0b42-4008-80da-97b19df9f25f · outbound
Scaling 4D Representations Hud- son, Thomas Albert Keck, Joao Carreira, Alexey Dosovit- skiy, Mehdi S
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 5d725e0d-58d4-492f-93fc-3c7d05034ef2 · outbound
Scaling 4D Representations An- ticipating the future by watching unlabeled video
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation a828f910-15b3-46a5-9d5b-589b1b30b86c · outbound
Scaling 4D Representations Videomae v2: Scaling video masked autoencoders with dual masking
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c361c815-92a8-4ac2-89f9-9204a67a5105 · outbound
Scaling 4D Representations Masked video distillation: Rethinking masked feature mod- eling for self-supervised video representation learning
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation f575aaac-9dc3-426e-9772-03bd552b2705 · outbound
Scaling 4D Representations Dust3r: Geometric 3d vi- sion made easy
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation b4072c0e-7f1c-4142-9901-00ed740a4c4e · outbound
Scaling 4D Representations Unsupervised learning of visual representations using videos
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b443d20-db6f-4640-aa81-a1a7d4a29a83 · outbound
Scaling 4D Representations Internvideo: General video foundation models via generative and discriminative learning
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 1e6ea0ac-238f-432b-8a31-1b7cad8c415e · outbound
Scaling 4D Representations Less is more: Consistent video depth estimation with masked frames modeling
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 830813fd-2280-4028-a09f-52dcb43d744d · outbound
Scaling 4D Representations Controlling space and time with dif- fusion models
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 75cdb228-4dec-4e78-9775-9b3c396b71e3 · outbound
Scaling 4D Representations Slow feature analysis: Unsupervised learning of invariances
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c514da1a-3ac7-45eb-9771-de72d24d97c7 · outbound
Scaling 4D Representations Depth anything: Unleashing the power of large-scale unlabeled data
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 5985c939-6421-429b-acfc-fe1d587550ef · outbound
Scaling 4D Representations Sigmoid loss for language image pre-training
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 12071b3b-2248-4965-9ac5-f858106e9b2a · outbound
Scaling 4D Representations A general protocol to probe large vision models for 3d physical understanding
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 26a37d42-9f15-4813-9b9d-6a0050b9e5b6 · outbound
Scaling 4D Representations Videoprism: A foundational visual encoder for video understanding
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation f2a9b05a-c703-41e1-a744-00cf8fe2d95d · outbound
Scaling 4D Representations Taking something from somewhere
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 647ea506-fb21-412f-8709-191a86556497 · inbound
From Image to Video: An Empirical Study of Diffusion Representations Scaling 4D Representations
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6cadd0b9-0116-42a5-a924-64ae95092733 · inbound
MoSiC: Optimal-Transport Motion Trajectory for Dense Self-Supervised Learning Scaling 4D Representations
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f95764d-646c-414b-9d2c-358331e7599a · inbound
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning Scaling 4D Representations
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation cd2fcdce-93e6-443e-a1d2-e60e744413ca · inbound
SciVid: Cross-Domain Evaluation of Video Models in Scientific Applications Scaling 4D Representations
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 987e4f4f-0610-45b9-a150-53cfdb66061e · inbound
Frozen Forecasting: A Unified Evaluation Scaling 4D Representations
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation b38649bd-72f1-4353-8d6a-7c3454b70b61 · inbound
Unique Lives, Shared World: Learning from Single-Life Videos Scaling 4D Representations
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17bdcb93-8b3e-47f5-953b-7f19a69424b2 · inbound
LA-Pose: Latent Action Pretraining Meets Pose Estimation Scaling 4D Representations
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 7f655eea-b5fc-4541-8a3f-5952a9598b38 · inbound
Towards Data-Efficient Video Pre-training with Frozen Image Foundation Models Scaling 4D Representations
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 394c74aa-a514-4f79-96b3-0293b0ebf52c · inbound
Where Do We (Not) Need Temporal Context in Low-Resource Video Task Adaptation? Scaling 4D Representations
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation a3458577-a4a7-4f4c-a660-a63981166db6 · inbound
Gen4U: Unifying Video Generation and Understanding via Diffusion Scaling 4D Representations
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 50b5a864-d3ae-4dc9-a564-937d49a0c200 · inbound
SeeSE3: Emergence of 3D Space in Vision Features Scaling 4D Representations
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b3a923d-c9c3-4986-9f24-86606157a485 · inbound
Self-Supervised Learning of Structured Dynamics from Videos Scaling 4D Representations
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 623c607f-1b6e-4707-bb06-c685e28ce65f · inbound
A Controlled Study of Self-Supervised Image and Video Pretraining under Limited Resources Scaling 4D Representations
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.