Pith. sign in

Paper Citation Record · LEDGER

Self-Supervised Learning of Structured Dynamics from Videos

As of 17 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 0 inbound Pith citation observations for arXiv:2607.21576.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.21576 v1

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T07:07:44.009137Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

47 of 47 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved46
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fcf5aec1-d427-4934-92b1-88dbb6c38a08 · outbound

This paper cites V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning.

Self-Supervised Learning of Structured Dynamics from Videos V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:42.700451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:42.700451Z digest=sha256:ca9d6ab0d4893fc5f100d33f34a2e0c4cd0ba9a08a0528c9754d55d19ce87d23

Observation 645df39c-c401-4de0-81ef-3b53daab1ac7 · outbound

This paper cites Back to the Features: DINO as a Foundation for Video World Models.

Self-Supervised Learning of Structured Dynamics from Videos Back to the Features: DINO as a Foundation for Video World Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:42.787898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:42.787898Z digest=sha256:7fc7bf200e8678547172f167d3ad48444a0171c045eaf3868be340e9fa65d30f

Observation e57671f3-34fe-4506-9642-2ad1495ef298 · outbound

This paper cites Revisiting feature prediction for learning visual representations from video.Transactions on Machine Learning Research, 2024.

Self-Supervised Learning of Structured Dynamics from Videos Revisiting feature prediction for learning visual representations from video.Transactions on Machine Learning Research, 2024

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:42.875142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:42.875142Z digest=sha256:8c5e438448677be2177b73e38b2814dfc4be34377d529db8610414e506b1b3d8

Observation ff4ba9dd-6950-4f73-9f34-ac9508a804ce · outbound

This paper cites VFMF: World modeling by forecasting vision foundation model features.arXiv preprint arXiv:2512.11225,.

Self-Supervised Learning of Structured Dynamics from Videos VFMF: World modeling by forecasting vision foundation model features.arXiv preprint arXiv:2512.11225,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:42.965780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:42.965780Z digest=sha256:ae0c36e2cdff365b8901c1a59013d4d432eea2a478244ebcf429c22038a3c982

Observation de75f00c-62f0-4962-9043-dd50da99aee3 · outbound

This paper cites TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment.

Self-Supervised Learning of Structured Dynamics from Videos TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:43.056117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:43.056117Z digest=sha256:e1131b822c24a130a595803cd89de9c60a4918b2bb3c837f932110b0bc28bdba

Observation e902e301-623b-4c03-9de9-fc397afdd936 · outbound

This paper cites A Short Note on the Kinetics-700 Human Action Dataset.

Self-Supervised Learning of Structured Dynamics from Videos A Short Note on the Kinetics-700 Human Action Dataset

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:43.147125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:43.147125Z digest=sha256:2b975600d7a7219d08dad6d2078ede7fca50c48c062bd1d6cddc45c2cfb36068

Observation 9b3a923d-c9c3-4986-9f24-86606157a485 · outbound

This paper cites Scaling 4D Representations.

Self-Supervised Learning of Structured Dynamics from Videos Scaling 4D Representations

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:43.307749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:43.307749Z digest=sha256:0ba758ba3a5c7b86161e78fbef8ef57d8c6dab1d42fd26043285df74db04d95d

Observation fe9778c6-5fa5-4553-aa9b-0415daf664ba · outbound

This paper cites Vision transformers need registers.

Self-Supervised Learning of Structured Dynamics from Videos Vision transformers need registers

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:43.455509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:43.455509Z digest=sha256:563861a2f3f8603867332eddc2821fab8437451962034c127d4a6ac412d70062

Observation 672b183b-0146-4b90-b179-08ac7841accb · outbound

This paper cites Probing the 3d awareness of visual foundation models.

Self-Supervised Learning of Structured Dynamics from Videos Probing the 3d awareness of visual foundation models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:43.538067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:43.538067Z digest=sha256:51e77a65594c3c92b7925d7ab397216b57f40891d1199b05be1fedf55ce7ed22

Observation 8c1e3832-5293-4856-a12b-ae26f0fa47d4 · outbound

This paper cites Scalable pre-training of large autoregressive image models.

Self-Supervised Learning of Structured Dynamics from Videos Scalable pre-training of large autoregressive image models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:43.612958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:43.612958Z digest=sha256:778138a8a2e61048ba47369c1496fe8ef73023e4ce076049d9c68c1c29f36af8

Observation 4f62d957-313a-43c0-bc5f-819ab38c4242 · outbound

This paper cites Multimodal autoregressive pre-training of large vision encoders.

Self-Supervised Learning of Structured Dynamics from Videos Multimodal autoregressive pre-training of large vision encoders

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:43.715885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:43.715885Z digest=sha256:4b7ec98e1c927a0b890e2340692c57d485ea099c74a879c1e8bf372355103933

Observation 594d648b-9a2c-4095-a313-23e54a26b8d7 · outbound

This paper cites Learning latent action world models in the wild.arXiv preprint arXiv:2601.05230,.

Self-Supervised Learning of Structured Dynamics from Videos Learning latent action world models in the wild.arXiv preprint arXiv:2601.05230,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:43.835002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:43.835002Z digest=sha256:eff091b3f0f599928c27dc66b23f4f5997fcbb5f48770b121cb772a907e7a77e

Observation 21ed614d-83ad-45e1-ad25-fb18370dd025 · outbound

This paper cites something something.

Self-Supervised Learning of Structured Dynamics from Videos something something

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:43.848667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:43.848667Z digest=sha256:87c84fe8369e6cf73dcc28eac5b8c9c988aa7ec45e7413bfe9dbcc4bde81f5db

Observation 1a5c7dea-afcc-4dae-959e-3bfa5d651e49 · outbound

This paper cites an unresolved cited work.

Self-Supervised Learning of Structured Dynamics from Videos Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:43.920666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:43.920666Z digest=sha256:f03a070e166283ddccda9150bbcad5052da462072bec77667895f789f2287473

Observation f873e337-e9d3-47b8-bf4b-d10e18b844d4 · outbound

This paper cites Siamese masked autoencoders.Advances in Neural Information Processing Systems, 36:40676–40693, 2023.

Self-Supervised Learning of Structured Dynamics from Videos Siamese masked autoencoders.Advances in Neural Information Processing Systems, 36:40676–40693, 2023

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:43.934201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:43.934201Z digest=sha256:1e8684543473be80c5324909c5ffb61ecd50e8b82992fe0f3e9ed5451eba91a8

Observation 547a2166-ea33-4e0a-bb48-c3eee0d8b9d3 · outbound

This paper cites Masked autoencoders are scalable vision learners.

Self-Supervised Learning of Structured Dynamics from Videos Masked autoencoders are scalable vision learners

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:43.936695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:43.936695Z digest=sha256:277f267dc364159fb3335b932f69fc033fca016b8fe25dd7b99e429b97bf5eef

Observation 22fc5f9e-72c4-4c09-b583-80c4ac259aef · outbound

This paper cites Rotary position embedding for vision transformer.

Self-Supervised Learning of Structured Dynamics from Videos Rotary position embedding for vision transformer

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:43.938975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:43.938975Z digest=sha256:b6d8894257f58517a53212c4b11132661db0c68013421a85c07d3a39ca38bf02

Observation c1d0eed5-b78e-4a04-97f1-8097b2ed3d86 · outbound

This paper cites VGGT4D: Mining motion cues in visual geometry transformers for 4d scene reconstruction.arXiv preprint arXiv:2511.19971, 2025.

Self-Supervised Learning of Structured Dynamics from Videos VGGT4D: Mining motion cues in visual geometry transformers for 4d scene reconstruction.arXiv preprint arXiv:2511.19971, 2025

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:43.941401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:43.941401Z digest=sha256:f0688017ffb912eb26d3d01c548aea414eaaa6394a5d8c81db67d660657bf1cd

Observation a8425ea1-c34c-4e13-bdcd-5392d11716e8 · outbound

This paper cites A Frame is Worth One Token: Efficient Generative World Modeling with Delta Tokens.

Self-Supervised Learning of Structured Dynamics from Videos A Frame is Worth One Token: Efficient Generative World Modeling with Delta Tokens

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:43.943838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:43.943838Z digest=sha256:9e1decafe3fb1b765a6e47e298386e3dae82224381cf14d2c694a090d584248f

Observation b3c5fe62-51b3-41a6-b595-f5cc7558a7ac · outbound

This paper cites Depth Anything 3: Recovering the Visual Space from Any Views.

Self-Supervised Learning of Structured Dynamics from Videos Depth Anything 3: Recovering the Visual Space from Any Views

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:43.946535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:43.946535Z digest=sha256:e3ce902776e84788f335f4014a1cebd9824eb9fc945031f28c2f696ed8f4343d

Observation 28d0b606-de80-41da-bdf2-c07efc56d5bc · outbound

This paper cites Towards Understanding Camera Motions in Any Video.

Self-Supervised Learning of Structured Dynamics from Videos Towards Understanding Camera Motions in Any Video

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:43.948987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:43.948987Z digest=sha256:d654f52c9db8dec592a1139fe4909bff32d1b662065c333cd97e224dec677d87

Observation f03464ab-f014-4252-9354-1a64f1527cb4 · outbound

This paper cites DL3DV-10k: A large-scale scene dataset for deep learning-based 3d vision.

Self-Supervised Learning of Structured Dynamics from Videos DL3DV-10k: A large-scale scene dataset for deep learning-based 3d vision

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:43.951383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:43.951383Z digest=sha256:bcd253071f27e8c477fb986718c65dad4a48378efe106a7e3150bdc5f21bc779

Observation 43ae3062-fec4-4b39-99d4-c39099bfd6d5 · outbound

This paper cites 4RC: 4D Reconstruction via Conditional Querying Anytime and Anywhere.

Self-Supervised Learning of Structured Dynamics from Videos 4RC: 4D Reconstruction via Conditional Querying Anytime and Anywhere

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:43.953662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:43.953662Z digest=sha256:54eab56fd2681df1a47cb21716c993bda4765bdfa3bef03aec20102ba7ddcb30

Observation 7a833b1c-186f-4aab-974b-2a97f257b8d4 · outbound

This paper cites V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning.

Self-Supervised Learning of Structured Dynamics from Videos V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:43.956666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:43.956666Z digest=sha256:3eb610c80a1a769b1311c7cf6d19ee4c8a90331699f1b4c01d8e40131273921a

Observation 34d14f7c-97fd-43c8-ac36-5edf4ac34380 · outbound

This paper cites DINOv2: Learning robust visual features without supervision.Transactions on Machine Learning Research, 2023.

Self-Supervised Learning of Structured Dynamics from Videos DINOv2: Learning robust visual features without supervision.Transactions on Machine Learning Research, 2023

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:43.958940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:43.958940Z digest=sha256:54adaf7d5835b5cf90ebec0e84d8748380fb5a853e56cd738a6e1e61fab0a818

Observation 182abe4a-1f14-4b48-82db-e75561513901 · outbound

This paper cites The 2017 DAVIS Challenge on Video Object Segmentation.

Self-Supervised Learning of Structured Dynamics from Videos The 2017 DAVIS Challenge on Video Object Segmentation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:43.961142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:43.961142Z digest=sha256:9b0b7395c3c597f46ed2aab459fd0f202503d23a6e623c3e925d4f17dd4d5c48

Observation 4583a188-4775-43e9-a355-cbe4adf0db60 · outbound

This paper cites Time does tell: Self-supervised time-tuning of dense image representations.

Self-Supervised Learning of Structured Dynamics from Videos Time does tell: Self-supervised time-tuning of dense image representations

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:43.963461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:43.963461Z digest=sha256:9d84ddbbfddf6727b0d56f28c23ec6c32ced6aec30bd533f4bc4f3aabb8ab425

Observation 450b75fa-98fa-4a39-89df-87ec8183382e · outbound

This paper cites MoSiC: Optimal-transport motion trajectory for dense self- supervised learning.

Self-Supervised Learning of Structured Dynamics from Videos MoSiC: Optimal-transport motion trajectory for dense self- supervised learning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:43.965731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:43.965731Z digest=sha256:0d0cc775e62c93cc5c625465d785163810bdff79e4db7eca5ba5c523aed943c9

Observation fa113fe4-b7ac-4da2-bf39-3462434d4f0a · outbound

This paper cites DINOv3.

Self-Supervised Learning of Structured Dynamics from Videos DINOv3

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:43.967955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:43.967955Z digest=sha256:237223101ee9a62940f6c6046d30d6d8d774e5ac7f3f2d2c455cf97fde73ef74

Observation 8ea85e0e-05bb-4463-92be-8ee667b13647 · outbound

This paper cites Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063, 2024.

Self-Supervised Learning of Structured Dynamics from Videos Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063, 2024

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:43.970153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:43.970153Z digest=sha256:56eb86bbd9b885258e2ed1bd8aecf5bb2ee34bf7daf23203c9dbcc7ef7846955

Observation 2645d77d-e259-461a-a17a-7ac3cebee984 · outbound

This paper cites VideoMAE: Masked autoencoders are data-efficient learners for self-supervised video pre-training.Advances in neural information processing systems, 35:10078–10093, 2022.

Self-Supervised Learning of Structured Dynamics from Videos VideoMAE: Masked autoencoders are data-efficient learners for self-supervised video pre-training.Advances in neural information processing systems, 35:10078–10093, 2022

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:43.972365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:43.972365Z digest=sha256:2251cdf81ebccffa2a16bec11aa6b990c37c3a1ce5220971480fcaa71d8ec030

Observation 2a851265-df0a-4b5a-9b3a-ae8248a7c565 · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

Self-Supervised Learning of Structured Dynamics from Videos SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:43.974627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:43.974627Z digest=sha256:bc87fbbbcfa2dd421073a66ab3b7506a6c4d5a52a7234b67cf9b33be2ecbd7a3

Observation 5aa556d7-3cee-41b0-8ff0-63f37f49b2c4 · outbound

This paper cites Is imagenet worth 1 video? learning strong image encoders from 1 long unlabelled video.

Self-Supervised Learning of Structured Dynamics from Videos Is imagenet worth 1 video? learning strong image encoders from 1 long unlabelled video

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:43.977072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:43.977072Z digest=sha256:1a2ed139ff6f3b63a6785e9ad1a44d100c968f78374bec452218a7bd3f084eba

Observation 99103198-8a40-45ba-9adf-1fe3682195b8 · outbound

This paper cites PooDLe: Pooled and dense self-supervised learning from naturalistic videos.

Self-Supervised Learning of Structured Dynamics from Videos PooDLe: Pooled and dense self-supervised learning from naturalistic videos

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:43.979519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:43.979519Z digest=sha256:1ea1a1faec2e1a9d49abf10809b6e96683789506ba65bbedfdd6b7faacbaacfe

Observation 709e7f6f-3816-4c21-9a39-cb847bbe7d88 · outbound

This paper cites VGGT: Visual geometry grounded transformer.

Self-Supervised Learning of Structured Dynamics from Videos VGGT: Visual geometry grounded transformer

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:43.981930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:43.981930Z digest=sha256:42ac14874426d88ec0384c8e6cb0680df5e97739d3587ba5d5519aba7175b195

Observation c973aec4-f6bb-4ab7-8b38-b14ba09dcfab · outbound

This paper cites VideoMAE v2: Scaling video masked autoencoders with dual masking.

Self-Supervised Learning of Structured Dynamics from Videos VideoMAE v2: Scaling video masked autoencoders with dual masking

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:43.984503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:43.984503Z digest=sha256:b59c27f2e85cf310f63c6be5f83de4c025acb2e22b6cbe62cbe5dcb75a48c836

Observation 89c50610-fdc6-4295-9893-a0331c937ea6 · outbound

This paper cites Continuous 3d perception model with persistent state.

Self-Supervised Learning of Structured Dynamics from Videos Continuous 3d perception model with persistent state

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:43.986832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:43.986832Z digest=sha256:3af7b66702055ba093982a79994fe44e2c28646f3b8337e2caefcfedd2dc9125

Observation 58664195-27a9-44fd-b628-84dbfd619214 · outbound

This paper cites DUSt3R: Geometric 3d vision made easy.

Self-Supervised Learning of Structured Dynamics from Videos DUSt3R: Geometric 3d vision made easy

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:43.989146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:43.989146Z digest=sha256:9dd0b4697ba7389d045abb0685c1717e78a93c960753890c4a3bc60769b392fc

Observation bd6afa90-1371-427b-b1c3-e89398834c6e · outbound

This paper cites $\pi^3$: Permutation-Equivariant Visual Geometry Learning.

Self-Supervised Learning of Structured Dynamics from Videos $\pi^3$: Permutation-Equivariant Visual Geometry Learning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:43.991386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:43.991386Z digest=sha256:5a2386a2b1ab621afe0ec7744266880861caf7d9e0d311855bc521d5f53eec31

Observation fde9c463-4331-40e4-ab44-ae8814e8eb1e · outbound

This paper cites CroCo: Self-supervised pre-training for 3d vision tasks by cross-view completion.Advances in Neural Information Processing Systems, 35:3502–3516, 2022.

Self-Supervised Learning of Structured Dynamics from Videos CroCo: Self-supervised pre-training for 3d vision tasks by cross-view completion.Advances in Neural Information Processing Systems, 35:3502–3516, 2022

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:43.993570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:43.993570Z digest=sha256:3b1e7fae51933ead6a38ca079fd1de9a9e2aaf1eb830689e96bc59c2729578ce

Observation 892cc710-3da3-4e36-81e7-21a4b860634c · outbound

This paper cites YouTube-VOS: Sequence-to-sequence video object segmentation.

Self-Supervised Learning of Structured Dynamics from Videos YouTube-VOS: Sequence-to-sequence video object segmentation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:43.995703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:43.995703Z digest=sha256:e776183b31feabf7e09c8333854961a1372d302aab43d3996f63856a0d574036

Observation 35b55820-2081-4f19-8147-b4dd6e62b70f · outbound

This paper cites YouTube-VOS: A Large-Scale Video Object Segmentation Benchmark.

Self-Supervised Learning of Structured Dynamics from Videos YouTube-VOS: A Large-Scale Video Object Segmentation Benchmark

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:43.997932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:43.997932Z digest=sha256:972ad44e935e8955e59f52e2baccc735841e989448a88d0e4b393ba581b11e85

Observation c8f125c4-0e5e-4172-b6aa-e76d34d9bae5 · outbound

This paper cites In pursuit of pixel supervision for visual pre-training.arXiv preprint arXiv:2512.15715, 2025.

Self-Supervised Learning of Structured Dynamics from Videos In pursuit of pixel supervision for visual pre-training.arXiv preprint arXiv:2512.15715, 2025

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:44.000290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:44.000290Z digest=sha256:71165d6498107b4f13b56c6311e76f5553eea52af4b2de42ef8559cf980d17ef

Observation f04e3e93-a621-4f88-bbe7-a42d78715838 · outbound

This paper cites Sigmoid loss for language image pre-training.

Self-Supervised Learning of Structured Dynamics from Videos Sigmoid loss for language image pre-training

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:44.002629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:44.002629Z digest=sha256:e1e753418460d05dc2454936506f51d7bc3b5b31486a882845a6d5a92e69d990

Observation e869fc90-2e3a-43d9-89d8-815c453154a3 · outbound

This paper cites MonST3R: A simple approach for estimating geometry in the presence of motion.

Self-Supervised Learning of Structured Dynamics from Videos MonST3R: A simple approach for estimating geometry in the presence of motion

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:44.004881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:44.004881Z digest=sha256:d8ca2c37511f4ae6eefa6857e38b61eed87c89a8e0f76b265aaf3630996fca11

Observation 87e6ebbf-0bee-495f-86a2-ceff86ce66f4 · outbound

This paper cites DINO-WM: World models on pre-trained visual features enable zero-shot planning.

Self-Supervised Learning of Structured Dynamics from Videos DINO-WM: World models on pre-trained visual features enable zero-shot planning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:44.007097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:44.007097Z digest=sha256:8aa1b709ea7df7ac44ba396dae022cf0e3937bcf5d72f374ecbe2e1dedf31c05

Observation bfebb87d-d8e2-43b6-85dc-94f29411fda9 · outbound

This paper cites Recurrent video masked autoencoders.

Self-Supervised Learning of Structured Dynamics from Videos Recurrent video masked autoencoders

Reference 47

Resolution
malformed identifier
no resolver link, observed 2026-08-01T07:07:44.009137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:44.009137Z digest=sha256:085aea6f27e1677e79e3cd648d68b1c2a9338beafe1147dbb47a6ef85f4891b1

Pith citing papers

No inbound Pith citation observations are available.