Pith. sign in

Paper Citation Record · LEDGER

Self-Supervised Learning of Structured Dynamics from Videos

As of 10 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 0 inbound Pith citation observations for arXiv:2607.21576.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.21576 v1

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T07:07:44.009137Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

47 of 47 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved46
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fcf5aec1-d427-4934-92b1-88dbb6c38a08 · outbound

This paper cites V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning.

Self-Supervised Learning of Structured Dynamics from Videos V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:42.700451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:42.700451Z digest=sha256:277879558cb086df998a8df56402a708f9808d7c00dba78757ef99403b024ba8

Observation 645df39c-c401-4de0-81ef-3b53daab1ac7 · outbound

This paper cites Back to the Features: DINO as a Foundation for Video World Models.

Self-Supervised Learning of Structured Dynamics from Videos Back to the Features: DINO as a Foundation for Video World Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:42.787898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:42.787898Z digest=sha256:6c4f911878178ea6a48b3ea1b536d50cf36ef446ae258de981fd0d3d11dd2054

Observation e57671f3-34fe-4506-9642-2ad1495ef298 · outbound

This paper cites Revisiting feature prediction for learning visual representations from video.Transactions on Machine Learning Research, 2024.

Self-Supervised Learning of Structured Dynamics from Videos Revisiting feature prediction for learning visual representations from video.Transactions on Machine Learning Research, 2024

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:42.875142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:42.875142Z digest=sha256:8eb51afdc7c1b374f77cddedd09fa187481c0721ca5467cc356b509eafa6a89d

Observation ff4ba9dd-6950-4f73-9f34-ac9508a804ce · outbound

This paper cites VFMF: World modeling by forecasting vision foundation model features.arXiv preprint arXiv:2512.11225,.

Self-Supervised Learning of Structured Dynamics from Videos VFMF: World modeling by forecasting vision foundation model features.arXiv preprint arXiv:2512.11225,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:42.965780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:42.965780Z digest=sha256:435f0ec6d1014aeb0008e8017afe3d1261ed70893c7cd4c65e051873a507c316

Observation de75f00c-62f0-4962-9043-dd50da99aee3 · outbound

This paper cites TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment.

Self-Supervised Learning of Structured Dynamics from Videos TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:43.056117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:43.056117Z digest=sha256:fca9cf679738bbce93984d735c2ff7cb6090ccc33076bf969b130bad3ca661be

Observation e902e301-623b-4c03-9de9-fc397afdd936 · outbound

This paper cites A Short Note on the Kinetics-700 Human Action Dataset.

Self-Supervised Learning of Structured Dynamics from Videos A Short Note on the Kinetics-700 Human Action Dataset

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:43.147125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:43.147125Z digest=sha256:534eadf16262db183ee65684289ea8e0fef1c6bf16fd9b71b08b815553f11e6e

Observation 9b3a923d-c9c3-4986-9f24-86606157a485 · outbound

This paper cites Scaling 4D Representations.

Self-Supervised Learning of Structured Dynamics from Videos Scaling 4D Representations

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:43.307749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:43.307749Z digest=sha256:c308681383989002e674e6e176595104087d24817e542f7c2c0e4e12610bf445

Observation fe9778c6-5fa5-4553-aa9b-0415daf664ba · outbound

This paper cites Vision transformers need registers.

Self-Supervised Learning of Structured Dynamics from Videos Vision transformers need registers

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:43.455509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:43.455509Z digest=sha256:7636daf26b6a15adc4e04fff758f65c4020d086e5433bc02ebd82e6d3097bc60

Observation 672b183b-0146-4b90-b179-08ac7841accb · outbound

This paper cites Probing the 3d awareness of visual foundation models.

Self-Supervised Learning of Structured Dynamics from Videos Probing the 3d awareness of visual foundation models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:43.538067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:43.538067Z digest=sha256:6448aa051fd9506d0825f6850495811d0cd80eceec0b0a8239d2b8975ee73a06

Observation 8c1e3832-5293-4856-a12b-ae26f0fa47d4 · outbound

This paper cites Scalable pre-training of large autoregressive image models.

Self-Supervised Learning of Structured Dynamics from Videos Scalable pre-training of large autoregressive image models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:43.612958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:43.612958Z digest=sha256:4899727faa4ea67dd435139fbc4296433bfbad39a1b64ab087eb40477f1864e0

Observation 4f62d957-313a-43c0-bc5f-819ab38c4242 · outbound

This paper cites Multimodal autoregressive pre-training of large vision encoders.

Self-Supervised Learning of Structured Dynamics from Videos Multimodal autoregressive pre-training of large vision encoders

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:43.715885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:43.715885Z digest=sha256:b6ecedd5155a91df6040f58925705f90337056c3fa093c63a68d3176c8a2ce9e

Observation 594d648b-9a2c-4095-a313-23e54a26b8d7 · outbound

This paper cites Learning latent action world models in the wild.arXiv preprint arXiv:2601.05230,.

Self-Supervised Learning of Structured Dynamics from Videos Learning latent action world models in the wild.arXiv preprint arXiv:2601.05230,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:43.835002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:43.835002Z digest=sha256:78c1abeaf6c1eb9fbff5d4e97485edf8e1607c35299943b9ca0072581de0f59a

Observation 21ed614d-83ad-45e1-ad25-fb18370dd025 · outbound

This paper cites something something.

Self-Supervised Learning of Structured Dynamics from Videos something something

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:43.848667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:43.848667Z digest=sha256:dbed1577556509b6d12c9a6359266467b2368a5f0fce0339c3442f3aeef7bf76

Observation 1a5c7dea-afcc-4dae-959e-3bfa5d651e49 · outbound

This paper cites an unresolved cited work.

Self-Supervised Learning of Structured Dynamics from Videos Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:43.920666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:43.920666Z digest=sha256:f59d9fb327d364a2efb6e795a4c2ca35168081fcf222642c69e0f30fe19c9cab

Observation f873e337-e9d3-47b8-bf4b-d10e18b844d4 · outbound

This paper cites Siamese masked autoencoders.Advances in Neural Information Processing Systems, 36:40676–40693, 2023.

Self-Supervised Learning of Structured Dynamics from Videos Siamese masked autoencoders.Advances in Neural Information Processing Systems, 36:40676–40693, 2023

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:43.934201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:43.934201Z digest=sha256:cc0fba75fb3fa834f9e98df111e2d05de5c7bbff7ead6ed782b6ffc288a569d4

Observation 547a2166-ea33-4e0a-bb48-c3eee0d8b9d3 · outbound

This paper cites Masked autoencoders are scalable vision learners.

Self-Supervised Learning of Structured Dynamics from Videos Masked autoencoders are scalable vision learners

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:43.936695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:43.936695Z digest=sha256:d4b64df98a9bb1e3f9100f1cd4c509ebfeb8edce8e3681027cba40ad8b22c7c8

Observation 22fc5f9e-72c4-4c09-b583-80c4ac259aef · outbound

This paper cites Rotary position embedding for vision transformer.

Self-Supervised Learning of Structured Dynamics from Videos Rotary position embedding for vision transformer

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:43.938975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:43.938975Z digest=sha256:287c85334d029094faae2b4c0f4c20b6b86d02176c26fe930a2dca63c3e7073f

Observation c1d0eed5-b78e-4a04-97f1-8097b2ed3d86 · outbound

This paper cites VGGT4D: Mining motion cues in visual geometry transformers for 4d scene reconstruction.arXiv preprint arXiv:2511.19971, 2025.

Self-Supervised Learning of Structured Dynamics from Videos VGGT4D: Mining motion cues in visual geometry transformers for 4d scene reconstruction.arXiv preprint arXiv:2511.19971, 2025

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:43.941401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:43.941401Z digest=sha256:89d82535092201b177e79f3f06a9aa2ea8cd12b7320804ad5379eec65ca95126

Observation a8425ea1-c34c-4e13-bdcd-5392d11716e8 · outbound

This paper cites A Frame is Worth One Token: Efficient Generative World Modeling with Delta Tokens.

Self-Supervised Learning of Structured Dynamics from Videos A Frame is Worth One Token: Efficient Generative World Modeling with Delta Tokens

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:43.943838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:43.943838Z digest=sha256:562cfa69abf838c533d9dcc73ea2691f41cf1aeb717d9c904ebd782e20cd3ee2

Observation b3c5fe62-51b3-41a6-b595-f5cc7558a7ac · outbound

This paper cites Depth Anything 3: Recovering the Visual Space from Any Views.

Self-Supervised Learning of Structured Dynamics from Videos Depth Anything 3: Recovering the Visual Space from Any Views

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:43.946535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:43.946535Z digest=sha256:8a9116071159fa746edce87d3cd5c457a014e97ee3c8a4be02a917e5331c0bad

Observation 28d0b606-de80-41da-bdf2-c07efc56d5bc · outbound

This paper cites Towards Understanding Camera Motions in Any Video.

Self-Supervised Learning of Structured Dynamics from Videos Towards Understanding Camera Motions in Any Video

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:43.948987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:43.948987Z digest=sha256:eb7b22f0636f3cdd9c00bda91c5978ca09aef21d4e5055fb28bd70b70d16024a

Observation f03464ab-f014-4252-9354-1a64f1527cb4 · outbound

This paper cites DL3DV-10k: A large-scale scene dataset for deep learning-based 3d vision.

Self-Supervised Learning of Structured Dynamics from Videos DL3DV-10k: A large-scale scene dataset for deep learning-based 3d vision

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:43.951383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:43.951383Z digest=sha256:1654d04054736909ea18115a20f12c8c114ebee78dbe2eda2b2ccc1de7378aff

Observation 43ae3062-fec4-4b39-99d4-c39099bfd6d5 · outbound

This paper cites 4RC: 4D Reconstruction via Conditional Querying Anytime and Anywhere.

Self-Supervised Learning of Structured Dynamics from Videos 4RC: 4D Reconstruction via Conditional Querying Anytime and Anywhere

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:43.953662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:43.953662Z digest=sha256:46d60f49813193fd710b42915f5d3c9e3e6b120b81a7b2ad3ad84c9dc8dcb79e

Observation 7a833b1c-186f-4aab-974b-2a97f257b8d4 · outbound

This paper cites V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning.

Self-Supervised Learning of Structured Dynamics from Videos V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:43.956666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:43.956666Z digest=sha256:7603aec6311f45ca6d7445062809c83f213b20dc7985eecc88c8a13ddbcfd5e0

Observation 34d14f7c-97fd-43c8-ac36-5edf4ac34380 · outbound

This paper cites DINOv2: Learning robust visual features without supervision.Transactions on Machine Learning Research, 2023.

Self-Supervised Learning of Structured Dynamics from Videos DINOv2: Learning robust visual features without supervision.Transactions on Machine Learning Research, 2023

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:43.958940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:43.958940Z digest=sha256:a4ba45eb092db36902d3016e1d57a3c4257886c82a66190b9a96e9acd38ed9d9

Observation 182abe4a-1f14-4b48-82db-e75561513901 · outbound

This paper cites The 2017 DAVIS Challenge on Video Object Segmentation.

Self-Supervised Learning of Structured Dynamics from Videos The 2017 DAVIS Challenge on Video Object Segmentation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:43.961142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:43.961142Z digest=sha256:46e8a6616c76b6f96c2b0490ee671fe774b1d4ff2bcf74ade61a995a0b9c82a5

Observation 4583a188-4775-43e9-a355-cbe4adf0db60 · outbound

This paper cites Time does tell: Self-supervised time-tuning of dense image representations.

Self-Supervised Learning of Structured Dynamics from Videos Time does tell: Self-supervised time-tuning of dense image representations

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:43.963461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:43.963461Z digest=sha256:5c80287f064fc292ac3cc03393dbd9e87568780fca746ffe59328f5d973c7b8b

Observation 450b75fa-98fa-4a39-89df-87ec8183382e · outbound

This paper cites MoSiC: Optimal-transport motion trajectory for dense self- supervised learning.

Self-Supervised Learning of Structured Dynamics from Videos MoSiC: Optimal-transport motion trajectory for dense self- supervised learning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:43.965731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:43.965731Z digest=sha256:3c3eab535413ff487de7f17d6959a9aeb0896a42140a583c4813b868f53e58e8

Observation fa113fe4-b7ac-4da2-bf39-3462434d4f0a · outbound

This paper cites DINOv3.

Self-Supervised Learning of Structured Dynamics from Videos DINOv3

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:43.967955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:43.967955Z digest=sha256:60a8313fb114274943cc18c7ede63e8d9c3e4a964e6babe14c3e84a60dc270aa

Observation 8ea85e0e-05bb-4463-92be-8ee667b13647 · outbound

This paper cites Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063, 2024.

Self-Supervised Learning of Structured Dynamics from Videos Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063, 2024

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:43.970153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:43.970153Z digest=sha256:86cdffe0dc43564a3d4c604f65524453d1a82f17e101badb6ff457661afb4e0e

Observation 2645d77d-e259-461a-a17a-7ac3cebee984 · outbound

This paper cites VideoMAE: Masked autoencoders are data-efficient learners for self-supervised video pre-training.Advances in neural information processing systems, 35:10078–10093, 2022.

Self-Supervised Learning of Structured Dynamics from Videos VideoMAE: Masked autoencoders are data-efficient learners for self-supervised video pre-training.Advances in neural information processing systems, 35:10078–10093, 2022

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:43.972365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:43.972365Z digest=sha256:351962c6464ab1b08a8a9bc027ab12f0c25ae93e1eae8c7cbf5d70eb80c29d8f

Observation 2a851265-df0a-4b5a-9b3a-ae8248a7c565 · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

Self-Supervised Learning of Structured Dynamics from Videos SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:43.974627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:43.974627Z digest=sha256:59b84102f309a8ff2ad781e5f07be1181445b1700d75ea1f2d52e1fc78116bdb

Observation 5aa556d7-3cee-41b0-8ff0-63f37f49b2c4 · outbound

This paper cites Is imagenet worth 1 video? learning strong image encoders from 1 long unlabelled video.

Self-Supervised Learning of Structured Dynamics from Videos Is imagenet worth 1 video? learning strong image encoders from 1 long unlabelled video

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:43.977072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:43.977072Z digest=sha256:652b2927df2bea213571777cbacbddc950954cc755b8cfcae54934beef35466b

Observation 99103198-8a40-45ba-9adf-1fe3682195b8 · outbound

This paper cites PooDLe: Pooled and dense self-supervised learning from naturalistic videos.

Self-Supervised Learning of Structured Dynamics from Videos PooDLe: Pooled and dense self-supervised learning from naturalistic videos

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:43.979519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:43.979519Z digest=sha256:d3e96e12f908334d06e7d26309e510c6e5882faac812101f476435d6d6fae2cc

Observation 709e7f6f-3816-4c21-9a39-cb847bbe7d88 · outbound

This paper cites VGGT: Visual geometry grounded transformer.

Self-Supervised Learning of Structured Dynamics from Videos VGGT: Visual geometry grounded transformer

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:43.981930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:43.981930Z digest=sha256:a1c604868c6263717caee043aafa22ec389205e607c9bb88aa24546dccb9a645

Observation c973aec4-f6bb-4ab7-8b38-b14ba09dcfab · outbound

This paper cites VideoMAE v2: Scaling video masked autoencoders with dual masking.

Self-Supervised Learning of Structured Dynamics from Videos VideoMAE v2: Scaling video masked autoencoders with dual masking

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:43.984503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:43.984503Z digest=sha256:7cf91ed3df6c08c42578319bdc35354ecb3046d466804b080459ea9d92df3001

Observation 89c50610-fdc6-4295-9893-a0331c937ea6 · outbound

This paper cites Continuous 3d perception model with persistent state.

Self-Supervised Learning of Structured Dynamics from Videos Continuous 3d perception model with persistent state

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:43.986832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:43.986832Z digest=sha256:58ad8ebba5c497ccb528f38671f71e5f5ff787a63b0bd94c9c5440dba8748b42

Observation 58664195-27a9-44fd-b628-84dbfd619214 · outbound

This paper cites DUSt3R: Geometric 3d vision made easy.

Self-Supervised Learning of Structured Dynamics from Videos DUSt3R: Geometric 3d vision made easy

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:43.989146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:43.989146Z digest=sha256:dd02d625b6df89cd226cb8c54dc563fae4f2ccc8830238968bc6f303b8019023

Observation bd6afa90-1371-427b-b1c3-e89398834c6e · outbound

This paper cites $\pi^3$: Permutation-Equivariant Visual Geometry Learning.

Self-Supervised Learning of Structured Dynamics from Videos $\pi^3$: Permutation-Equivariant Visual Geometry Learning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:43.991386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:43.991386Z digest=sha256:547286d7a4ca8159da87cb47c28516a9d4f5b587da4e716f8d6c1ef393cafd79

Observation fde9c463-4331-40e4-ab44-ae8814e8eb1e · outbound

This paper cites CroCo: Self-supervised pre-training for 3d vision tasks by cross-view completion.Advances in Neural Information Processing Systems, 35:3502–3516, 2022.

Self-Supervised Learning of Structured Dynamics from Videos CroCo: Self-supervised pre-training for 3d vision tasks by cross-view completion.Advances in Neural Information Processing Systems, 35:3502–3516, 2022

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:43.993570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:43.993570Z digest=sha256:358df2919685bd0178c3fb988bcb5ca5cdebb76e6c2521c2ac18ca1f7a78e759

Observation 892cc710-3da3-4e36-81e7-21a4b860634c · outbound

This paper cites YouTube-VOS: Sequence-to-sequence video object segmentation.

Self-Supervised Learning of Structured Dynamics from Videos YouTube-VOS: Sequence-to-sequence video object segmentation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:43.995703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:43.995703Z digest=sha256:d3be367e567f2e0b16bc367170be6c4b493abbe94b2d6a044540b820597b7b2d

Observation 35b55820-2081-4f19-8147-b4dd6e62b70f · outbound

This paper cites YouTube-VOS: A Large-Scale Video Object Segmentation Benchmark.

Self-Supervised Learning of Structured Dynamics from Videos YouTube-VOS: A Large-Scale Video Object Segmentation Benchmark

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:43.997932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:43.997932Z digest=sha256:eb1a38ab8de3c58b2be46dd2088996f78ccb0ea0b9dc229dd50d017f0ecdee24

Observation c8f125c4-0e5e-4172-b6aa-e76d34d9bae5 · outbound

This paper cites In pursuit of pixel supervision for visual pre-training.arXiv preprint arXiv:2512.15715, 2025.

Self-Supervised Learning of Structured Dynamics from Videos In pursuit of pixel supervision for visual pre-training.arXiv preprint arXiv:2512.15715, 2025

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:44.000290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:44.000290Z digest=sha256:3ff913e91046b445c971ce98e240806baac6f0fd3236d85a989e85b0ccaf7010

Observation f04e3e93-a621-4f88-bbe7-a42d78715838 · outbound

This paper cites Sigmoid loss for language image pre-training.

Self-Supervised Learning of Structured Dynamics from Videos Sigmoid loss for language image pre-training

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:44.002629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:44.002629Z digest=sha256:c55392c9578d46179d7320699734178cd422ab3d96c5d767dadd722afa531f9c

Observation e869fc90-2e3a-43d9-89d8-815c453154a3 · outbound

This paper cites MonST3R: A simple approach for estimating geometry in the presence of motion.

Self-Supervised Learning of Structured Dynamics from Videos MonST3R: A simple approach for estimating geometry in the presence of motion

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:44.004881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:44.004881Z digest=sha256:f5a5055c1c61f30e911811f01041f2d08be873cf2ba9809e03ba1d35e4c23001

Observation 87e6ebbf-0bee-495f-86a2-ceff86ce66f4 · outbound

This paper cites DINO-WM: World models on pre-trained visual features enable zero-shot planning.

Self-Supervised Learning of Structured Dynamics from Videos DINO-WM: World models on pre-trained visual features enable zero-shot planning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:44.007097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:44.007097Z digest=sha256:2813dfcb1f9d0f6663d3f27e3c71f5d5026b2fcf155636e4b3f69aa9888e258a

Observation bfebb87d-d8e2-43b6-85dc-94f29411fda9 · outbound

This paper cites Recurrent video masked autoencoders.

Self-Supervised Learning of Structured Dynamics from Videos Recurrent video masked autoencoders

Reference 47

Resolution
malformed identifier
no resolver link, observed 2026-08-01T07:07:44.009137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:44.009137Z digest=sha256:cd9c40969e527d1e678b7143ef1c46cc5fa0de381d45ff2bf052e735a5e46aa7

Pith citing papers

No inbound Pith citation observations are available.