Pith. sign in

Paper Citation Record · LEDGER

Scaling 4D Representations

As of 23 August 2026, this Paper Citation Record lists 83 of 83 outbound references and 13 inbound Pith citation observations for arXiv:2412.15212.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.15212 v2

Coverage vector

measured 83 of 83 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T11:36:44.784379Z

measured 96 of 96 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T14:58:05.471418Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T00:26:39.628379Z

Reference resolution

83 of 83 outbound references displayed

  • verified exact0
  • verified fuzzy68
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0b339532-bb5e-4af6-a620-76975105b77b · outbound

This paper cites Learning to see by moving.

Scaling 4D Representations Learning to see by moving

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:44.354073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:36:44.354073Z digest=sha256:6cea6e51c32d6d9041a4f5da781ac534af2471e250541a9a69331abd3ad58cc5

Observation cf470e0f-5277-4e63-8bb8-17d39113ebec · outbound

This paper cites Flamingo: a visual language model for few-shot learning.NeurIPS, 2022.

Scaling 4D Representations Flamingo: a visual language model for few-shot learning.NeurIPS, 2022

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:44.359767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:36:44.359767Z digest=sha256:ba18e942fd5c313b734da480a00213cd8630a5f11a2b2906bc73887b84665cb5

Observation 662649da-014c-44dc-9414-82ada87ce246 · outbound

This paper cites Self-supervised learning by cross-modal audio-video clustering.

Scaling 4D Representations Self-supervised learning by cross-modal audio-video clustering

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:44.365223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:36:44.365223Z digest=sha256:7055bf1286538f1fe266adc263362d9643c5ebd668636c60e689f96ed2244c27

Observation c10bbcd5-9c71-4960-8f22-2645104fbdd0 · outbound

This paper cites Look, listen and learn.

Scaling 4D Representations Look, listen and learn

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:36:46.386223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T11:36:44.370818Z digest=sha256:2409937fcda6472675a1c9d9b484ec55a64b5df529f3400db50c8fdbea4111d0

Observation c658eb3e-2ea4-46b3-a254-1705fe88234c · outbound

This paper cites Objects that sound.

Scaling 4D Representations Objects that sound

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:36:46.368100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T11:36:44.376859Z digest=sha256:895bc2cd1acd86588db1ce928a0622a7daa018121d93f68dcdad4ce9d543c1db

Observation 16179761-9a9e-4210-a727-73d0c9be023b · outbound

This paper cites Vivit: A video vi- sion transformer.

Scaling 4D Representations Vivit: A video vi- sion transformer

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:36:46.348568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T11:36:44.382654Z digest=sha256:03159bba5aff2e1b9d5573fb83a79b917d4be87b6ecd4e85b152076be5a262c3

Observation 7eddce4d-481e-42b1-8e33-4682642feebd · outbound

This paper cites Yuille, Trevor Darrell, Jitendra Malik, and Alexei A.

Scaling 4D Representations Yuille, Trevor Darrell, Jitendra Malik, and Alexei A

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:36:46.327250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T11:36:44.388743Z digest=sha256:c1905b717e91397c9ad19103311af1b8932533ea903dc3531b880441d97e3b32

Observation a0c4a0cc-5f7d-425e-a74c-82815ff07527 · outbound

This paper cites Revisiting feature prediction for learning visual rep- resentations from video.

Scaling 4D Representations Revisiting feature prediction for learning visual rep- resentations from video

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:36:46.308947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T11:36:44.394216Z digest=sha256:2872f03509700850b850291d5425eacca9a4ae890c9b2b528f39c7aeebbdff0c

Observation 2ca458b1-eaac-4e82-8075-37dd30dac47c · outbound

This paper cites Physion: Evaluating physical prediction from vision in humans and machines.

Scaling 4D Representations Physion: Evaluating physical prediction from vision in humans and machines

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:36:46.289261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T11:36:44.399820Z digest=sha256:dce22ebcd366490a034231f5713fda8fbca581464638814cccabc856b5101335

Observation 952ee2be-95a3-42b3-a3b1-2f1536bec186 · outbound

This paper cites Zoedepth: Zero-shot transfer by com- bining relative and metric depth.

Scaling 4D Representations Zoedepth: Zero-shot transfer by com- bining relative and metric depth

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:36:46.261251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T11:36:44.405025Z digest=sha256:b179597d516465722c0fa9931612357a8814afaa58f3c07bb01b45e77713d9ac

Observation 470119a5-a728-47d2-b01a-33f65d853cf0 · outbound

This paper cites Deep regression on manifolds: a 3d rotation case study.

Scaling 4D Representations Deep regression on manifolds: a 3d rotation case study

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:36:46.238400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T11:36:44.410003Z digest=sha256:1e0f07c92c099867ad39aaf88438203732fd8ecc7ddc67ec7ae0b44499f2a32f

Observation e7ab0d43-18e1-48cd-a0a1-431000193af7 · outbound

This paper cites Activitynet: A large-scale video benchmark for human activity understanding.

Scaling 4D Representations Activitynet: A large-scale video benchmark for human activity understanding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:44.414894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:36:44.414894Z digest=sha256:824dcd9650df42a45f3da6a51bee813383df85caf2ef2fd754ffa1385ba4a1ff

Observation e27a8fe5-b26e-4900-81b2-85849eaa52e1 · outbound

This paper cites Generative pre- training from pixels.

Scaling 4D Representations Generative pre- training from pixels

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:36:46.209135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T11:36:44.420227Z digest=sha256:76fb8ee72d1395543fbe64e152f8eb0d6f8199ecba7353692299b57d69411c0a

Observation 78290ec3-c7e5-4978-a6ff-79e2078fb82e · outbound

This paper cites A simple framework for contrastive learning of visual representations.

Scaling 4D Representations A simple framework for contrastive learning of visual representations

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:36:46.187482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T11:36:44.424999Z digest=sha256:8015e5bf41a482f8ad84d9a7d58eae09b810dcef457f57959a54f141c69a6a32

Observation 1b2718de-c323-4f1a-954f-b0760812c3d6 · outbound

This paper cites PaLI-3 vision language models: Smaller, faster, stronger.

Scaling 4D Representations PaLI-3 vision language models: Smaller, faster, stronger

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:36:46.163254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T11:36:44.430032Z digest=sha256:f50d1d40a2e5f2b70fabd6c41809983010cce6e18e3d985ac7603725c203488c

Observation f7725220-fe7d-4be8-96c1-e8f4e56e71f3 · outbound

This paper cites A unified architecture for natural language processing: Deep neural networks with multitask learning.

Scaling 4D Representations A unified architecture for natural language processing: Deep neural networks with multitask learning

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:36:46.142695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T11:36:44.435072Z digest=sha256:235e5404f30ed87c1234464ef516b4fb8535219af43e4dc313301e50739870ba

Observation ebcf9b9b-d7fa-4559-a9a8-482ceb41f14d · outbound

This paper cites Scannet: Richly-annotated 3d reconstructions of indoor scenes.

Scaling 4D Representations Scannet: Richly-annotated 3d reconstructions of indoor scenes

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:36:46.124391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T11:36:44.439902Z digest=sha256:a738e05aef0992ce665c06cc431cf86acabbf5bb5b7a4dac1787afba41f0b8ac

Observation c7a9af97-80d3-422b-9df6-2facdef4b3c2 · outbound

This paper cites Scaling egocentric vision: The epic-kitchens dataset.

Scaling 4D Representations Scaling egocentric vision: The epic-kitchens dataset

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:36:46.107048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T11:36:44.445719Z digest=sha256:0f4f555dd23d20ac2ae5302add720f4f6cc4dd4ed7c138967f67fd50cbfc8b4c

Observation 50e07444-1252-4e4e-9661-0d83d8d4a623 · outbound

This paper cites Scaling vision transformers to 22 billion pa- rameters.

Scaling 4D Representations Scaling vision transformers to 22 billion pa- rameters

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:36:46.086169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T11:36:44.450707Z digest=sha256:2b42f60c77c46310c12e50e4afcea600e63134bbb28b5381a756bd21308e359b

Observation 4b68ec81-0a88-4fc0-8921-ee4363d57d59 · outbound

This paper cites Unsuper- vised visual representation learning by context prediction.

Scaling 4D Representations Unsuper- vised visual representation learning by context prediction

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:44.456233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:36:44.456233Z digest=sha256:13d0caae9ba154107c977695d47bee71ee6c03d2ca788858b9514bd454110be3

Observation 2f9c645e-3faf-4d17-a87a-a946352dbe79 · outbound

This paper cites TAP-vid: A bench- mark for tracking any point in a video.

Scaling 4D Representations TAP-vid: A bench- mark for tracking any point in a video

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:36:46.047510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T11:36:44.461739Z digest=sha256:1149331718d0cf7ae0707d8187ddcddeef3a6049405139e365ed22b1b1f37117

Observation 01d6782e-17ff-437c-840f-d051952c183f · outbound

This paper cites TAPIR: Tracking any point with per-frame initialization and temporal refinement.

Scaling 4D Representations TAPIR: Tracking any point with per-frame initialization and temporal refinement

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:36:46.028211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T11:36:44.466748Z digest=sha256:e04baadb883db2201631a1bc2cac7fa9f725186b29741ff3db99f480cde00ea8

Observation 980b072f-a4a6-4b6e-85d1-10c0b558138a · outbound

This paper cites Boot- sTAP: Bootstrapped training for tracking any point.

Scaling 4D Representations Boot- sTAP: Bootstrapped training for tracking any point

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:36:46.009921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T11:36:44.472024Z digest=sha256:82c0fbb58dbb0a5bb18e70b5ef746074108e31d0ee95df4df5967e4f02fa6d73

Observation d85b95bb-494b-4c9b-a766-2e1751f1d87b · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

Scaling 4D Representations An image is worth 16x16 words: Transformers for image recognition at scale

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:36:45.986631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T11:36:44.477427Z digest=sha256:548b230a8af13c0f003fae9c25a52d25a6551f0a775853219dd9de1057b5df4f

Observation 8149e4d6-62b2-46ba-90aa-2aee56ab9d2c · outbound

This paper cites Prob- ing the 3d awareness of visual foundation models.

Scaling 4D Representations Prob- ing the 3d awareness of visual foundation models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:36:45.967667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T11:36:44.483266Z digest=sha256:3881787bb57caca1b302fe3db383436ef77a796440be47f0e2e953781cc3ff7e

Observation 11b57ec8-2b55-496c-81c8-ad1ab73851a7 · outbound

This paper cites Scalable pre- training of large autoregressive image models.

Scaling 4D Representations Scalable pre- training of large autoregressive image models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:36:45.946847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T11:36:44.488711Z digest=sha256:d8345580f553514c9400a33f34a76852b669da3e950b3b75aaee65767d83da41

Observation cbdcdae6-b054-4dca-82d1-2f5c2517b874 · outbound

This paper cites SA Vi++: Towards end-to-end object-centric learning from real-world videos.

Scaling 4D Representations SA Vi++: Towards end-to-end object-centric learning from real-world videos

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:36:45.927939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T11:36:44.493907Z digest=sha256:5deb115e67668546df9e48106851bf57f5eb33f6634e6ee0d6d9c8e3560633c5

Observation a6f3db81-a5af-404c-9abb-4bffa133347b · outbound

This paper cites Spatiotemporal residual networks for video action recogni- tion.

Scaling 4D Representations Spatiotemporal residual networks for video action recogni- tion

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:36:45.910275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T11:36:44.498819Z digest=sha256:ab81f7b3dc3d17085ed3b28e90caf825e38ca32b22bf09799ebc484877467ebf

Observation 1cc776fa-20ed-484d-a36b-7c45bbcad05f · outbound

This paper cites Slowfast networks for video recognition.

Scaling 4D Representations Slowfast networks for video recognition

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:44.503894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:36:44.503894Z digest=sha256:8c70b6c2be464eb75d5c3bd7e9e46dd1d97ea8d35ade0e762004195724341392

Observation b951d335-bee8-4660-b415-b4ad64cfb981 · outbound

This paper cites A large-scale study on unsupervised spatiotemporal representation learning.

Scaling 4D Representations A large-scale study on unsupervised spatiotemporal representation learning

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:36:45.881938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T11:36:44.508721Z digest=sha256:c864d881635202a185c26bd04ef34a76d2080dcd56da0846dddd78c6b84fe472

Observation 4ac01222-34ab-43e4-96ec-48e8fc11e549 · outbound

This paper cites Distributed hier- archical processing in the primate cerebral cortex.

Scaling 4D Representations Distributed hier- archical processing in the primate cerebral cortex

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:36:45.863127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T11:36:44.513732Z digest=sha256:a7e6caf998f1c7aad417b7be7d4067faf9b7cb24518d99768d65e2ac8f6da0f8

Observation 8529e2cd-b473-4fe8-a8ec-f95a577ef7a0 · outbound

This paper cites Learning invariance from transformation se- quences.

Scaling 4D Representations Learning invariance from transformation se- quences

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:36:45.841178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T11:36:44.518633Z digest=sha256:ee54c65839cbdfe2fbec65ccb0ccea6904c4cfe10ad4c5f6698db573a1fcdfdd

Observation 5f749498-3321-4791-a6b3-8fc2c4a6662c · outbound

This paper cites an unresolved cited work.

Scaling 4D Representations Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-11T11:36:45.823038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T11:36:44.524798Z digest=sha256:d0b3fdd45b7377746f48e47eec7d32f5c1d56fab4590a1630537857f04415324

Observation de41f93e-0a26-4c5d-88fe-6dc04e181cce · outbound

This paper cites Learn- ing to linearize under uncertainty.

Scaling 4D Representations Learn- ing to linearize under uncertainty

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:36:45.807276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T11:36:44.529883Z digest=sha256:9b4f878473b7c604a31f95f1ac7cb70c8c9eb175bd104958e029cb1628bc5fae

Observation 6a837ddb-613d-4c10-981b-8d4a1751f540 · outbound

This paper cites The” something something” video database for learning and evaluating visual common sense.

Scaling 4D Representations The” something something” video database for learning and evaluating visual common sense

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:36:45.789753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T11:36:44.534714Z digest=sha256:8a0a0c92754864a4331a339cc026231d8b3bef5b0818e43fde8d8e9bc31e44f3

Observation 944b4d25-d168-4266-9f56-f4e0face1b11 · outbound

This paper cites an unresolved cited work.

Scaling 4D Representations Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-11T11:36:45.766954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T11:36:44.539521Z digest=sha256:1a10ba9f86d6387607cf09082df47b4c98c5d955638743f975d3dc3599a0c061

Observation 8b952943-1a6a-43ad-b817-66005dd3982e · outbound

This paper cites Memory- augmented dense predictive coding for video representation learning.

Scaling 4D Representations Memory- augmented dense predictive coding for video representation learning

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:36:45.748370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T11:36:44.544826Z digest=sha256:51c812ba8e3b6e56a3666861633e4176231c00ba4ad1b442adb689b34e27f8c6

Observation 264c93e1-577b-42c3-9099-229716f64b24 · outbound

This paper cites Self- supervised co-training for video representation learning.

Scaling 4D Representations Self- supervised co-training for video representation learning

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:36:45.729882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T11:36:44.549996Z digest=sha256:1af0c95197e9b6a67574b09f3d788744cf43ec2d3918f4fd8bd1044bb86f2189

Observation cd76d221-6c15-49cd-88f3-65c09b1f0f85 · outbound

This paper cites Masked autoencoders are scalable vision learners.

Scaling 4D Representations Masked autoencoders are scalable vision learners

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:36:45.712791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T11:36:44.555139Z digest=sha256:9122efe5ab89ca8be5281c8a715c1c88146148cf2abd4658a4df4119630af422

Observation 43151a63-3d22-4585-bb0a-0e4fba8352f4 · outbound

This paper cites Data-efficient image recognition with contrastive predictive coding.

Scaling 4D Representations Data-efficient image recognition with contrastive predictive coding

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:36:45.692499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T11:36:44.560183Z digest=sha256:0c8953b3185d3f34211004b78f181d6db676c3304f3d511d9d6883182701ff1a

Observation 1600f371-b1eb-4762-a851-64da1dfddb98 · outbound

This paper cites Representation learn- ing with video deep infomax.

Scaling 4D Representations Representation learn- ing with video deep infomax

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:36:45.671888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T11:36:44.565039Z digest=sha256:d2d8f7bb93bdf74cc953b72e1504186d5de0efcee2f4468c12fc25efcaa48c7e

Observation 9a4e9834-cdee-46a2-af31-099304396f39 · outbound

This paper cites The kinetics human action video dataset, 2017.

Scaling 4D Representations The kinetics human action video dataset, 2017

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:36:45.652634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T11:36:44.570042Z digest=sha256:c14de69117f099cae8ba658284cd65a2802b29075d618ce96fa1c4244149c477

Observation 336990ec-141c-4e48-8f93-b07cf87d402f · outbound

This paper cites Condi- tional object-centric learning from video.

Scaling 4D Representations Condi- tional object-centric learning from video

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:36:45.635274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T11:36:44.575002Z digest=sha256:cf973cda4a20aba1976736758a3da046fbe47c47e073896851c1c454e89445ff

Observation ffd50336-2e97-4534-9526-22ee3a190c0e · outbound

This paper cites Coopera- tive learning of audio and video models from self-supervised synchronization.

Scaling 4D Representations Coopera- tive learning of audio and video models from self-supervised synchronization

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:36:45.618463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T11:36:44.580047Z digest=sha256:ae1e6c8732685c0661fb4b3566bcf692acb7a20d9c6a8c248226763259e72c56

Observation 9b3a51b1-a16d-4ff7-9196-c70c1e81d226 · outbound

This paper cites Decoupled weight decay regularization.

Scaling 4D Representations Decoupled weight decay regularization

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:44.585082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:36:44.585082Z digest=sha256:d19084980a1f095e3f38c3d19391a018952dc273f9a98d1d4d1d747d8b5bced6

Observation 7118bcf7-3c7e-4615-bfc4-265bf3cd8f56 · outbound

This paper cites Object vision and spatial vision: two cortical path- ways.

Scaling 4D Representations Object vision and spatial vision: two cortical path- ways

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:36:45.590723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T11:36:44.590119Z digest=sha256:76fb0a077e258877574904e47c68482728a43195034c7ea82a64b0e5887fb53a

Observation 043fbb96-2f40-4c0c-b452-658bcc1ea143 · outbound

This paper cites Deep learning from temporal coherence in video.

Scaling 4D Representations Deep learning from temporal coherence in video

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:36:45.570025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T11:36:44.595390Z digest=sha256:6f09737f42344710e189e1695ad039fcd5198cc500e768effbb9dc2e7bb616cc

Observation fae716d6-0ba1-49bf-9880-919e8337cdfc · outbound

This paper cites Audio- visual instance discrimination with cross-modal agreement.

Scaling 4D Representations Audio- visual instance discrimination with cross-modal agreement

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:36:45.543544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T11:36:44.600415Z digest=sha256:0326e05de0904f1b6395e8eb24304f7beb6c6916128190900caebd9428f1abfa

Observation b74c613d-5ea4-49af-b0d5-fb8aa63ba36c · outbound

This paper cites Atlas: End- to-end 3d scene reconstruction from posed images.

Scaling 4D Representations Atlas: End- to-end 3d scene reconstruction from posed images

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:44.605220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:36:44.605220Z digest=sha256:14828fb13d6ef03458432d4a1de8d05e809ef9eeea8d7c9e65cbcaafab634474

Observation f1a85c6a-e7e0-4a0b-8d3d-ac99c50bfb22 · outbound

This paper cites an unresolved cited work.

Scaling 4D Representations Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-11T11:36:45.513807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T11:36:44.610251Z digest=sha256:d6cb369e1d85c5d8eba82e0d9f5dd06fa0476d01a81c99733a6d0bd599bf3b5c

Observation c3d11847-36cb-4cf0-b3ba-cc83e4ae29a5 · outbound

This paper cites Fully sharded data parallel: faster ai training with fewer gpus.

Scaling 4D Representations Fully sharded data parallel: faster ai training with fewer gpus

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:36:45.494406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T11:36:44.615109Z digest=sha256:b7ea3c1c60f6d6f81ab7bc5545bffb82208ef2d4d11ab5c7e9621091a637b128

Observation 6094bae1-83ff-4d21-91e8-ce3e0944b628 · outbound

This paper cites Audio-visual scene analysis with self-supervised multisensory features.

Scaling 4D Representations Audio-visual scene analysis with self-supervised multisensory features

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:36:45.461267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T11:36:44.620015Z digest=sha256:9f89bdaecb08b0e962f078fed9a72ca92da1d6729cc13c88e085d4f861852181

Observation 73e05ce9-cdde-4a70-9aff-550ffe691b5a · outbound

This paper cites Context encoders: Feature learning by inpainting.

Scaling 4D Representations Context encoders: Feature learning by inpainting

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:36:45.432140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T11:36:44.625046Z digest=sha256:f9aaf609b5d38d9cd1bb6741be119b588ee48c8430d2ed7e5b3f6d4b8913b633

Observation 92e1b67d-3531-45b7-af13-56a231503c45 · outbound

This paper cites Learning features by watching ob- jects move.

Scaling 4D Representations Learning features by watching ob- jects move

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:36:45.411765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T11:36:44.630083Z digest=sha256:771e385311814aed2100b275e3a3215af3fbcf7f075c3414af18e1007129c429

Observation 65957157-1fdd-4a39-a0fd-c759560f712c · outbound

This paper cites Perception test: A diagnostic benchmark for mul- timodal video models.

Scaling 4D Representations Perception test: A diagnostic benchmark for mul- timodal video models

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:36:45.389270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T11:36:44.634804Z digest=sha256:90c371e2bc30bcb01c9de6ea87ed224a5af079ff7b1c05317f2e878c044a4091

Observation e6b5101f-2a7a-464c-b2c9-8edb3d41da86 · outbound

This paper cites Asano, Ruth Fong, Jo ˜ao F.

Scaling 4D Representations Asano, Ruth Fong, Jo ˜ao F

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:36:45.365374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T11:36:44.639791Z digest=sha256:2f6f559580f77838b957bf4ec2f8d31f28f6691b29355c8fba810051b58547ba

Observation 96ae4680-7f6b-4366-9b2d-4559f66f0d26 · outbound

This paper cites Spatiotempo- ral contrastive video representation learning.

Scaling 4D Representations Spatiotempo- ral contrastive video representation learning

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:36:45.342771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T11:36:44.644737Z digest=sha256:7f90e86f91810127940aa2e7781a6d09dba472088abd83f7744e0a2c3768f409

Observation 28b8b729-232a-4b07-8567-cf22db7bea2b · outbound

This paper cites Improving language understanding by generative pre-training.

Scaling 4D Representations Improving language understanding by generative pre-training

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:36:45.325299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T11:36:44.649622Z digest=sha256:fbc88c4c4f10dbef42850e3c556bdc32e15ef4c12cfbbb7049d53288333c89a7

Observation a69fee24-1b84-4747-af8c-6f155bb275cc · outbound

This paper cites Learning transferable visual models from natural language supervision, 2021.

Scaling 4D Representations Learning transferable visual models from natural language supervision, 2021

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:44.654983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:36:44.654983Z digest=sha256:808cbeca126088057af558bfb96e5856054f7012d38f5efe0072551a596d272b

Observation 514204a2-b968-4065-a126-d4023f49adf3 · outbound

This paper cites Vi- sion transformers for dense prediction.

Scaling 4D Representations Vi- sion transformers for dense prediction

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:44.660314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:36:44.660314Z digest=sha256:c28bc03c5e94807785568db9dc4c4a7f0929f1f664f372098e74ab33bd20b28c

Observation 1ead10f5-7ab3-4fe3-86ed-aeb7edd1b75f · outbound

This paper cites Video (lan- guage) modeling: a baseline for generative models of natural videos.

Scaling 4D Representations Video (lan- guage) modeling: a baseline for generative models of natural videos

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:36:45.277112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T11:36:44.665462Z digest=sha256:5e2ecd2c1bf5847f798ac0d3ec4ce3eac7845d08cc9163fd56bff07720ce33b6

Observation 486ee18c-3a64-4bcc-89aa-9fb47a615e7f · outbound

This paper cites Broaden your views for self-supervised video learning.

Scaling 4D Representations Broaden your views for self-supervised video learning

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:36:45.260379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T11:36:44.670899Z digest=sha256:b436fdccae045fbc59b68a6141b1dd98055861be77bed29e271803f146e15d39

Observation 26dd5823-a6e0-4fa5-87b1-43a83b349148 · outbound

This paper cites Learning to localize sound source in visual scenes.

Scaling 4D Representations Learning to localize sound source in visual scenes

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:36:45.242197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T11:36:44.675919Z digest=sha256:e64b7caf0ab65514dd502a2d3cf382b476e8384b1f53864bc0206400640d12bd

Observation 4fcba358-e2cb-44ac-9b09-72b27f264076 · outbound

This paper cites Two-stream con- volutional networks for action recognition in videos.

Scaling 4D Representations Two-stream con- volutional networks for action recognition in videos

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:36:45.224949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T11:36:44.681099Z digest=sha256:a4c8e4e83ef66b5d955974c1bc960ff6d4c4689a9714f59c1f3b26c84feb0d6b

Observation 8fe5f959-0955-4d4d-90ad-2085c89c9540 · outbound

This paper cites A Short Note on the Kinetics-700-2020 Human Action Dataset.

Scaling 4D Representations A Short Note on the Kinetics-700-2020 Human Action Dataset

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:44.686114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:36:44.686114Z digest=sha256:4e9e2d9d03275bdc1081c196e45e2c91a21c4a3912c42c80a7820a6cb804cb24

Observation ac78d9a9-ec84-435d-9575-4d8af7d2e6de · outbound

This paper cites Scalability in perception for autonomous driving: Waymo open dataset.

Scaling 4D Representations Scalability in perception for autonomous driving: Waymo open dataset

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:36:45.207170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T11:36:44.691436Z digest=sha256:308b8886f8644e65cdc3a7564bf6dcfc7c4e83acb74d1bccc4140e5df01a466f

Observation 2616d095-999d-4c51-a0ef-5ea8d62de70e · outbound

This paper cites Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training.

Scaling 4D Representations Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:36:45.186125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T11:36:44.696334Z digest=sha256:092901d110e7a996b9951069d52f51122d4989a80ff8e2f6cfefaf3c7b8bedd7

Observation 39d41e55-08ef-43df-9ac7-34f1a761a08f · outbound

This paper cites RealEstate10K.

Scaling 4D Representations RealEstate10K

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:36:45.163517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T11:36:44.701207Z digest=sha256:b237ea3628548bf237712d944a38afc15557daab30a4aa3b6eee3343907156f2

Observation fd0f2790-0b42-4008-80da-97b19df9f25f · outbound

This paper cites Hud- son, Thomas Albert Keck, Joao Carreira, Alexey Dosovit- skiy, Mehdi S.

Scaling 4D Representations Hud- son, Thomas Albert Keck, Joao Carreira, Alexey Dosovit- skiy, Mehdi S

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:36:45.144915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T11:36:44.706491Z digest=sha256:9ed2375e2f9b4c0852e55d74682e5aa974f7c20e8e91284e60cebc2ecd8ed644

Observation 5d725e0d-58d4-492f-93fc-3c7d05034ef2 · outbound

This paper cites An- ticipating the future by watching unlabeled video.

Scaling 4D Representations An- ticipating the future by watching unlabeled video

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:36:45.127424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T11:36:44.711464Z digest=sha256:b307adef9e07b7a99306664765676c29b2356c6feffea7168ef97fadecd245d4

Observation a828f910-15b3-46a5-9d5b-589b1b30b86c · outbound

This paper cites Videomae v2: Scaling video masked autoencoders with dual masking.

Scaling 4D Representations Videomae v2: Scaling video masked autoencoders with dual masking

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:36:45.108616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T11:36:44.716074Z digest=sha256:5a7cdeccc0b9a98e755c43e6d2e9efed3e1ac1134d1074da37e142ca16cc8ce7

Observation c361c815-92a8-4ac2-89f9-9204a67a5105 · outbound

This paper cites Masked video distillation: Rethinking masked feature mod- eling for self-supervised video representation learning.

Scaling 4D Representations Masked video distillation: Rethinking masked feature mod- eling for self-supervised video representation learning

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:36:45.082943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T11:36:44.720987Z digest=sha256:219a529c0292b8a2cfbbbb6751347c330f6a88bb8d1c61d2d5efee3017d91b8b

Observation f575aaac-9dc3-426e-9772-03bd552b2705 · outbound

This paper cites Dust3r: Geometric 3d vi- sion made easy.

Scaling 4D Representations Dust3r: Geometric 3d vi- sion made easy

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:36:45.058181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T11:36:44.727013Z digest=sha256:cf999dc3e825f6c5fc2f87a2a1f0d5e409ed52bc30de44f8195e91402d198d86

Observation b4072c0e-7f1c-4142-9901-00ed740a4c4e · outbound

This paper cites Unsupervised learning of visual representations using videos.

Scaling 4D Representations Unsupervised learning of visual representations using videos

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:44.732519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:36:44.732519Z digest=sha256:ada0640ad8e26165c2071d1ecb2f4777bc6eede96b9fefee172fb5b02aaddd55

Observation 0b443d20-db6f-4640-aa81-a1a7d4a29a83 · outbound

This paper cites Internvideo: General video foundation models via generative and discriminative learning.

Scaling 4D Representations Internvideo: General video foundation models via generative and discriminative learning

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:36:45.023854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T11:36:44.737783Z digest=sha256:5b7b8879fc3c611752cf90d24f6a6b15ce3cd44c01566123bd0186716d842725

Observation 1e6ea0ac-238f-432b-8a31-1b7cad8c415e · outbound

This paper cites Less is more: Consistent video depth estimation with masked frames modeling.

Scaling 4D Representations Less is more: Consistent video depth estimation with masked frames modeling

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:36:45.005827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T11:36:44.743101Z digest=sha256:e28d4c75383456121f7adf79e77a2c4a85edfd6a747e64f4a61a03aba6132cc3

Observation 830813fd-2280-4028-a09f-52dcb43d744d · outbound

This paper cites Controlling space and time with dif- fusion models.

Scaling 4D Representations Controlling space and time with dif- fusion models

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:36:44.987316Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T11:36:44.748910Z digest=sha256:3c850714f65b425e367e5ec19aa92710d4e0d8e3d764782114ed03b96a8c9220

Observation 75cdb228-4dec-4e78-9775-9b3c396b71e3 · outbound

This paper cites Slow feature analysis: Unsupervised learning of invariances.

Scaling 4D Representations Slow feature analysis: Unsupervised learning of invariances

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:36:44.963970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T11:36:44.754204Z digest=sha256:3ddad565a1be491b87906c560745a907e0add2357cb9ce032ca4692fd52c362d

Observation c514da1a-3ac7-45eb-9771-de72d24d97c7 · outbound

This paper cites Depth anything: Unleashing the power of large-scale unlabeled data.

Scaling 4D Representations Depth anything: Unleashing the power of large-scale unlabeled data

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:36:44.944943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T11:36:44.759162Z digest=sha256:862d2f82e00142cb74f37542063bb5454920e8872b2b279dbc0d1e45366ef7b6

Observation 5985c939-6421-429b-acfc-fe1d587550ef · outbound

This paper cites Sigmoid loss for language image pre-training.

Scaling 4D Representations Sigmoid loss for language image pre-training

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:36:44.927363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T11:36:44.764313Z digest=sha256:10eea56afe35eadb0290d8b046fa3ba76b2c660b916c251823b07c7606b58abd

Observation 12071b3b-2248-4965-9ac5-f858106e9b2a · outbound

This paper cites A general protocol to probe large vision models for 3d physical understanding.

Scaling 4D Representations A general protocol to probe large vision models for 3d physical understanding

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:36:44.909064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T11:36:44.769749Z digest=sha256:8e442c265c9eacc1b5837322cd52558e287ce04a1fd19c00a45a940db8abd3f2

Observation 26a37d42-9f15-4813-9b9d-6a0050b9e5b6 · outbound

This paper cites Videoprism: A foundational visual encoder for video understanding.

Scaling 4D Representations Videoprism: A foundational visual encoder for video understanding

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:36:44.889195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T11:36:44.775000Z digest=sha256:4322d7faf131cc0275aafa2043d4fcd824cad73233a406302c8ba48edb08ae13

Observation f2a9b05a-c703-41e1-a744-00cf8fe2d95d · outbound

This paper cites Taking something from somewhere.

Scaling 4D Representations Taking something from somewhere

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:36:44.863759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T11:36:44.784379Z digest=sha256:ee49e7345cb0ecdf7bdbd64a382b749add8e401a5ff81a66fab6ad90892f117c

Pith citing papers

Observation 647ea506-fb21-412f-8709-191a86556497 · inbound

From Image to Video: An Empirical Study of Diffusion Representations cites this paper.

From Image to Video: An Empirical Study of Diffusion Representations Scaling 4D Representations

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T14:10:55.171296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:10:55.171296Z digest=sha256:2e6938d5d0104172f40bb286d7f87e216fef35aa56310e2753543a96e42ad0e7

Observation 6cadd0b9-0116-42a5-a924-64ae95092733 · inbound

MoSiC: Optimal-Transport Motion Trajectory for Dense Self-Supervised Learning cites this paper.

MoSiC: Optimal-Transport Motion Trajectory for Dense Self-Supervised Learning Scaling 4D Representations

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T05:11:16.394161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:11:16.394161Z digest=sha256:eb6cbef74a8ecbe9d298ae3886e12b33f62aacca53e762be90f7476b23a58b90

Observation 3f95764d-646c-414b-9d2c-358331e7599a · inbound

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning cites this paper.

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning Scaling 4D Representations

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:33:50.730654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T00:33:50.471804Z digest=sha256:53576cd295e1c977e52a431004275435d3eb7ab8863a7bd20ec6096085ba8b6c

Observation cd2fcdce-93e6-443e-a1d2-e60e744413ca · inbound

SciVid: Cross-Domain Evaluation of Video Models in Scientific Applications cites this paper.

SciVid: Cross-Domain Evaluation of Video Models in Scientific Applications Scaling 4D Representations

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:06.382164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:06.382164Z digest=sha256:fb8d28f0ea5c31f388f757e01eb718b2ca09a7cd8630d38a64a1101963ff714d

Observation 987e4f4f-0610-45b9-a150-53cfdb66061e · inbound

Frozen Forecasting: A Unified Evaluation cites this paper.

Frozen Forecasting: A Unified Evaluation Scaling 4D Representations

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-19T03:42:57.307979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T03:42:54.620069Z digest=sha256:e5eacfdaddb8d05eb198958ef673a5ad6c5b13a0b40e3582d85a2c5e0c396c3c

Observation b38649bd-72f1-4353-8d6a-7c3454b70b61 · inbound

Unique Lives, Shared World: Learning from Single-Life Videos cites this paper.

Unique Lives, Shared World: Learning from Single-Life Videos Scaling 4D Representations

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T18:43:43.563351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:43:43.563351Z digest=sha256:b8ab36537b2db532c879783be8110e26eacd25a165ec2e0954afee0fdb9d38de

Observation 17bdcb93-8b3e-47f5-953b-7f19a69424b2 · inbound

LA-Pose: Latent Action Pretraining Meets Pose Estimation cites this paper.

LA-Pose: Latent Action Pretraining Meets Pose Estimation Scaling 4D Representations

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:51:28.839842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-07T08:57:27.503349Z digest=sha256:56d7931ca49cd21d7b9997af1762159f42e996fa1bbaf7194416086a43802e45

Observation 7f655eea-b5fc-4541-8a3f-5952a9598b38 · inbound

Towards Data-Efficient Video Pre-training with Frozen Image Foundation Models cites this paper.

Towards Data-Efficient Video Pre-training with Frozen Image Foundation Models Scaling 4D Representations

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:33:13.031523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-20T10:28:42.027773Z digest=sha256:baeaff66b958fe663219ecc247c868dcb60d8255b4493bf8fafdac85fc542aa9

Observation 394c74aa-a514-4f79-96b3-0293b0ebf52c · inbound

Where Do We (Not) Need Temporal Context in Low-Resource Video Task Adaptation? cites this paper.

Where Do We (Not) Need Temporal Context in Low-Resource Video Task Adaptation? Scaling 4D Representations

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T02:56:28.965993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T10:31:10.409674Z digest=sha256:5cbaecf988dcd05f143e7c5a5b904dc37282212e80c3e71169f57eec63c1ccd9

Observation a3458577-a4a7-4f4c-a660-a63981166db6 · inbound

Gen4U: Unifying Video Generation and Understanding via Diffusion cites this paper.

Gen4U: Unifying Video Generation and Understanding via Diffusion Scaling 4D Representations

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T00:26:39.630147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-10T00:16:43.190961Z digest=sha256:0b1ca0c9a304ff93af9fbeae2bf5dc06ebfbdc9e6359e34e462978ec30f4b4e5

Observation 50b5a864-d3ae-4dc9-a564-937d49a0c200 · inbound

SeeSE3: Emergence of 3D Space in Vision Features cites this paper.

SeeSE3: Emergence of 3D Space in Vision Features Scaling 4D Representations

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-02T02:49:58.712379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:49:58.712379Z digest=sha256:ca78bb7436f8aa19a757bc87886c8fa60c582759a9565210eab10aa388b1efa6

Observation 9b3a923d-c9c3-4986-9f24-86606157a485 · inbound

Self-Supervised Learning of Structured Dynamics from Videos cites this paper.

Self-Supervised Learning of Structured Dynamics from Videos Scaling 4D Representations

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:43.307749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:43.307749Z digest=sha256:0ba758ba3a5c7b86161e78fbef8ef57d8c6dab1d42fd26043285df74db04d95d

Observation 623c607f-1b6e-4707-bb06-c685e28ce65f · inbound

A Controlled Study of Self-Supervised Image and Video Pretraining under Limited Resources cites this paper.

A Controlled Study of Self-Supervised Image and Video Pretraining under Limited Resources Scaling 4D Representations

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T14:58:05.471418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:58:05.471418Z digest=sha256:fed58cc5fc22d650d13b910366c5acc371a322aab8e52e7c972680900ac9d650