Pith. sign in

Paper Citation Record · LEDGER

FRAME: Pre-Training Video Feature Representations via Anticipation and Memory

As of 8 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 0 inbound Pith citation observations for arXiv:2506.05543.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.05543 v1

Coverage vector

measured 57 of 57 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:27:00.912867Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

57 of 57 outbound references displayed

  • verified exact10
  • verified fuzzy9
  • unresolved35
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3747041c-8cd9-492d-be05-5258c858e53d · outbound

This paper cites Self-supervised Object-Centric Learning for Videos.

FRAME: Pre-Training Video Feature Representations via Anticipation and Memory Self-supervised Object-Centric Learning for Videos

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-07T10:27:03.244802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:26:56.075692Z digest=sha256:f44491b356cb6f2483ea30b43872178bc8de0c3e7f72e3d1a65931acd3d38958

Observation dec122db-7350-433f-8049-759b4eb3a591 · outbound

This paper cites Fully-Convolutional Siamese Networks for Object Tracking.

FRAME: Pre-Training Video Feature Representations via Anticipation and Memory Fully-Convolutional Siamese Networks for Object Tracking

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-07T10:27:03.060266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:26:56.143077Z digest=sha256:71059f027880af3a53b053c4197055d8ec9bc98a21cff231a59f82ff2c84b404

Observation 241f286c-143a-41d6-aad5-4fc4f5104a5b · outbound

This paper cites an unresolved cited work.

FRAME: Pre-Training Video Feature Representations via Anticipation and Memory Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:27:04.002367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:26:56.207833Z digest=sha256:cf6e122efb740ec2a26016e065389b8503d874a30c2638f0255b9f944efd114e

Observation 9d7786a1-03c0-4b7c-9345-89b60baab0f6 · outbound

This paper cites Emerging Properties in Self-Supervised Vision Transformers.

FRAME: Pre-Training Video Feature Representations via Anticipation and Memory Emerging Properties in Self-Supervised Vision Transformers

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T10:26:56.327573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:26:56.327573Z digest=sha256:91dd023963e03efbff9aef6c340a6fb4300deead45ab47f509a7d678423d56a8

Observation 136b593e-054b-40c0-8cba-1e6c5ac50521 · outbound

This paper cites A Simple Framework for Contrastive Learning of Visual Representations.

FRAME: Pre-Training Video Feature Representations via Anticipation and Memory A Simple Framework for Contrastive Learning of Visual Representations

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T10:26:56.430091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:26:56.430091Z digest=sha256:93f161fc9bda110f4342f3a3ede62dcd32ad91b080049e950277b337da93dce8

Observation 5c9830a1-1da1-4750-ab2b-2dfec7727654 · outbound

This paper cites Exploring Simple Siamese Representation Learning.

FRAME: Pre-Training Video Feature Representations via Anticipation and Memory Exploring Simple Siamese Representation Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T10:26:56.652735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:26:56.652735Z digest=sha256:d4705f66850b40e222963763d8c4e7fb099f2783848ecc50f490efc10a9708ff

Observation cca99645-80ee-476b-819f-e7e2ec483e01 · outbound

This paper cites Tracking Anything with Decoupled Video Segmentation.

FRAME: Pre-Training Video Feature Representations via Anticipation and Memory Tracking Anything with Decoupled Video Segmentation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T10:26:56.755323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:26:56.755323Z digest=sha256:e61c40581cec382d681fb077252947a21370171990301392a8588222255aad5c

Observation f613318c-2866-40df-a95f-55976169af5f · outbound

This paper cites Putting the Object Back into Video Object Segmentation.

FRAME: Pre-Training Video Feature Representations via Anticipation and Memory Putting the Object Back into Video Object Segmentation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T10:26:56.844695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:26:56.844695Z digest=sha256:9f5d365dcea317611358d972b79accaf880b9de5dfa8de14d3e4f3f5e12b416e

Observation 8272b971-cc30-4a9e-b18a-7d640c83a150 · outbound

This paper cites an unresolved cited work.

FRAME: Pre-Training Video Feature Representations via Anticipation and Memory Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:27:03.994420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:26:56.935230Z digest=sha256:990ba5bc6bcb6c558c04896ebc2a2abfca39bbad077684a61b5473866d094a35

Observation 8725609c-f12c-4f4b-882d-156bb4d77ae9 · outbound

This paper cites Doersch, A.

FRAME: Pre-Training Video Feature Representations via Anticipation and Memory Doersch, A

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:03.985668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:26:57.058328Z digest=sha256:d53f40881353f0f8bcfc7f194b3dbe4a4a56654cccdcd87cfd4ea249131e55b7

Observation 0433ac2b-f287-49f3-a4bf-5ea004b69a6b · outbound

This paper cites Eymaël, R.

FRAME: Pre-Training Video Feature Representations via Anticipation and Memory Eymaël, R

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:03.977742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:26:57.154505Z digest=sha256:d39232c0bd740d40f9b9783bf830d19049bd45ad196ec19a47073f2a46d1cebf

Observation e26bfb12-99b0-4ff8-91e5-7094c69b37fa · outbound

This paper cites A Large-Scale Study on Unsupervised Spatiotemporal Representation Learning.

FRAME: Pre-Training Video Feature Representations via Anticipation and Memory A Large-Scale Study on Unsupervised Spatiotemporal Representation Learning

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-07T10:27:02.867254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:26:57.235226Z digest=sha256:ec7b394b5093e13eed186bf227f2b46381b7fba63a71b3dbfee7438e72153dc5

Observation ce2c6fce-6f4e-4aac-82f3-a201add82bc6 · outbound

This paper cites Grauman, A.

FRAME: Pre-Training Video Feature Representations via Anticipation and Memory Grauman, A

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:03.968950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:26:57.306125Z digest=sha256:f35042540fab1ce9e9b70939e675d5652178e4b48f3ef6bcd325755ae2cebc9b

Observation 6d00ed5d-2980-49be-8e70-6a311b2d186d · outbound

This paper cites Gupta, J.

FRAME: Pre-Training Video Feature Representations via Anticipation and Memory Gupta, J

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:03.960068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:26:57.390376Z digest=sha256:0618fd08206176ebc1f48271c7d05df4c40c8cd73478b60725a83e304aeab166

Observation a4c8d04b-9d96-4fb1-af2c-d3cdb9c0ef66 · outbound

This paper cites Momentum Contrast for Unsupervised Visual Representation Learning.

FRAME: Pre-Training Video Feature Representations via Anticipation and Memory Momentum Contrast for Unsupervised Visual Representation Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T10:26:57.500232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:26:57.500232Z digest=sha256:2936e825082f4c19f0572b58862a8065060ba0c426fa220c79b766da9e9d1c6e

Observation f4ab7a08-36a7-4184-bfe3-f99d8183d1d3 · outbound

This paper cites an unresolved cited work.

FRAME: Pre-Training Video Feature Representations via Anticipation and Memory Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:27:03.951006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:26:57.569273Z digest=sha256:c024012a9a003e8bafd0653c9f2b6f98ae14fffdc794588f9e6390dff520f4f0

Observation e4a6811e-1fd1-463e-b401-2314274aeabd · outbound

This paper cites an unresolved cited work.

FRAME: Pre-Training Video Feature Representations via Anticipation and Memory Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:27:03.942775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:26:57.640645Z digest=sha256:fedcbc332d6b0a5b47e10540cdbdeec76d8c0eb9d6555ed61a468f3f68f7af4e

Observation 15ee4353-4502-46f3-b31a-18fb3f4163e9 · outbound

This paper cites Space-Time Correspondence as a Contrastive Random Walk.

FRAME: Pre-Training Video Feature Representations via Anticipation and Memory Space-Time Correspondence as a Contrastive Random Walk

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T10:26:57.863321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:26:57.863321Z digest=sha256:1604ad858a3bd11fa65fce28162db2dfa69dd8c9a729d2e80e03ef76c53c3481

Observation f54fb717-9f43-494a-948f-8ae6e567440a · outbound

This paper cites Jhuang, J.

FRAME: Pre-Training Video Feature Representations via Anticipation and Memory Jhuang, J

Reference 22

Resolution
malformed identifier
no resolver link, observed 2026-08-07T10:26:57.927831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:26:57.927831Z digest=sha256:f4a0affa67a7466d85b30165ca5e75216f466c8553e77e5d7bc16201b712b42c

Observation 63aa4cfe-383e-47b3-849f-dd5d24a31edf · outbound

This paper cites CoTracker: It is Better to Track Together.

FRAME: Pre-Training Video Feature Representations via Anticipation and Memory CoTracker: It is Better to Track Together

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T10:26:57.990243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:26:57.990243Z digest=sha256:37c7ea6fa9f4cef326551cdac5d4fc2ca5dcd58b03f09fdc48b5f232b8eea836

Observation d94900eb-8934-406d-93ce-d7c2fc3a4cfe · outbound

This paper cites The Kinetics Human Action Video Dataset.

FRAME: Pre-Training Video Feature Representations via Anticipation and Memory The Kinetics Human Action Video Dataset

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T10:26:58.052557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:26:58.052557Z digest=sha256:c042c7c76e3ac0b0e9f24ef4cc07ab5c1ab64ac8b510eac04356c01fb409d935

Observation 61509798-0e3f-4c3b-b2ff-d3fa8e119566 · outbound

This paper cites Khosla, S.

FRAME: Pre-Training Video Feature Representations via Anticipation and Memory Khosla, S

Reference 25

Resolution
verified exact
raw_fallback, observed 2026-08-07T10:27:02.610398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:26:58.101630Z digest=sha256:87fa53b48de579337411290df610d4ef62ebca3afd264a72af5e6ab759ba145e

Observation 3fd45c79-3869-4050-b1c2-fc2e153ac534 · outbound

This paper cites Segment Anything.

FRAME: Pre-Training Video Feature Representations via Anticipation and Memory Segment Anything

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T10:26:58.145168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:26:58.145168Z digest=sha256:59488545fd53b30132c360ca63ef58bea7598a795896bec03d35a4b0afb64d68

Observation 2daaabb7-c22a-41f2-8edf-9f8b1818bb7e · outbound

This paper cites Kuehne, H.

FRAME: Pre-Training Video Feature Representations via Anticipation and Memory Kuehne, H

Reference 27

Resolution
malformed identifier
no resolver link, observed 2026-08-07T10:26:58.201275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:26:58.201275Z digest=sha256:6162bfd271af4c87ecb2ff9a7580f9fe33a8da937b2d86a6bd04adf390e79043

Observation 344394e4-3429-49a4-89a9-158b4aaa81a7 · outbound

This paper cites Joint-task Self-supervised Learning for Temporal Correspondence.

FRAME: Pre-Training Video Feature Representations via Anticipation and Memory Joint-task Self-supervised Learning for Temporal Correspondence

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-08-07T10:27:02.318566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:26:58.239798Z digest=sha256:7703e13f91f600d73a8ba27103a5ebd8023901eacc68b85542a143438bfd6e18

Observation afe415da-e1c5-4bf7-9549-add591fc2b49 · outbound

This paper cites an unresolved cited work.

FRAME: Pre-Training Video Feature Representations via Anticipation and Memory Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:27:03.934310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:26:58.315087Z digest=sha256:f808ace2a19c22337c415aec902c733b0615ec005c46dd0d143adf12d18c3672

Observation 5d8db677-0ef6-42f1-8878-dd5e2aada7c8 · outbound

This paper cites an unresolved cited work.

FRAME: Pre-Training Video Feature Representations via Anticipation and Memory Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T10:26:58.347979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:26:58.347979Z digest=sha256:323e8802e7528ff58c64802e965307499df4aa115cd7d44c299cb69b970b1e59

Observation 916fd0e2-a0a5-4594-bcf8-77e410fa51c5 · outbound

This paper cites Misra, C.

FRAME: Pre-Training Video Feature Representations via Anticipation and Memory Misra, C

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:03.926676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:26:58.418402Z digest=sha256:6fbf21d2dbc2993c73d8bb245071781408ccd8f0171e1c88e106cff72d49e113

Observation 3b7a744a-7df0-48b2-b03e-9fbcea87adef · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

FRAME: Pre-Training Video Feature Representations via Anticipation and Memory DINOv2: Learning Robust Visual Features without Supervision

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T10:26:58.495985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:26:58.495985Z digest=sha256:c78714c18825e95e70c64d76a08c0559bad55e4c99feb9c57de82274267293da

Observation 42c73411-c453-4acc-be82-4f89be80b56f · outbound

This paper cites The 2017 DAVIS Challenge on Video Object Segmentation.

FRAME: Pre-Training Video Feature Representations via Anticipation and Memory The 2017 DAVIS Challenge on Video Object Segmentation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T10:26:58.556924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:26:58.556924Z digest=sha256:565f21b79bdf133d52e5c6b42911d8d99e30b335742ad93d577c50b1d2312ea4

Observation 76778535-a78d-40ee-991e-bd8795000eb3 · outbound

This paper cites an unresolved cited work.

FRAME: Pre-Training Video Feature Representations via Anticipation and Memory Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:27:03.918396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:26:58.650435Z digest=sha256:643ec94be935556f62a762c82bf3c71a536a1e7a75ecd0538c53fd09da8b7dae

Observation c6cc08d7-bbfa-4765-9ecf-de748d330a6f · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

FRAME: Pre-Training Video Feature Representations via Anticipation and Memory Learning Transferable Visual Models From Natural Language Supervision

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T10:26:58.719677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:26:58.719677Z digest=sha256:02f76c432eb189ac47bb0dafa35c2bddfbadec30a3b0dca3643a32973d1caef1

Observation 12096951-7ea5-481a-97c3-1c5b1f02b4f4 · outbound

This paper cites Ranasinghe, M.

FRAME: Pre-Training Video Feature Representations via Anticipation and Memory Ranasinghe, M

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:03.910214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:26:58.802843Z digest=sha256:8d80cabe10f313347047c82b8b47de852931fff2b55f1c836e9584aba0d657fa

Observation 652e3848-feb2-48d4-a04c-a985cb022826 · outbound

This paper cites Ranzinger, G.

FRAME: Pre-Training Video Feature Representations via Anticipation and Memory Ranzinger, G

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:03.901689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:26:58.847210Z digest=sha256:e16d31a3c2ee788867cf051f0619398566cf3ebf2fe1b5cae0dcd1615eeea98f

Observation 7a21a2bd-f4da-43b5-b5d5-203e765557a6 · outbound

This paper cites Fine-tuned CLIP Models are Efficient Video Learners.

FRAME: Pre-Training Video Feature Representations via Anticipation and Memory Fine-tuned CLIP Models are Efficient Video Learners

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-08-07T10:27:02.098834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:26:58.906976Z digest=sha256:7820f38911fbf8fb310640c306c3522e8a83d3c9847c31b538fea20b2257d913

Observation cd207a04-67fe-4ffa-a901-cf4d916bc411 · outbound

This paper cites an unresolved cited work.

FRAME: Pre-Training Video Feature Representations via Anticipation and Memory Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:27:03.893723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:26:58.961025Z digest=sha256:79446ea0be429c48730114df749749db840290bbe6ed27e1880e928bd6be6cb6

Observation 59beb273-b3c3-423f-9071-09fde54c3d36 · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

FRAME: Pre-Training Video Feature Representations via Anticipation and Memory SAM 2: Segment Anything in Images and Videos

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T10:26:59.026576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:26:59.026576Z digest=sha256:7476bceca257ccf262a1766c95ee3b0aba0e73b26fd478c506023a6a80263cf7

Observation 608eebaa-edbd-464e-950f-02c5d95ccf1d · outbound

This paper cites Salehi, E.

FRAME: Pre-Training Video Feature Representations via Anticipation and Memory Salehi, E

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:03.885315Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:26:59.099479Z digest=sha256:0da387ac13723e5c778f688069dd6dd14b0c7c21e7d98bffb12878399c69b42f

Observation 1ce80e29-79f4-4392-8aa9-50d17e1f2385 · outbound

This paper cites Sameni, K.

FRAME: Pre-Training Video Feature Representations via Anticipation and Memory Sameni, K

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:03.877127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:26:59.204553Z digest=sha256:7452078c91f22d31c288eb7f8aa7d848866590487260351bb4801f42d776cfe5

Observation b57555d6-8a01-4aba-bf93-5214b165c6dc · outbound

This paper cites Time-Contrastive Networks: Self-Supervised Learning from Video.

FRAME: Pre-Training Video Feature Representations via Anticipation and Memory Time-Contrastive Networks: Self-Supervised Learning from Video

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T10:26:59.275835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:26:59.275835Z digest=sha256:734240790aec817b373c52d8c5dd58310c6b877225c62487ed317152b2a54516

Observation 0621d52f-41f0-41d0-9c51-022ea67875d6 · outbound

This paper cites Region-Based Representations Revisited.

FRAME: Pre-Training Video Feature Representations via Anticipation and Memory Region-Based Representations Revisited

Reference 44

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T10:27:01.941110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:26:59.295328Z digest=sha256:6fc153ac6c23e20206319cb200046b9d5d7cb63000f5eb53c286711220e26da9

Observation 1c7a520e-a0e3-412a-ba9e-825a4f371110 · outbound

This paper cites UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild.

FRAME: Pre-Training Video Feature Representations via Anticipation and Memory UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T10:26:59.394650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:26:59.394650Z digest=sha256:5b1590e658cc88166ab83199568990665d2657924b6f95c272285fb02abb568f

Observation 3a3ce278-b861-4ab0-a19b-46b89a94b4c6 · outbound

This paper cites VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training.

FRAME: Pre-Training Video Feature Representations via Anticipation and Memory VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T10:26:59.459159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:26:59.459159Z digest=sha256:ca9c3d33a990520cdd8a6587005651e3c527f49ff6c9269206ea45cfbb3b2ff1

Observation fb5b97ec-663e-4590-9d54-37f6a8ed3860 · outbound

This paper cites DINO-Tracker: Taming DINO for Self-Supervised Point Tracking in a Single Video.

FRAME: Pre-Training Video Feature Representations via Anticipation and Memory DINO-Tracker: Taming DINO for Self-Supervised Point Tracking in a Single Video

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-08-07T10:27:01.699803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:26:59.573734Z digest=sha256:a6d8bba892959e0d0fc4898d62b703b37c04b0359f6b1f11006155d1742dfd7b

Observation 2f9870ff-2b68-403d-8327-29aecad9838a · outbound

This paper cites Valmadre, L.

FRAME: Pre-Training Video Feature Representations via Anticipation and Memory Valmadre, L

Reference 48

Resolution
verified exact
doi, observed 2026-08-07T10:27:01.113263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:26:59.674156Z digest=sha256:c05c8d27e1a9d9456487baa105b6633ddc890817a3609c084ab1f900d11ae298

Observation f133875f-465d-4055-a12b-a1b2bf6f951e · outbound

This paper cites ActionCLIP: A New Paradigm for Video Action Recognition.

FRAME: Pre-Training Video Feature Representations via Anticipation and Memory ActionCLIP: A New Paradigm for Video Action Recognition

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T10:26:59.796244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:26:59.796244Z digest=sha256:0660a08d2eed4006af541b30066343bd84527b0afafec7d2638281e8b57ec750

Observation 2dc564fe-bf92-48f4-99b6-c24d41690e48 · outbound

This paper cites an unresolved cited work.

FRAME: Pre-Training Video Feature Representations via Anticipation and Memory Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:27:03.854216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:26:59.872388Z digest=sha256:c9c169f61c77388e8c3635aa187e47c39763a4147a87e61ee7621752266e9f04

Observation f4b16e50-ff0d-48d3-86ab-c1d5d1630fc5 · outbound

This paper cites an unresolved cited work.

FRAME: Pre-Training Video Feature Representations via Anticipation and Memory Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:27:03.635499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:26:59.957771Z digest=sha256:9659256ee12a5eaf396b5fa24d26e2ef75b180ee27500761db4ca2b32791d488

Observation 66795952-0924-49bf-8abd-8f5f9b65fd93 · outbound

This paper cites an unresolved cited work.

FRAME: Pre-Training Video Feature Representations via Anticipation and Memory Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:27:03.472517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:27:00.064803Z digest=sha256:8ec8c1cf7460d6ce46772d5089728c022f6f409780631e601c91fdf5fd3f0097

Observation 89638095-e7d9-4b61-9ddd-6e42162a3f61 · outbound

This paper cites Mask Propagation for Efficient Video Semantic Segmentation.

FRAME: Pre-Training Video Feature Representations via Anticipation and Memory Mask Propagation for Efficient Video Semantic Segmentation

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:00.201356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:00.201356Z digest=sha256:a4dfb72d5000cafc0fed2ca1b6a38605ce36de8fe19a3cc26a6ab31ebdc2e4c8

Observation ff2ce6b6-5e65-47c3-b5fe-22e9f8f83219 · outbound

This paper cites What Should Not Be Contrastive in Contrastive Learning.

FRAME: Pre-Training Video Feature Representations via Anticipation and Memory What Should Not Be Contrastive in Contrastive Learning

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:00.313883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:00.313883Z digest=sha256:81603254dda0a0dc5964dee590e9f878a9f54931fdff70c0f453f4d39d47e4b0

Observation fff358bb-25ac-4658-9b82-c9e02b7bb5a2 · outbound

This paper cites Rethinking Self-supervised Correspondence Learning: A Video Frame-level Similarity Perspective.

FRAME: Pre-Training Video Feature Representations via Anticipation and Memory Rethinking Self-supervised Correspondence Learning: A Video Frame-level Similarity Perspective

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-08-07T10:27:01.507949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:27:00.453266Z digest=sha256:86a5fba4aa4e5cecdb601811ac6a6b04be67926b4bf6415751a784feab1c8f21

Observation d28aa240-261e-4ce8-9022-5197386c229c · outbound

This paper cites PIDNet: A Real-time Semantic Segmentation Network Inspired by PID Controllers.

FRAME: Pre-Training Video Feature Representations via Anticipation and Memory PIDNet: A Real-time Semantic Segmentation Network Inspired by PID Controllers

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:00.565630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:00.565630Z digest=sha256:3121c0080b18c2ac695fcd728eb45428de1f4370e2f3601fc70f0c047d522f98

Observation c0f9be20-457d-4a22-923c-6fe834b27fb4 · outbound

This paper cites YouTube-VOS: A Large-Scale Video Object Segmentation Benchmark.

FRAME: Pre-Training Video Feature Representations via Anticipation and Memory YouTube-VOS: A Large-Scale Video Object Segmentation Benchmark

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:00.674386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:00.674386Z digest=sha256:2769e28bef9602c3144d6deba835d2e6c5d6e06e8f5fceb0284bb1b21ac38b0f

Observation f065b8b8-0aa2-4922-9686-68ad0db7794d · outbound

This paper cites DVIS++: Improved Decoupled Framework for Universal Video Segmentation.

FRAME: Pre-Training Video Feature Representations via Anticipation and Memory DVIS++: Improved Decoupled Framework for Universal Video Segmentation

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:00.801887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:00.801887Z digest=sha256:058e1a593148ae79f4676f6f48f5e37e443a1bc19f89a2651b368698cceb991b

Observation 932e4234-7db9-49b8-819a-e7ead73a5e3c · outbound

This paper cites Adaptive Temporal Encoding Network for Video Instance-level Human Parsing.

FRAME: Pre-Training Video Feature Representations via Anticipation and Memory Adaptive Temporal Encoding Network for Video Instance-level Human Parsing

Reference 59

Resolution
verified exact
local_arxiv, observed 2026-08-07T10:27:01.307432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:27:00.912867Z digest=sha256:7843ca30d875dd8d84bbdeb1611dcb2e8a9940232c6c717401ab8cfdcb5cd6f2

Observation d678e2af-5d5f-4479-b05a-e88dd39b194c · outbound

This paper cites Masked Autoencoders Are Scalable Vision Learners.

FRAME: Pre-Training Video Feature Representations via Anticipation and Memory Masked Autoencoders Are Scalable Vision Learners

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T10:26:57.705213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:26:57.705213Z digest=sha256:3780634435734efc622014e64baa48ba22999bf3ddc5bcc91dc5ebdc06fa80f7

Pith citing papers

No inbound Pith citation observations are available.