Pith. sign in

Paper Citation Record · LEDGER

P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture

As of 6 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 0 inbound Pith citation observations for arXiv:2606.23256.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.23256 v1

Coverage vector

measured 54 of 54 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-26T08:50:10.781971Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

54 of 54 outbound references displayed

  • verified exact13
  • verified fuzzy0
  • unresolved41
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 30762952-72a7-48bd-bc06-7c3a8b3e541d · outbound

This paper cites Vivit: A video vision transformer.

P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture Vivit: A video vision transformer

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-26T08:50:10.781971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:50:10.781971Z digest=sha256:a0064d002b1e22076ad615c9943a0be4d4d4cca830ce845c3a70fc6fb3a12e25

Observation 06fe4d81-9776-4d8d-9179-cfaee721a028 · outbound

This paper cites Hiervl: Learning hierarchical video-language embeddings.

P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture Hiervl: Learning hierarchical video-language embeddings

Reference 2

Resolution
unresolved
no resolver link, observed 2026-06-26T08:50:10.781971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:50:10.781971Z digest=sha256:5e725b092762e9ce1292af040079b517040c8add4f95aebc7c7c1fa7a0d66668

Observation eaf34c33-27d3-4183-832f-2c3259b86b69 · outbound

This paper cites V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning.

P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-04T10:29:44.920591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T08:50:10.781971Z digest=sha256:4ef8e818edd34ca225716b707b4addd435061a719c4b1077334ef610c5bd7c78

Observation 3da0e6e3-d6c1-474f-85d2-583c6f207800 · outbound

This paper cites How much temporal long-term context is needed for action segmentation? InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 10351–10361, 2023.

P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture How much temporal long-term context is needed for action segmentation? InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 10351–10361, 2023

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-26T08:50:10.781971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:50:10.781971Z digest=sha256:221ee9bc029340d42981ef46d98e837a2c0bac20fad59062a02778f52130f449

Observation 23d89db5-1e21-4915-aeb7-475487c61feb · outbound

This paper cites My view is the best view: Procedure learning from egocentric videos.

P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture My view is the best view: Procedure learning from egocentric videos

Reference 5

Resolution
unresolved
no resolver link, observed 2026-06-26T08:50:10.781971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:50:10.781971Z digest=sha256:336d287e58719a41dc9f96f54d124282cab274085f802b130673daff1c9c248e

Observation b0c1c981-328a-4bf3-9c43-b96cdbc7e24b · outbound

This paper cites United we stand, divided we fall: Unitygraph for unsupervised procedure learning from videos.

P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture United we stand, divided we fall: Unitygraph for unsupervised procedure learning from videos

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-26T08:50:10.781971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:50:10.781971Z digest=sha256:978595af6ac5b09250ee4b15b02bc2d9ab0ac8df76ab10b9829b1783b000785c

Observation e8b411a7-43f5-43b1-9a15-60e3436c6fa6 · outbound

This paper cites Revisiting Feature Prediction for Learning Visual Representations from Video.

P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture Revisiting Feature Prediction for Learning Visual Representations from Video

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-07-04T10:29:44.931885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T08:50:10.781971Z digest=sha256:34766f8555fc015d03f3a6037b5de2a0b153b06ea2dc099a314b55e439340282

Observation 8b162020-9d2a-4d11-bcfd-b631d8e2a47f · outbound

This paper cites Unified fully and timestamp supervised temporal action segmentation via sequence to sequence translation.

P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture Unified fully and timestamp supervised temporal action segmentation via sequence to sequence translation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-26T08:50:10.781971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:50:10.781971Z digest=sha256:5f24fe0e96b721d1520c98b371ecafa56aef6bf958a30a635286f5dda57f54e1

Observation f6d52c73-000b-4da7-bb27-051020c1db05 · outbound

This paper cites an unresolved cited work.

P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-26T08:50:10.781971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:50:10.781971Z digest=sha256:132fd92662a5e6ff4f043cd3ce47da4f070345e29a4d00a0ec71482130cbf009

Observation 7bbae96d-5dc0-4b67-acff-61f53883c12f · outbound

This paper cites Quo vadis, action recognition? a new model and the kinetics dataset.

P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture Quo vadis, action recognition? a new model and the kinetics dataset

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-26T08:50:10.781971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:50:10.781971Z digest=sha256:632eaf14ff9609ac41887d88af006602d6f29f0d77fa6551ae3b9c58de08f0bc

Observation 88ba19b6-fd84-4cec-b005-5f0d3c09f131 · outbound

This paper cites Streaming videollms for real-time procedural video understanding.

P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture Streaming videollms for real-time procedural video understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-26T08:50:10.781971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:50:10.781971Z digest=sha256:c5c731dd07ea42bdf7347feae7ab31d16e92e2a64bb116411a203fded61f309c

Observation adfd7a1b-c00b-40d6-b48d-d34f0bb6a259 · outbound

This paper cites arXiv preprint arXiv:2512.10942 (2025).

P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture arXiv preprint arXiv:2512.10942 (2025)

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:29:44.907216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T08:50:10.781971Z digest=sha256:e182fc4f2143df9962d3e552a29c5815f7bb3b4d610ea3fd5aea238503ca72d9

Observation 71aa14d4-081b-4f1c-aa51-6904dfc229d1 · outbound

This paper cites Videollm-online: Online video large language model for streaming video.

P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture Videollm-online: Online video large language model for streaming video

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-26T08:50:10.781971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:50:10.781971Z digest=sha256:782231e0907250ea21fc5cf3987c565aab4c2872b5a4cbd49d72c5b6c5dfcc73

Observation 5364542d-b0f3-43cd-97e7-36697146a388 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-26T08:50:10.781971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:50:10.781971Z digest=sha256:402a29cf21a9aceb8ae1f8e73043aa4ef9849034cdba594511a2c7a80f245a98

Observation 00e27043-8aea-4f77-9b04-eed8cf749986 · outbound

This paper cites Unsupervised procedure learning via joint dynamic summarization.

P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture Unsupervised procedure learning via joint dynamic summarization

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-26T08:50:10.781971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:50:10.781971Z digest=sha256:0584a41e17b37ab3d1e4a689232bb51812287beb411f9f498b81a05e9cbeea19

Observation df23605c-390b-4893-b90b-c47ee46ae557 · outbound

This paper cites Multiscale vision transformers.

P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture Multiscale vision transformers

Reference 16

Resolution
unresolved
no resolver link, observed 2026-06-26T08:50:10.781971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:50:10.781971Z digest=sha256:cedb4b302f3ba5ebfa55b7d7b9a6f1bccdf8cc3c8ff2b9bb9854586a19a8074c

Observation 11e5cc62-821d-420f-b000-10e97d582ceb · outbound

This paper cites Ms-tcn: Multi-stage temporal convolutional network for action segmentation.

P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture Ms-tcn: Multi-stage temporal convolutional network for action segmentation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-26T08:50:10.781971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:50:10.781971Z digest=sha256:d99d8ed9c6241de0b28f174f653ee56178bbf171c227e442659039430b2537a3

Observation 3a446f45-2d4b-4628-a78d-1ca0422f4703 · outbound

This paper cites Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis.

P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-26T08:50:10.781971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:50:10.781971Z digest=sha256:25e352170e5e946470f6c3c3605c7952e02c0f1e457fc302996ea9e5e2d8aa82

Observation 1b034fb5-208c-4a7a-8c5f-5e15fc5b790e · outbound

This paper cites Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives.

P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-26T08:50:10.781971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:50:10.781971Z digest=sha256:a14047e9f3ecea3e3a8e2c74b605faced889a6fa4d66593cdf6037417a43f662

Observation d7018eb6-96e9-4cf5-bf6e-6fef59e1f2cc · outbound

This paper cites Ma-lmm: Memory-augmented large multimodal model for long-term video understanding.

P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture Ma-lmm: Memory-augmented large multimodal model for long-term video understanding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-26T08:50:10.781971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:50:10.781971Z digest=sha256:4aebc5d19d163842fda0133adbdabad26c0ddfd64ffa38bbaf2de511b836763a

Observation b061230b-40aa-434d-a92a-4c5602cf849b · outbound

This paper cites What Changed and What Could Have Changed? State-Change Counterfactuals for Procedure-Aware Video Representation Learning.

P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture What Changed and What Could Have Changed? State-Change Counterfactuals for Procedure-Aware Video Representation Learning

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:29:44.910139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T08:50:10.781971Z digest=sha256:375cb0158c2ed716fb36f52b013e1515f8501db5a4ceb85347403ce38c5c9a0c

Observation 7d9afdd9-8058-4dfc-bf36-35fc5bb7caed · outbound

This paper cites Video token merging for long video understanding.Advances in Neural Information Processing Systems, 37:13851–13871, 2024.

P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture Video token merging for long video understanding.Advances in Neural Information Processing Systems, 37:13851–13871, 2024

Reference 22

Resolution
unresolved
no resolver link, observed 2026-06-26T08:50:10.781971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:50:10.781971Z digest=sha256:d755ed617e651a380885c026fa3ee91a77a9c7955593f23828254b91b334f403

Observation 0162f33f-b6be-4529-becb-9eabebc5d2fa · outbound

This paper cites Temporal Reasoning Transfer from Text to Video.

P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture Temporal Reasoning Transfer from Text to Video

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:29:44.917793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T08:50:10.781971Z digest=sha256:b2e37fd9dfdc791d2952173761ed1c133562546c3fc62050726e7a21dd44243b

Observation 5bbda93a-6220-42db-bad6-88ebc778d723 · outbound

This paper cites IEEE Trans- actions on Pattern Analysis and Machine Intelligence pp.

P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture IEEE Trans- actions on Pattern Analysis and Machine Intelligence pp

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-06-26T08:59:15.355133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T08:50:10.781971Z digest=sha256:f764cb861134ac2079dbae2da5375abcd207666ecad648ce4860475b08c8d370

Observation f8fab75e-9cfe-40ad-9448-f9d310d6324c · outbound

This paper cites Tsm: Temporal shift module for efficient video understanding.

P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture Tsm: Temporal shift module for efficient video understanding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-06-26T08:50:10.781971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:50:10.781971Z digest=sha256:42f33c935fcdbe99a655124990f1d7291498cd46835b0521885244c832a9dfe5

Observation 9f976101-5c0b-4d29-a98c-06ede3757fb5 · outbound

This paper cites TempCompass: Do Video LLMs Really Understand Videos?.

P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture TempCompass: Do Video LLMs Really Understand Videos?

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-07-04T10:29:44.912158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T08:50:10.781971Z digest=sha256:7cfbe8d140844a6e494b6935d173d947d1bde4c9af2d75613a5e8fa8d4310803

Observation 1b79aa2e-9f30-45db-92e4-78d5dd7a1f0b · outbound

This paper cites Video swin transformer.

P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture Video swin transformer

Reference 27

Resolution
unresolved
no resolver link, observed 2026-06-26T08:50:10.781971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:50:10.781971Z digest=sha256:43477d45ccd0113b767705a49b9ccabd298cb246c6571a44cb5ea6b9c0e882e8

Observation a546a6aa-27a6-4418-b6c9-32ae967c7b36 · outbound

This paper cites Fact: Frame-action cross-attention temporal modeling for efficient action segmentation.

P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture Fact: Frame-action cross-attention temporal modeling for efficient action segmentation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-06-26T08:50:10.781971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:50:10.781971Z digest=sha256:5c01240e0ac247d57a19ab3f775c9a1aa5ab856229ae334ccc7965464f719408

Observation fead2f64-dc68-4440-8ad1-adfa1389bf70 · outbound

This paper cites Streamer: Streaming representation learning and event segmentation in a hierarchical manner.Advances in Neural Information Processing Systems, 36: 45694–45715, 2023.

P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture Streamer: Streaming representation learning and event segmentation in a hierarchical manner.Advances in Neural Information Processing Systems, 36: 45694–45715, 2023

Reference 29

Resolution
unresolved
no resolver link, observed 2026-06-26T08:50:10.781971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:50:10.781971Z digest=sha256:9dbb66978ed7172c40ee906f1475dac37bb70d3adf77164985cd8f5eeb8b1a89

Observation 9476243b-aa54-45e4-85f3-ef0d67dc355f · outbound

This paper cites V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning.

P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-07-04T10:29:44.926386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T08:50:10.781971Z digest=sha256:018c257dfd8a2feefa5e15976d2242db4328d04eb646dffb7f9214245187c75d

Observation 83678f8d-02ef-4142-8e04-514bee9b5d4d · outbound

This paper cites Video transformer network.

P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture Video transformer network

Reference 31

Resolution
unresolved
no resolver link, observed 2026-06-26T08:50:10.781971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:50:10.781971Z digest=sha256:b48443a064021c7115456c680459f43a1e4b402098ff1e8e1cdb1528c667edbe

Observation c11bbe3b-a71b-4b30-b643-19820b9c79e8 · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture Representation Learning with Contrastive Predictive Coding

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-07-04T10:29:44.912567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T08:50:10.781971Z digest=sha256:74a89d022cb9a453697e024ac584d3411461a70d291af51856b8e248e49aa0c4

Observation c51aedc1-5ff2-4ddc-bc71-e78bd51f5680 · outbound

This paper cites Hiero: understanding the hierarchy of human behavior enhances reasoning on egocentric videos.

P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture Hiero: understanding the hierarchy of human behavior enhances reasoning on egocentric videos

Reference 33

Resolution
unresolved
no resolver link, observed 2026-06-26T08:50:10.781971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:50:10.781971Z digest=sha256:c5b55792907cec2bb40d01354644a016546f69816ee04b77efb1f9d4b2bcda15

Observation 4ce97aa0-80c5-4523-80a0-8eaf7a7406b0 · outbound

This paper cites Omnia de egotempo: Benchmarking temporal understanding of multi-modal llms in egocentric videos.

P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture Omnia de egotempo: Benchmarking temporal understanding of multi-modal llms in egocentric videos

Reference 34

Resolution
unresolved
no resolver link, observed 2026-06-26T08:50:10.781971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:50:10.781971Z digest=sha256:736ff84554fa2935d9fce713e6dcbfe27928a2a409b888b2f783357ec37f0942

Observation cd01ca72-9e9d-4e0a-822e-3a5f0a10473f · outbound

This paper cites Egovlpv2: Egocentric video-language pre-training with fusion in the backbone.

P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture Egovlpv2: Egocentric video-language pre-training with fusion in the backbone

Reference 35

Resolution
unresolved
no resolver link, observed 2026-06-26T08:50:10.781971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:50:10.781971Z digest=sha256:2294a0c02545f571df6cb72cbbef91df4b3bc6ae890041a3772a670299d8acbc

Observation 6bca6163-4a9e-44c0-9974-be842e71cff4 · outbound

This paper cites Learning from untrimmed videos: Self-supervised video representation learning with hierarchical consistency.

P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture Learning from untrimmed videos: Self-supervised video representation learning with hierarchical consistency

Reference 36

Resolution
unresolved
no resolver link, observed 2026-06-26T08:50:10.781971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:50:10.781971Z digest=sha256:ecf10fc672348b98d2b74b97d546a4f84f14ab7b3816c6257e1a16857b8d44f8

Observation 0a58456e-931f-4ce1-a837-dae08856644d · outbound

This paper cites Understanding Long Videos with Multimodal Language Models.

P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture Understanding Long Videos with Multimodal Language Models

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:29:44.915306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T08:50:10.781971Z digest=sha256:dcf3de14a3da3e7ddb64c741b5043070e37a0f978aa41c9f136e18971c7e73ee

Observation bc406de2-06fd-4ddc-bae5-ee79b19bb3d7 · outbound

This paper cites Timechat: A time-sensitive multimodal large language model for long video understanding.

P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture Timechat: A time-sensitive multimodal large language model for long video understanding

Reference 38

Resolution
unresolved
no resolver link, observed 2026-06-26T08:50:10.781971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:50:10.781971Z digest=sha256:53bfb41bee4495d14e68b4b9a71cc2695a1183050c0032ad4e67af8bbe716591

Observation ef084bf9-244c-42c1-98ff-a33632dfc924 · outbound

This paper cites Assembly101: A large-scale multi-view video dataset for understanding procedural activities.

P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture Assembly101: A large-scale multi-view video dataset for understanding procedural activities

Reference 39

Resolution
unresolved
no resolver link, observed 2026-06-26T08:50:10.781971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:50:10.781971Z digest=sha256:fe9369c5b3d379071e51d6413da75c545bb2a5d3b77b3f8508a2d0f3b6a96620

Observation 3931da6f-60a2-49af-9c11-1c12b3e661a4 · outbound

This paper cites Video-xl: Extra-long vision language model for hour-scale video understanding.

P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture Video-xl: Extra-long vision language model for hour-scale video understanding

Reference 40

Resolution
unresolved
no resolver link, observed 2026-06-26T08:50:10.781971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:50:10.781971Z digest=sha256:b6f677ee7011f376cc1774150991bc0c62dfffe5118ef67ce29c742949e0d6f4

Observation d22fa270-00bf-4df9-9cdb-f4902d7533be · outbound

This paper cites C2f-tcn: A framework for semi-and fully-supervised temporal action segmentation.IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(10): 11484–11501, 2023.

P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture C2f-tcn: A framework for semi-and fully-supervised temporal action segmentation.IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(10): 11484–11501, 2023

Reference 41

Resolution
unresolved
no resolver link, observed 2026-06-26T08:50:10.781971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:50:10.781971Z digest=sha256:a12fd911ca2a9e71981e8577cf98d38c4e167752b03a098bf1a5afd014683e8f

Observation 46b509a0-3dd0-476b-807d-d1fe345c9914 · outbound

This paper cites Moviechat: From dense token to sparse memory for long video understanding.

P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture Moviechat: From dense token to sparse memory for long video understanding

Reference 42

Resolution
unresolved
no resolver link, observed 2026-06-26T08:50:10.781971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:50:10.781971Z digest=sha256:3e38bfdc9c044a6905fa77851a86d1351ceb621c51138383cc9b466113d29504

Observation 123104cd-b2ce-440a-8f35-f9ac36f28727 · outbound

This paper cites Moviechat+: Question- aware sparse memory for long video question answering.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025.

P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture Moviechat+: Question- aware sparse memory for long video question answering.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025

Reference 43

Resolution
unresolved
no resolver link, observed 2026-06-26T08:50:10.781971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:50:10.781971Z digest=sha256:fdd285ddc0090434eaab1bdabfe2f21b06eea5b0859e59eccf986bf9cb66e547

Observation b2b77dd9-732d-48b1-90ab-5559d9311043 · outbound

This paper cites Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063, 2024.

P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063, 2024

Reference 44

Resolution
unresolved
no resolver link, observed 2026-06-26T08:50:10.781971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:50:10.781971Z digest=sha256:0767971c63248bfe74717681e1ba96f8a48ae876d26a24ba1b7cb2f7e07dd349

Observation df19cd4c-1f83-4b27-97c6-3020bc27712c · outbound

This paper cites Koala: Key frame-conditioned long video-llm.

P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture Koala: Key frame-conditioned long video-llm

Reference 45

Resolution
unresolved
no resolver link, observed 2026-06-26T08:50:10.781971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:50:10.781971Z digest=sha256:996a47d3828be3c3e34df7019725120d6f682f92fd1719c1a5ecc818a654b118

Observation a501836b-12c1-4f49-8374-bbf318a41535 · outbound

This paper cites Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training.Advances in neural information processing systems, 35: 10078–10093, 2022.

P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training.Advances in neural information processing systems, 35: 10078–10093, 2022

Reference 46

Resolution
unresolved
no resolver link, observed 2026-06-26T08:50:10.781971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:50:10.781971Z digest=sha256:7bc4ffdd001e33699485cce2d5b91b0a00438a7354dd7766d846f40ed6391027

Observation a7d9282e-2770-4d41-8c3b-578a1d32c1c9 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-07-04T10:29:44.917802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T08:50:10.781971Z digest=sha256:8b6056f9703ed5f1104e524aef6031abbd99afbe7d880cba80c98398e12796f9

Observation d6c02184-a138-4520-95a1-de0289740a77 · outbound

This paper cites Longvideobench: A benchmark for long-context interleaved video-language understanding.Advances in Neural Information Processing Systems, 37: 28828–28857, 2024.

P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture Longvideobench: A benchmark for long-context interleaved video-language understanding.Advances in Neural Information Processing Systems, 37: 28828–28857, 2024

Reference 48

Resolution
unresolved
no resolver link, observed 2026-06-26T08:50:10.781971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:50:10.781971Z digest=sha256:9e9e6f46915722c0eecadecac050ee209cb7cb9025ea36a006b8cb79bd60bfcc

Observation 0c7eca5d-c4c6-47f5-abb5-08d848f2ad82 · outbound

This paper cites Videollm-mod: Efficient video-language streaming with mixture- of-depths vision computation.Advances in Neural Information Processing Systems, 37:109922–109947, 2024.

P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture Videollm-mod: Efficient video-language streaming with mixture- of-depths vision computation.Advances in Neural Information Processing Systems, 37:109922–109947, 2024

Reference 49

Resolution
unresolved
no resolver link, observed 2026-06-26T08:50:10.781971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:50:10.781971Z digest=sha256:1b93af74b037dac2aabdd2fa3ffb3797ad905beb4bb6c8d5d2a44dbed9bae7d3

Observation beb185d8-738a-44b3-a01f-5ef2ba2254cb · outbound

This paper cites Hierarchical self-supervised representation learning for movie understanding.

P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture Hierarchical self-supervised representation learning for movie understanding

Reference 50

Resolution
unresolved
no resolver link, observed 2026-06-26T08:50:10.781971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:50:10.781971Z digest=sha256:ce4362e7ab3359b2fb8791b4f33df7ca75e8a2d6972e0dcf23855f8b1c68332b

Observation 4f24f1a1-63c1-4daf-b518-729a621678e3 · outbound

This paper cites ASFormer: Transformer for Action Segmentation.

P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture ASFormer: Transformer for Action Segmentation

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:29:44.923552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T08:50:10.781971Z digest=sha256:e90bb1f603cf197cc4ed915c71f9587202a76d0e45e162432da17ed1ff16c67f

Observation 77c22baa-aadc-4027-8c66-000783c501fc · outbound

This paper cites Learning procedure-aware video representation from instructional videos and their narrations.

P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture Learning procedure-aware video representation from instructional videos and their narrations

Reference 52

Resolution
unresolved
no resolver link, observed 2026-06-26T08:50:10.781971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:50:10.781971Z digest=sha256:6e3d9f12bc44618babd6ccd3b7818ecb493e854395737a2fbc77c28a046b67eb

Observation b063b89b-3244-4784-9e16-80dc3c6b9334 · outbound

This paper cites Procedure-aware pretraining for instructional video understanding.

P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture Procedure-aware pretraining for instructional video understanding

Reference 53

Resolution
unresolved
no resolver link, observed 2026-06-26T08:50:10.781971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:50:10.781971Z digest=sha256:f0c0be15222ec2802b9662a6f226dcb7c394a9823f1171a52dfb0c20ce380458

Observation 7d1a366b-694e-4372-a3c8-d2394234f837 · outbound

This paper cites an unresolved cited work.

P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture Unresolved cited work

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:29:44.929335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T08:50:10.781971Z digest=sha256:12d6ae1013e883a90ca950c5c62b0271a093f4763c70db95cadf45d650471be2

Pith citing papers

No inbound Pith citation observations are available.