Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-26T08:50:10.781971Z
Paper Citation Record · LEDGER
As of 6 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 0 inbound Pith citation observations for arXiv:2606.23256.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-26T08:50:10.781971Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
54 of 54 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 30762952-72a7-48bd-bc06-7c3a8b3e541d · outbound
P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture Vivit: A video vision transformer
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06fe4d81-9776-4d8d-9179-cfaee721a028 · outbound
P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture Hiervl: Learning hierarchical video-language embeddings
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eaf34c33-27d3-4183-832f-2c3259b86b69 · outbound
P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3da0e6e3-d6c1-474f-85d2-583c6f207800 · outbound
P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture How much temporal long-term context is needed for action segmentation? InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 10351–10361, 2023
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23d89db5-1e21-4915-aeb7-475487c61feb · outbound
P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture My view is the best view: Procedure learning from egocentric videos
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0c1c981-328a-4bf3-9c43-b96cdbc7e24b · outbound
P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture United we stand, divided we fall: Unitygraph for unsupervised procedure learning from videos
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8b411a7-43f5-43b1-9a15-60e3436c6fa6 · outbound
P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture Revisiting Feature Prediction for Learning Visual Representations from Video
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8b162020-9d2a-4d11-bcfd-b631d8e2a47f · outbound
P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture Unified fully and timestamp supervised temporal action segmentation via sequence to sequence translation
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6d52c73-000b-4da7-bb27-051020c1db05 · outbound
P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture Unresolved cited work
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7bbae96d-5dc0-4b67-acff-61f53883c12f · outbound
P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture Quo vadis, action recognition? a new model and the kinetics dataset
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88ba19b6-fd84-4cec-b005-5f0d3c09f131 · outbound
P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture Streaming videollms for real-time procedural video understanding
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation adfd7a1b-c00b-40d6-b48d-d34f0bb6a259 · outbound
P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture arXiv preprint arXiv:2512.10942 (2025)
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 71aa14d4-081b-4f1c-aa51-6904dfc229d1 · outbound
P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture Videollm-online: Online video large language model for streaming video
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5364542d-b0f3-43cd-97e7-36697146a388 · outbound
P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00e27043-8aea-4f77-9b04-eed8cf749986 · outbound
P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture Unsupervised procedure learning via joint dynamic summarization
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df23605c-390b-4893-b90b-c47ee46ae557 · outbound
P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture Multiscale vision transformers
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11e5cc62-821d-420f-b000-10e97d582ceb · outbound
P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture Ms-tcn: Multi-stage temporal convolutional network for action segmentation
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a446f45-2d4b-4628-a78d-1ca0422f4703 · outbound
P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b034fb5-208c-4a7a-8c5f-5e15fc5b790e · outbound
P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7018eb6-96e9-4cf5-bf6e-6fef59e1f2cc · outbound
P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture Ma-lmm: Memory-augmented large multimodal model for long-term video understanding
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b061230b-40aa-434d-a92a-4c5602cf849b · outbound
P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture What Changed and What Could Have Changed? State-Change Counterfactuals for Procedure-Aware Video Representation Learning
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 7d9afdd9-8058-4dfc-bf36-35fc5bb7caed · outbound
P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture Video token merging for long video understanding.Advances in Neural Information Processing Systems, 37:13851–13871, 2024
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0162f33f-b6be-4529-becb-9eabebc5d2fa · outbound
P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture Temporal Reasoning Transfer from Text to Video
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5bbda93a-6220-42db-bad6-88ebc778d723 · outbound
P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture IEEE Trans- actions on Pattern Analysis and Machine Intelligence pp
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f8fab75e-9cfe-40ad-9448-f9d310d6324c · outbound
P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture Tsm: Temporal shift module for efficient video understanding
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f976101-5c0b-4d29-a98c-06ede3757fb5 · outbound
P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture TempCompass: Do Video LLMs Really Understand Videos?
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1b79aa2e-9f30-45db-92e4-78d5dd7a1f0b · outbound
P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture Video swin transformer
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a546a6aa-27a6-4418-b6c9-32ae967c7b36 · outbound
P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture Fact: Frame-action cross-attention temporal modeling for efficient action segmentation
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fead2f64-dc68-4440-8ad1-adfa1389bf70 · outbound
P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture Streamer: Streaming representation learning and event segmentation in a hierarchical manner.Advances in Neural Information Processing Systems, 36: 45694–45715, 2023
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9476243b-aa54-45e4-85f3-ef0d67dc355f · outbound
P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 83678f8d-02ef-4142-8e04-514bee9b5d4d · outbound
P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture Video transformer network
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c11bbe3b-a71b-4b30-b643-19820b9c79e8 · outbound
P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture Representation Learning with Contrastive Predictive Coding
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c51aedc1-5ff2-4ddc-bc71-e78bd51f5680 · outbound
P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture Hiero: understanding the hierarchy of human behavior enhances reasoning on egocentric videos
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ce97aa0-80c5-4523-80a0-8eaf7a7406b0 · outbound
P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture Omnia de egotempo: Benchmarking temporal understanding of multi-modal llms in egocentric videos
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd01ca72-9e9d-4e0a-822e-3a5f0a10473f · outbound
P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture Egovlpv2: Egocentric video-language pre-training with fusion in the backbone
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6bca6163-4a9e-44c0-9974-be842e71cff4 · outbound
P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture Learning from untrimmed videos: Self-supervised video representation learning with hierarchical consistency
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a58456e-931f-4ce1-a837-dae08856644d · outbound
P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture Understanding Long Videos with Multimodal Language Models
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation bc406de2-06fd-4ddc-bae5-ee79b19bb3d7 · outbound
P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture Timechat: A time-sensitive multimodal large language model for long video understanding
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef084bf9-244c-42c1-98ff-a33632dfc924 · outbound
P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture Assembly101: A large-scale multi-view video dataset for understanding procedural activities
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3931da6f-60a2-49af-9c11-1c12b3e661a4 · outbound
P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture Video-xl: Extra-long vision language model for hour-scale video understanding
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d22fa270-00bf-4df9-9cdb-f4902d7533be · outbound
P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture C2f-tcn: A framework for semi-and fully-supervised temporal action segmentation.IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(10): 11484–11501, 2023
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46b509a0-3dd0-476b-807d-d1fe345c9914 · outbound
P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture Moviechat: From dense token to sparse memory for long video understanding
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 123104cd-b2ce-440a-8f35-f9ac36f28727 · outbound
P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture Moviechat+: Question- aware sparse memory for long video question answering.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2b77dd9-732d-48b1-90ab-5559d9311043 · outbound
P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063, 2024
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df19cd4c-1f83-4b27-97c6-3020bc27712c · outbound
P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture Koala: Key frame-conditioned long video-llm
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a501836b-12c1-4f49-8374-bbf318a41535 · outbound
P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training.Advances in neural information processing systems, 35: 10078–10093, 2022
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7d9282e-2770-4d41-8c3b-578a1d32c1c9 · outbound
P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d6c02184-a138-4520-95a1-de0289740a77 · outbound
P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture Longvideobench: A benchmark for long-context interleaved video-language understanding.Advances in Neural Information Processing Systems, 37: 28828–28857, 2024
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c7eca5d-c4c6-47f5-abb5-08d848f2ad82 · outbound
P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture Videollm-mod: Efficient video-language streaming with mixture- of-depths vision computation.Advances in Neural Information Processing Systems, 37:109922–109947, 2024
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation beb185d8-738a-44b3-a01f-5ef2ba2254cb · outbound
P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture Hierarchical self-supervised representation learning for movie understanding
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f24f1a1-63c1-4daf-b518-729a621678e3 · outbound
P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture ASFormer: Transformer for Action Segmentation
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 77c22baa-aadc-4027-8c66-000783c501fc · outbound
P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture Learning procedure-aware video representation from instructional videos and their narrations
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b063b89b-3244-4784-9e16-80dc3c6b9334 · outbound
P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture Procedure-aware pretraining for instructional video understanding
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d1a366b-694e-4372-a3c8-d2394234f837 · outbound
P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture Unresolved cited work
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
No inbound Pith citation observations are available.