Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T00:26:05.265658Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 1 inbound Pith citation observation for arXiv:2506.14356.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T00:26:05.265658Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-30T21:29:27.063028Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-06-30T21:35:04.650377Z
54 of 54 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 3b3039c5-4f63-43b2-96a1-d4c4fc357c19 · outbound
EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Quo vadis, action recognition? a new model and the kinetics dataset,
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3b05827e-9bcb-4504-b918-91617a462caa · outbound
EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Video summarization through reinforcement learning with a 3d spatio- temporal u-net,
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b2e2b62d-3ff3-42f4-b510-693cd4de7b11 · outbound
EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Deep attention network for egocentric action recognition,
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 416ae5dc-9eea-4ae5-8c37-38a17c72361d · outbound
EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Training a Large Video Model on a Single Machine in a Day
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c5d6b164-6f58-4f03-a66e-b3565b0bde8d · outbound
EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c6c2b82-1304-4e98-8301-7975bacc3446 · outbound
EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization InternVideo2: Scaling Foundation Models for Multimodal Video Understanding
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2d5163b-29bd-4182-9040-32cfece62edd · outbound
EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization EgoVLPv2: Egocentric video-language pre-training with fusion in the backbone,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 82d49f24-f819-402c-81c9-472ae702a1bd · outbound
EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Egocentric video-language pretraining,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 144cc8dd-9c57-48e0-8412-8e9a04012dba · outbound
EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Improving semantic video retrieval models by training with a relevance-aware online mining strategy,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 04357bdc-9f02-452b-ae9e-54018f9be793 · outbound
EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Learning video representations from large language models,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f9ec9a56-fbec-4170-bf4d-b6170093b1b6 · outbound
EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization EgoVideo: Exploring Egocentric Foundation Model and Downstream Adaptation
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93e8b51b-8004-400e-9a20-d4c80012e913 · outbound
EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Videomae: Masked autoen- coders are data-efficient learners for self-supervised video pre-training,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3aa1e723-af16-4d5f-8aa8-4345f828d67c · outbound
EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4cda262c-684a-4f56-9402-0d6eabd98073 · outbound
EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Internvid: A large-scale video-text dataset for multimodal understanding and generation,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 69ab1970-654d-4c41-a237-ef7ac463104c · outbound
EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7fbc78e4-1606-4303-b25e-ece1889a162a · outbound
EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7176760f-561d-43fe-bc91-9b9f7c99fb15 · outbound
EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c2c0630-67ea-4f02-96fb-a8e0cd32a142 · outbound
EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization EVA-02: A Visual Representation for Neon Genesis
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45c01c0e-2c0c-4446-ae29-20e8d6fed1c2 · outbound
EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Multi- similarity loss with general pair weighting for deep metric learning,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ab8d33a3-8344-4b78-99e6-7feda72268a3 · outbound
EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Ego4d: Around the world in 3,000 hours of egocentric video,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7aed2aee-2f71-4fe6-bcdd-308d381bc970 · outbound
EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Rescaling egocentric vision: Collection, pipeline and challenges for epic-kitchens- 100,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7ba327eb-6879-4ede-adc7-5b81d3a058b7 · outbound
EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Scaling egocentric vision: The epic-kitchens dataset,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f8f59c61-0302-4d25-aa0e-c79e170886c6 · outbound
EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Charades-Ego: A Large-Scale Dataset of Paired Third and First Person Videos
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c7b3501-3185-4670-a377-336c7b19a095 · outbound
EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Long short-term memory,
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f86341f-085c-4b6c-9496-678cb767245d · outbound
EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Is space-time attention all you need for video understanding?
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8a3e82d7-031b-4fea-b2b7-9e35cac358e4 · outbound
EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Frozen in time: A joint video and image encoder for end-to-end retrieval,
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df4130a7-8ddd-434b-a195-251c34c4c424 · outbound
EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization VideoMAE V2: Scaling video masked autoencoders with dual masking,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 47013c40-a70d-4d51-b498-12e7e211c017 · outbound
EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Flamingo: a visual language model for few-shot learning,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 46fab9ae-e11f-4030-a325-dc9b400980ad · outbound
EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Roformer: En- hanced transformer with rotary position embedding,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d3b604c5-7e92-45b2-bcb2-12430e34096d · outbound
EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72d0da19-90bd-49a7-bad9-1dce2fcd7bc2 · outbound
EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization VideoRoPE: What Makes for Good Video Rotary Position Embedding?
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1db5d195-264f-4332-b8ae-69d2eef1df4c · outbound
EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Supervised contrastive learn- ing,
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a6dcc2cf-ea91-4430-a219-604ca220f721 · outbound
EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Parameter-free deep multi-modal clustering with reliable contrastive learning,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 34495b16-3787-4229-aac2-8c1013cfe167 · outbound
EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Cross-modal contrastive learning network for few-shot action recognition,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2423d424-3315-4b21-a30c-f5c861e37277 · outbound
EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Representation Learning with Contrastive Predictive Coding
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30bce6e3-7aa2-4791-a3b7-8242b8c22b93 · outbound
EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization End-to-end learning of visual representations from uncurated instructional videos,
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dd7dbc9d-7868-4a71-ae29-fec6e005ca2f · outbound
EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Facenet: A unified embed- ding for face recognition and clustering,
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5858fc87-7a64-427c-8795-10f7672704a3 · outbound
EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Circle loss: A unified perspective of pair similarity optimization,
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 49ed26cd-60af-4674-93e2-0eb323b8581d · outbound
EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Relevance-based margin for contrastively-trained video retrieval mod- els,
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 677b1418-3e49-410a-a9de-4e81b4db830a · outbound
EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Masked video distillation: Rethinking masked feature modeling for self-supervised video representation learning,
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2f8337f0-1222-4470-8bbf-3b7e10e8222a · outbound
EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Fine-grained action retrieval through multiple parts-of-speech embeddings,
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation aa116c9e-61d8-4509-b642-0b598643aea5 · outbound
EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization On semantic similarity in video retrieval,
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 431349d9-3713-4553-afbb-2d7961ed415d · outbound
EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Egocentric Video-Language Pretraining @ EPIC-KITCHENS-100 Multi-Instance Retrieval Challenge 2022
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d653bc96-b238-49b9-bc71-1b6c966c223c · outbound
EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Collecting highly parallel data for paraphrase evaluation,
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f784b997-ab31-4fbf-b6d6-9a2195580094 · outbound
EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Epic-fusion: Audio-visual temporal binding for egocentric action recognition,
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1b72421b-4411-4f93-bd51-689cb203ecfb · outbound
EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Language models are unsupervised multitask learners,
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a5978af-2eed-4efd-9175-fb775f39d593 · outbound
EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization HowTo100M: Learning a text-video embedding by watching hundred million narrated video clips,
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cc0c3d76-bd7f-44da-b907-3baf74be4d47 · outbound
EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Rethinking spatiotem- poral feature learning: Speed-accuracy trade-offs in video classification,
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b475f87a-9c8c-4317-b1b7-9243d14519ca · outbound
EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Learning transferable visual models from natural language supervi- sion,
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e9c8fe45-b40d-420a-b7cb-4a383d5660b0 · outbound
EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Hiervl: Learning hierarchical video-language embeddings,
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c85f5577-4f51-4526-b62d-8f395bb174e9 · outbound
EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization SViTT-Ego: A Sparse Video-Text Transformer for Egocentric Video
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5ecc0798-5e90-4dfe-858a-40557e35fb32 · outbound
EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Decoupled weight decay regularization
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4be5b5db-e73a-4a44-8640-2a6d06e242af · outbound
EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 784e0a32-bc55-4749-9325-564a43c2f4c0 · outbound
EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization RoBERTa: A Robustly Optimized BERT Pretraining Approach
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f36432fc-d3bc-424a-92ef-4a6a6846fb18 · inbound
EARL: Towards a Unified Analysis-Guided Reinforcement Learning Framework for Egocentric Interaction Reasoning and Pixel Grounding EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.