Pith. sign in

Paper Citation Record · LEDGER

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization

As of 7 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 1 inbound Pith citation observation for arXiv:2506.14356.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.14356 v1

Coverage vector

measured 54 of 54 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:26:05.265658Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-30T21:29:27.063028Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T21:35:04.650377Z

Reference resolution

54 of 54 outbound references displayed

  • verified exact3
  • verified fuzzy34
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3b3039c5-4f63-43b2-96a1-d4c4fc357c19 · outbound

This paper cites Quo vadis, action recognition? a new model and the kinetics dataset,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Quo vadis, action recognition? a new model and the kinetics dataset,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:10.550525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:25:59.431559Z digest=sha256:fc281720d1baf79c1831aaf58425220a00c01a1bc50241723bdaee08791dd94c

Observation 3b05827e-9bcb-4504-b918-91617a462caa · outbound

This paper cites Video summarization through reinforcement learning with a 3d spatio- temporal u-net,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Video summarization through reinforcement learning with a 3d spatio- temporal u-net,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:10.535398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:25:59.542659Z digest=sha256:84f424f08f6cd36b5fe9c8e79f8a7dea73e46d11c95df0ccae408e5d8385a186

Observation b2e2b62d-3ff3-42f4-b510-693cd4de7b11 · outbound

This paper cites Deep attention network for egocentric action recognition,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Deep attention network for egocentric action recognition,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:10.521628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:25:59.700863Z digest=sha256:57f9dd30c1ec3a67023c9570c1587d4810c1627a9ebf87cfe76114bdbb8830f4

Observation 416ae5dc-9eea-4ae5-8c37-38a17c72361d · outbound

This paper cites Training a Large Video Model on a Single Machine in a Day.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Training a Large Video Model on a Single Machine in a Day

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:26:05.951232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:25:59.818943Z digest=sha256:48cfa995afe05da8f85c3f1840bd0feb83399b4c864b29b872be7d123abd4fe1

Observation c5d6b164-6f58-4f03-a66e-b3565b0bde8d · outbound

This paper cites InternVideo: General Video Foundation Models via Generative and Discriminative Learning.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T00:25:59.990074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:25:59.990074Z digest=sha256:f8990f93aedcf3bcae236f09713c5848ddcb1a200204ecd45ccffcbfccdf7b7b

Observation 0c6c2b82-1304-4e98-8301-7975bacc3446 · outbound

This paper cites InternVideo2: Scaling Foundation Models for Multimodal Video Understanding.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:00.161920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:26:00.161920Z digest=sha256:281c054908cf0595cf93ff32839100527bf12c5763f4f696ff8b53dcd1e02669

Observation a2d5163b-29bd-4182-9040-32cfece62edd · outbound

This paper cites EgoVLPv2: Egocentric video-language pre-training with fusion in the backbone,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization EgoVLPv2: Egocentric video-language pre-training with fusion in the backbone,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:10.507959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:26:00.334603Z digest=sha256:803af3625a28ad2d94f982ef0892f5325740ad9749018750267c7ba81c3b979f

Observation 82d49f24-f819-402c-81c9-472ae702a1bd · outbound

This paper cites Egocentric video-language pretraining,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Egocentric video-language pretraining,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:10.495279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:26:00.477538Z digest=sha256:8f0e3a353cf32dd4769574a18d6d7859cd214d1405f63a6adf9e7cae4f16dd23

Observation 144cc8dd-9c57-48e0-8412-8e9a04012dba · outbound

This paper cites Improving semantic video retrieval models by training with a relevance-aware online mining strategy,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Improving semantic video retrieval models by training with a relevance-aware online mining strategy,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:10.481056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:26:00.576434Z digest=sha256:ef0f7e6059d851442312b0b5bb2b87c92cf2e86e35fecc1b60d46c815a01891d

Observation 04357bdc-9f02-452b-ae9e-54018f9be793 · outbound

This paper cites Learning video representations from large language models,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Learning video representations from large language models,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:10.467688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:26:00.716989Z digest=sha256:ddaa3e32c0ad43a6bf5661e6d17a972eabde5506b6f535a388ee4335837fd52a

Observation f9ec9a56-fbec-4170-bf4d-b6170093b1b6 · outbound

This paper cites EgoVideo: Exploring Egocentric Foundation Model and Downstream Adaptation.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization EgoVideo: Exploring Egocentric Foundation Model and Downstream Adaptation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:00.852320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:26:00.852320Z digest=sha256:589521088b80ddc68f3fc931ad2743f3a3eeb2f19bde7abe34f5e191a02b6945

Observation 93e8b51b-8004-400e-9a20-d4c80012e913 · outbound

This paper cites Videomae: Masked autoen- coders are data-efficient learners for self-supervised video pre-training,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Videomae: Masked autoen- coders are data-efficient learners for self-supervised video pre-training,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:10.453114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:26:01.017157Z digest=sha256:a51fd4eadc456193332cdd559daf35053f0f7cae8583d4dd5fbd0008e00a3d94

Observation 3aa1e723-af16-4d5f-8aa8-4345f828d67c · outbound

This paper cites LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:01.224115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:26:01.224115Z digest=sha256:4bad383c7afe9a4bdd6d2ceca5157b0abf7e8354d8a341dfbbf752373e1d8cd0

Observation 4cda262c-684a-4f56-9402-0d6eabd98073 · outbound

This paper cites Internvid: A large-scale video-text dataset for multimodal understanding and generation,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Internvid: A large-scale video-text dataset for multimodal understanding and generation,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:10.437990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:26:01.393406Z digest=sha256:2b7a4a519b4416e12a7b21fbe2680a55cd1e08efb506d8ed23e1dd18f043617a

Observation 69ab1970-654d-4c41-a237-ef7ac463104c · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:01.534507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:26:01.534507Z digest=sha256:ce0e699d6e4ce2288acaaca522f9ed528273ddcefc9531a085ae9f80174e63dd

Observation 7fbc78e4-1606-4303-b25e-ece1889a162a · outbound

This paper cites EVA-CLIP: Improved Training Techniques for CLIP at Scale.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:01.692337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:26:01.692337Z digest=sha256:a5ba8b1da9e3314f7a37e90a3f3213e943cb5d2257680cab6ea335293a4e91ec

Observation 7176760f-561d-43fe-bc91-9b9f7c99fb15 · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:01.868488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:26:01.868488Z digest=sha256:3dace60233b14486c37c30eee5b5513a153a95265b6e91e02b6c247054f65ba1

Observation 0c2c0630-67ea-4f02-96fb-a8e0cd32a142 · outbound

This paper cites EVA-02: A Visual Representation for Neon Genesis.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization EVA-02: A Visual Representation for Neon Genesis

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:01.964467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:26:01.964467Z digest=sha256:226eca485b4a4e932207a7afb15bd9c58287427a2a8ab74f0273f804da24f695

Observation 45c01c0e-2c0c-4446-ae29-20e8d6fed1c2 · outbound

This paper cites Multi- similarity loss with general pair weighting for deep metric learning,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Multi- similarity loss with general pair weighting for deep metric learning,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:10.424959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:26:02.137317Z digest=sha256:56afb0616bf5da4d0b7ed44f1d12dcd5d913a15e64a0efc0fa4eb2a654bbecb0

Observation ab8d33a3-8344-4b78-99e6-7feda72268a3 · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Ego4d: Around the world in 3,000 hours of egocentric video,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:10.410562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:26:02.208447Z digest=sha256:db9ffef159221990c1f6232ba539554d4f398dbde7471143637d550537e6248e

Observation 7aed2aee-2f71-4fe6-bcdd-308d381bc970 · outbound

This paper cites Rescaling egocentric vision: Collection, pipeline and challenges for epic-kitchens- 100,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Rescaling egocentric vision: Collection, pipeline and challenges for epic-kitchens- 100,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:10.218130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:26:02.280538Z digest=sha256:02abac478622bc66f9d0208bf8e089660ed3311c8b0635a2f715e09bf14ee33a

Observation 7ba327eb-6879-4ede-adc7-5b81d3a058b7 · outbound

This paper cites Scaling egocentric vision: The epic-kitchens dataset,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Scaling egocentric vision: The epic-kitchens dataset,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:09.831272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:26:02.395824Z digest=sha256:96f0097ba8388f46c02fce4b435b0d4d43852137adb64a0befaa3e3cd2704c0e

Observation f8f59c61-0302-4d25-aa0e-c79e170886c6 · outbound

This paper cites Charades-Ego: A Large-Scale Dataset of Paired Third and First Person Videos.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Charades-Ego: A Large-Scale Dataset of Paired Third and First Person Videos

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:02.474819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:26:02.474819Z digest=sha256:cfdadf3318522a88b770aa4d2c28fa9132c116d59bb93cc89b97d569e644c46c

Observation 7c7b3501-3185-4670-a377-336c7b19a095 · outbound

This paper cites Long short-term memory,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Long short-term memory,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:02.579192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:26:02.579192Z digest=sha256:42486bbba8d68a47d0abd034f910f885ad653bd7b3d042f04767c5d672f01dc3

Observation 0f86341f-085c-4b6c-9496-678cb767245d · outbound

This paper cites Is space-time attention all you need for video understanding?.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Is space-time attention all you need for video understanding?

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:09.719279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:26:02.688759Z digest=sha256:fdda97a0c017c3b437e32a5e0279df5109b9fdf550cbd61858362fa286512ffa

Observation 8a3e82d7-031b-4fea-b2b7-9e35cac358e4 · outbound

This paper cites Frozen in time: A joint video and image encoder for end-to-end retrieval,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Frozen in time: A joint video and image encoder for end-to-end retrieval,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:02.803661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:26:02.803661Z digest=sha256:1bc000f3969ecfd201086150cbe0fea4e2980458ec69cbf0668eb2cdfcb8987a

Observation df4130a7-8ddd-434b-a195-251c34c4c424 · outbound

This paper cites VideoMAE V2: Scaling video masked autoencoders with dual masking,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization VideoMAE V2: Scaling video masked autoencoders with dual masking,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:09.436074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:26:02.910107Z digest=sha256:dbfde216164909ce410fe0e0f7e01ef2742ff24d4a187d7578a1ffea7af18742

Observation 47013c40-a70d-4d51-b498-12e7e211c017 · outbound

This paper cites Flamingo: a visual language model for few-shot learning,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Flamingo: a visual language model for few-shot learning,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:09.120647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:26:02.983354Z digest=sha256:2871528df8e1df40e198a2830460152dcbd2b839aa40def68a72d326aefa30d3

Observation 46fab9ae-e11f-4030-a325-dc9b400980ad · outbound

This paper cites Roformer: En- hanced transformer with rotary position embedding,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Roformer: En- hanced transformer with rotary position embedding,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:08.908407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:26:03.094336Z digest=sha256:e4625686e6ded10e047488d58e9b122e553076ba56224dbbe3a32fc0c48800dd

Observation d3b604c5-7e92-45b2-bcb2-12430e34096d · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:03.203928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:26:03.203928Z digest=sha256:239645020492be242d730fc9e89de6b1fba5cdf626d068e5bff75f17236420ce

Observation 72d0da19-90bd-49a7-bad9-1dce2fcd7bc2 · outbound

This paper cites VideoRoPE: What Makes for Good Video Rotary Position Embedding?.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization VideoRoPE: What Makes for Good Video Rotary Position Embedding?

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:03.315947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:26:03.315947Z digest=sha256:d1d8f4f91edaead8ff8f8ca90cc1d2bb659364ceab3d2f01f3a35c71bd1ad253

Observation 1db5d195-264f-4332-b8ae-69d2eef1df4c · outbound

This paper cites Supervised contrastive learn- ing,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Supervised contrastive learn- ing,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:08.615395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:26:03.397899Z digest=sha256:54e46f748475cfc32a271460b2ff712a466b21c93965769591db0e849678654e

Observation a6dcc2cf-ea91-4430-a219-604ca220f721 · outbound

This paper cites Parameter-free deep multi-modal clustering with reliable contrastive learning,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Parameter-free deep multi-modal clustering with reliable contrastive learning,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:08.308703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:26:03.518745Z digest=sha256:1d2bb4688843976e1c0eba0335f35ff86499ab59d3dc4fd6d84209817cf5369e

Observation 34495b16-3787-4229-aac2-8c1013cfe167 · outbound

This paper cites Cross-modal contrastive learning network for few-shot action recognition,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Cross-modal contrastive learning network for few-shot action recognition,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:08.040114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:26:03.600867Z digest=sha256:6108f33d74a4f5ee6360c643901f66af7c5362545fffe01f66d55d7d3a62355b

Observation 2423d424-3315-4b21-a30c-f5c861e37277 · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Representation Learning with Contrastive Predictive Coding

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:03.695009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:26:03.695009Z digest=sha256:f9414b6fb286ea33e7a063118b7ff76fda809a71a7ad356ae3866ab5e1d9f6c5

Observation 30bce6e3-7aa2-4791-a3b7-8242b8c22b93 · outbound

This paper cites End-to-end learning of visual representations from uncurated instructional videos,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization End-to-end learning of visual representations from uncurated instructional videos,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:07.827552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:26:03.805158Z digest=sha256:133de033139a455cad7395538373b415bd547cc029c3f358247eea90a9cc75e8

Observation dd7dbc9d-7868-4a71-ae29-fec6e005ca2f · outbound

This paper cites Facenet: A unified embed- ding for face recognition and clustering,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Facenet: A unified embed- ding for face recognition and clustering,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:07.728889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:26:03.911202Z digest=sha256:6cc9292636d4b6e1ec03675006c9a6c09490221df4d4281534a2a5988d4c2071

Observation 5858fc87-7a64-427c-8795-10f7672704a3 · outbound

This paper cites Circle loss: A unified perspective of pair similarity optimization,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Circle loss: A unified perspective of pair similarity optimization,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:07.605657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:26:03.992358Z digest=sha256:7a2232142e35a036f24efeb1d22b5bfe73162c85bab080d6bd5eeee99d7011d9

Observation 49ed26cd-60af-4674-93e2-0eb323b8581d · outbound

This paper cites Relevance-based margin for contrastively-trained video retrieval mod- els,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Relevance-based margin for contrastively-trained video retrieval mod- els,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:07.497252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:26:04.100339Z digest=sha256:a1891ae2722cf59668cf81ce0a3c8bb3afd936104b26808b2789057a132d221e

Observation 677b1418-3e49-410a-a9de-4e81b4db830a · outbound

This paper cites Masked video distillation: Rethinking masked feature modeling for self-supervised video representation learning,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Masked video distillation: Rethinking masked feature modeling for self-supervised video representation learning,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:07.350471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:26:04.178813Z digest=sha256:21e5df569c3395b6344cd11308c383366e3c0552a2136e350238851e762181a0

Observation 2f8337f0-1222-4470-8bbf-3b7e10e8222a · outbound

This paper cites Fine-grained action retrieval through multiple parts-of-speech embeddings,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Fine-grained action retrieval through multiple parts-of-speech embeddings,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:07.212680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:26:04.268357Z digest=sha256:548279a2ab8c53e89ede84948f3bf197ddeaca4ad10e8c93d2bd2da82f004844

Observation aa116c9e-61d8-4509-b642-0b598643aea5 · outbound

This paper cites On semantic similarity in video retrieval,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization On semantic similarity in video retrieval,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:07.075302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:26:04.332494Z digest=sha256:fc560d7e9251185d63aa0aaa878c3a1068edf65e8d2c9e8f094b80501eae0a00

Observation 431349d9-3713-4553-afbb-2d7961ed415d · outbound

This paper cites Egocentric Video-Language Pretraining @ EPIC-KITCHENS-100 Multi-Instance Retrieval Challenge 2022.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Egocentric Video-Language Pretraining @ EPIC-KITCHENS-100 Multi-Instance Retrieval Challenge 2022

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:26:05.630065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:26:04.405236Z digest=sha256:db4676172913e536d3b0bf435bfcd3503b1457f696f6f4720c5df95377755d42

Observation d653bc96-b238-49b9-bc71-1b6c966c223c · outbound

This paper cites Collecting highly parallel data for paraphrase evaluation,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Collecting highly parallel data for paraphrase evaluation,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:06.939788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:26:04.491269Z digest=sha256:00275f1abcad423c1ae2010302d07c9d07a4eae84eaaf34e70cc4f8357cf8c74

Observation f784b997-ab31-4fbf-b6d6-9a2195580094 · outbound

This paper cites Epic-fusion: Audio-visual temporal binding for egocentric action recognition,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Epic-fusion: Audio-visual temporal binding for egocentric action recognition,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:06.834335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:26:04.577115Z digest=sha256:d037df8971a905b5d16a51a1c548b6447d6319be94acac45467023745f5effdd

Observation 1b72421b-4411-4f93-bd51-689cb203ecfb · outbound

This paper cites Language models are unsupervised multitask learners,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Language models are unsupervised multitask learners,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:04.664578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:26:04.664578Z digest=sha256:8d6ba1f5ecf5dc6e1d6c2b8535ba230bf4aee6545ea48399edca0c472fb56b94

Observation 7a5978af-2eed-4efd-9175-fb775f39d593 · outbound

This paper cites HowTo100M: Learning a text-video embedding by watching hundred million narrated video clips,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization HowTo100M: Learning a text-video embedding by watching hundred million narrated video clips,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:06.732224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:26:04.750789Z digest=sha256:e9e008256ed42ba3ba7cea482cd0fc5bf397d923aea4bdae32de2e44bc6f7f7d

Observation cc0c3d76-bd7f-44da-b907-3baf74be4d47 · outbound

This paper cites Rethinking spatiotem- poral feature learning: Speed-accuracy trade-offs in video classification,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Rethinking spatiotem- poral feature learning: Speed-accuracy trade-offs in video classification,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:06.594967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:26:04.815632Z digest=sha256:5d354b37edebc8571a34a88492f6e59f90cdb904401d6752c6817d9ef7096e9e

Observation b475f87a-9c8c-4317-b1b7-9243d14519ca · outbound

This paper cites Learning transferable visual models from natural language supervi- sion,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Learning transferable visual models from natural language supervi- sion,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:06.422516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:26:04.888141Z digest=sha256:8089df6ba380eaabd653c61219d0c94a402b3845811b20be87d15131c89abdf0

Observation e9c8fe45-b40d-420a-b7cb-4a383d5660b0 · outbound

This paper cites Hiervl: Learning hierarchical video-language embeddings,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Hiervl: Learning hierarchical video-language embeddings,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:06.261608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:26:04.957092Z digest=sha256:e7385f210bc968cd50e8b527f8cc5c4c8c26501ac8f8b8be5c94c078e4f99355

Observation c85f5577-4f51-4526-b62d-8f395bb174e9 · outbound

This paper cites SViTT-Ego: A Sparse Video-Text Transformer for Egocentric Video.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization SViTT-Ego: A Sparse Video-Text Transformer for Egocentric Video

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:26:05.459079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:26:05.023665Z digest=sha256:66b0863a35ea2a85aa396262540aec7af8445730df87b2c44e954615e28c9cc2

Observation 5ecc0798-5e90-4dfe-858a-40557e35fb32 · outbound

This paper cites Decoupled weight decay regularization.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Decoupled weight decay regularization

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:06.125884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:26:05.111484Z digest=sha256:300d3cd89fcd6133b9b7167f09e5f5d22906179f523d3baf5cfe7bd012c5df06

Observation 4be5b5db-e73a-4a44-8640-2a6d06e242af · outbound

This paper cites DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:05.197274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:26:05.197274Z digest=sha256:2ab3d84beddd5b4124020b1e2fd971817a8cf6ef6ea88d9f64e2c66849ee390e

Observation 784e0a32-bc55-4749-9325-564a43c2f4c0 · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:05.265658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:26:05.265658Z digest=sha256:0905861280797b419b4881ab591c1187f15b3b5702a138b4424b5def804013a1

Pith citing papers

Observation f36432fc-d3bc-424a-92ef-4a6a6846fb18 · inbound

EARL: Towards a Unified Analysis-Guided Reinforcement Learning Framework for Egocentric Interaction Reasoning and Pixel Grounding cites this paper.

EARL: Towards a Unified Analysis-Guided Reinforcement Learning Framework for Egocentric Interaction Reasoning and Pixel Grounding EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-06-30T21:35:04.652174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T21:29:27.063028Z digest=sha256:4beb8315857e8f1171f261310bf9432fc8ef5387b1351309a840fbe90ea999c8