Pith. sign in

Paper Citation Record · LEDGER

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization

As of 10 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 1 inbound Pith citation observation for arXiv:2506.14356.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.14356 v1

Coverage vector

measured 54 of 54 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:26:05.265658Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-30T21:29:27.063028Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T21:35:04.650377Z

Reference resolution

54 of 54 outbound references displayed

  • verified exact3
  • verified fuzzy34
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3b3039c5-4f63-43b2-96a1-d4c4fc357c19 · outbound

This paper cites Quo vadis, action recognition? a new model and the kinetics dataset,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Quo vadis, action recognition? a new model and the kinetics dataset,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:10.550525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:25:59.431559Z digest=sha256:145033ca513b392969ea0c11bc685737d9fd8c307c2b93bbe1a1d46c73ff0549

Observation 3b05827e-9bcb-4504-b918-91617a462caa · outbound

This paper cites Video summarization through reinforcement learning with a 3d spatio- temporal u-net,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Video summarization through reinforcement learning with a 3d spatio- temporal u-net,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:10.535398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:25:59.542659Z digest=sha256:bf77a94a1bad8e69c02eec3c1afa484022a2a325e2dbca5ae99d8610ff8f0c20

Observation b2e2b62d-3ff3-42f4-b510-693cd4de7b11 · outbound

This paper cites Deep attention network for egocentric action recognition,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Deep attention network for egocentric action recognition,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:10.521628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:25:59.700863Z digest=sha256:407380ff3990a50208b70c0f06103c820300d858c2eb59be2464b0fa5617edde

Observation 416ae5dc-9eea-4ae5-8c37-38a17c72361d · outbound

This paper cites Training a Large Video Model on a Single Machine in a Day.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Training a Large Video Model on a Single Machine in a Day

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:26:05.951232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:25:59.818943Z digest=sha256:d547954ae8b9093f225c84353ab654b4384dec3cea7478d2fbfe55d60a5f4b22

Observation c5d6b164-6f58-4f03-a66e-b3565b0bde8d · outbound

This paper cites InternVideo: General Video Foundation Models via Generative and Discriminative Learning.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T00:25:59.990074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:25:59.990074Z digest=sha256:97b5d90898e11f19d94442cfd05f983ebe6e2ed665978debcd5d3776811918f2

Observation 0c6c2b82-1304-4e98-8301-7975bacc3446 · outbound

This paper cites InternVideo2: Scaling Foundation Models for Multimodal Video Understanding.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:00.161920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:26:00.161920Z digest=sha256:27a4eb624927ad9eb81886055294ffdd4cbdaa07d22d46501e46ba9bd9ae8d9c

Observation a2d5163b-29bd-4182-9040-32cfece62edd · outbound

This paper cites EgoVLPv2: Egocentric video-language pre-training with fusion in the backbone,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization EgoVLPv2: Egocentric video-language pre-training with fusion in the backbone,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:10.507959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:26:00.334603Z digest=sha256:a4b0e1af043bd5ac05e389e0fd09d701df955d5fe1edb506d47368bd35551129

Observation 82d49f24-f819-402c-81c9-472ae702a1bd · outbound

This paper cites Egocentric video-language pretraining,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Egocentric video-language pretraining,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:10.495279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:26:00.477538Z digest=sha256:5276d9be5db466c619cd3c0ad7675416fbc9bbd3159961d84d677cfb27ecb0d7

Observation 144cc8dd-9c57-48e0-8412-8e9a04012dba · outbound

This paper cites Improving semantic video retrieval models by training with a relevance-aware online mining strategy,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Improving semantic video retrieval models by training with a relevance-aware online mining strategy,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:10.481056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:26:00.576434Z digest=sha256:de84ad589c5bb47b56f6ac093c31958a569b98e979015f6ecd886edaa1fb7017

Observation 04357bdc-9f02-452b-ae9e-54018f9be793 · outbound

This paper cites Learning video representations from large language models,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Learning video representations from large language models,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:10.467688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:26:00.716989Z digest=sha256:4f887c3cf32a35b6dbcbdbc653c4666f2e380b174c99f6910ca7e9f740dce64b

Observation f9ec9a56-fbec-4170-bf4d-b6170093b1b6 · outbound

This paper cites EgoVideo: Exploring Egocentric Foundation Model and Downstream Adaptation.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization EgoVideo: Exploring Egocentric Foundation Model and Downstream Adaptation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:00.852320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:26:00.852320Z digest=sha256:21313f026dad9ff58d4e4e67d4d59ef14a22ed82da5a258315bb34f12df79c43

Observation 93e8b51b-8004-400e-9a20-d4c80012e913 · outbound

This paper cites Videomae: Masked autoen- coders are data-efficient learners for self-supervised video pre-training,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Videomae: Masked autoen- coders are data-efficient learners for self-supervised video pre-training,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:10.453114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:26:01.017157Z digest=sha256:b4303defa565268bcfbf714ba133def5f59815441fc32a33179fcc1072fcb4a8

Observation 3aa1e723-af16-4d5f-8aa8-4345f828d67c · outbound

This paper cites LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:01.224115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:26:01.224115Z digest=sha256:a9078a069869fd52b2342a3a9728fd2c3d3d3074003ce71b8bade8fa44af8103

Observation 4cda262c-684a-4f56-9402-0d6eabd98073 · outbound

This paper cites Internvid: A large-scale video-text dataset for multimodal understanding and generation,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Internvid: A large-scale video-text dataset for multimodal understanding and generation,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:10.437990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:26:01.393406Z digest=sha256:864505632ecaeaf337cfbeea312cfd411d5cf31a9448797fe127b66467e88fec

Observation 69ab1970-654d-4c41-a237-ef7ac463104c · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:01.534507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:26:01.534507Z digest=sha256:59bc8735c7bae4fd2612307d8ba10e470b88091adfe380f2f014538857efcbf2

Observation 7fbc78e4-1606-4303-b25e-ece1889a162a · outbound

This paper cites EVA-CLIP: Improved Training Techniques for CLIP at Scale.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:01.692337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:26:01.692337Z digest=sha256:d925401ba7c5c98c98fbd0694d8df87ba049facdd329881e2edef219eba224c4

Observation 7176760f-561d-43fe-bc91-9b9f7c99fb15 · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:01.868488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:26:01.868488Z digest=sha256:645037cfb598dc94f1927d9a84fc89d82cfb953c5e621b94d0c8740dd8766783

Observation 0c2c0630-67ea-4f02-96fb-a8e0cd32a142 · outbound

This paper cites EVA-02: A Visual Representation for Neon Genesis.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization EVA-02: A Visual Representation for Neon Genesis

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:01.964467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:26:01.964467Z digest=sha256:6553a298e9050e7b07ea8ad9c207a2853948adc3b951efe2d0cedd79e0766aa5

Observation 45c01c0e-2c0c-4446-ae29-20e8d6fed1c2 · outbound

This paper cites Multi- similarity loss with general pair weighting for deep metric learning,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Multi- similarity loss with general pair weighting for deep metric learning,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:10.424959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:26:02.137317Z digest=sha256:a5202117bdd1e4845b1a9a41a901e5c6e453dcf47a7c70b54fe446385097220c

Observation ab8d33a3-8344-4b78-99e6-7feda72268a3 · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Ego4d: Around the world in 3,000 hours of egocentric video,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:10.410562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:26:02.208447Z digest=sha256:8f51dfb95c7803ec598f7823acf9733d9e1cb655a7ab66974426065c587f2f24

Observation 7aed2aee-2f71-4fe6-bcdd-308d381bc970 · outbound

This paper cites Rescaling egocentric vision: Collection, pipeline and challenges for epic-kitchens- 100,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Rescaling egocentric vision: Collection, pipeline and challenges for epic-kitchens- 100,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:10.218130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:26:02.280538Z digest=sha256:28d9fc8440abfe7deee9e8db33173fe77d2ce612e4760f6298755de58aaf2123

Observation 7ba327eb-6879-4ede-adc7-5b81d3a058b7 · outbound

This paper cites Scaling egocentric vision: The epic-kitchens dataset,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Scaling egocentric vision: The epic-kitchens dataset,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:09.831272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:26:02.395824Z digest=sha256:24e34af7f46fb56fb5c58574c9ab80d28a22991b0e6bebfcc2f7a0efca6368a3

Observation f8f59c61-0302-4d25-aa0e-c79e170886c6 · outbound

This paper cites Charades-Ego: A Large-Scale Dataset of Paired Third and First Person Videos.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Charades-Ego: A Large-Scale Dataset of Paired Third and First Person Videos

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:02.474819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:26:02.474819Z digest=sha256:f240141df29d2954057c9183ea39405d565c770744898344026ec11d71099960

Observation 7c7b3501-3185-4670-a377-336c7b19a095 · outbound

This paper cites Long short-term memory,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Long short-term memory,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:02.579192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:26:02.579192Z digest=sha256:8277cdd54e70155cd5069a3d7d26988b251783a2698fe8ec9c6f0aa3414ffef9

Observation 0f86341f-085c-4b6c-9496-678cb767245d · outbound

This paper cites Is space-time attention all you need for video understanding?.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Is space-time attention all you need for video understanding?

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:09.719279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:26:02.688759Z digest=sha256:d80acb379f09a20a4643aa76689259dafc05ef00cc44d453a0bb072b26cc43de

Observation 8a3e82d7-031b-4fea-b2b7-9e35cac358e4 · outbound

This paper cites Frozen in time: A joint video and image encoder for end-to-end retrieval,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Frozen in time: A joint video and image encoder for end-to-end retrieval,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:02.803661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:26:02.803661Z digest=sha256:0726a5b563f6b3a453220336144d36f34ef4044e4cd32e969112b89f7a9a3d08

Observation df4130a7-8ddd-434b-a195-251c34c4c424 · outbound

This paper cites VideoMAE V2: Scaling video masked autoencoders with dual masking,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization VideoMAE V2: Scaling video masked autoencoders with dual masking,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:09.436074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:26:02.910107Z digest=sha256:bc5ba41333c1ee5039294ca5bf74ca4dad6f1f6e9c08d31cdd013b3e5d8d7cab

Observation 47013c40-a70d-4d51-b498-12e7e211c017 · outbound

This paper cites Flamingo: a visual language model for few-shot learning,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Flamingo: a visual language model for few-shot learning,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:09.120647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:26:02.983354Z digest=sha256:9275f3ec9144afa811d6e97495c7d5f2b9fa558729b5605c57886acbcdbc600a

Observation 46fab9ae-e11f-4030-a325-dc9b400980ad · outbound

This paper cites Roformer: En- hanced transformer with rotary position embedding,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Roformer: En- hanced transformer with rotary position embedding,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:08.908407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:26:03.094336Z digest=sha256:f7108002b6af43f5c4a7e1a8e89af31a3c753f8f0409fd41f6cd0771546e3f56

Observation d3b604c5-7e92-45b2-bcb2-12430e34096d · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:03.203928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:26:03.203928Z digest=sha256:037fda000a70ea95b2ef352391ec52b28bfe78f51572258d0fa329c7fb2ef89d

Observation 72d0da19-90bd-49a7-bad9-1dce2fcd7bc2 · outbound

This paper cites VideoRoPE: What Makes for Good Video Rotary Position Embedding?.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization VideoRoPE: What Makes for Good Video Rotary Position Embedding?

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:03.315947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:26:03.315947Z digest=sha256:c14e01032e3e07be4eda2bfd6c2b87f8952eabcf9f4214339a3bce20e484cfc6

Observation 1db5d195-264f-4332-b8ae-69d2eef1df4c · outbound

This paper cites Supervised contrastive learn- ing,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Supervised contrastive learn- ing,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:08.615395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:26:03.397899Z digest=sha256:3a3726bc265461036867e75b1aa7e15e2639623f922af08a42f0a5df42b1425a

Observation a6dcc2cf-ea91-4430-a219-604ca220f721 · outbound

This paper cites Parameter-free deep multi-modal clustering with reliable contrastive learning,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Parameter-free deep multi-modal clustering with reliable contrastive learning,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:08.308703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:26:03.518745Z digest=sha256:fb37ea6978bcf10a4b667157295f9893c0d724a8ddcbc02885b8152d4390bc6c

Observation 34495b16-3787-4229-aac2-8c1013cfe167 · outbound

This paper cites Cross-modal contrastive learning network for few-shot action recognition,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Cross-modal contrastive learning network for few-shot action recognition,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:08.040114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:26:03.600867Z digest=sha256:4bb088013ad9ef854afa234be7ae2d167fea2eee5b87a2a15b05066e34ee2ab9

Observation 2423d424-3315-4b21-a30c-f5c861e37277 · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Representation Learning with Contrastive Predictive Coding

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:03.695009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:26:03.695009Z digest=sha256:393bbe753fa3c81dc9c114f09581065141f3f6f1dc221a6bfe9e1825cb0b18bf

Observation 30bce6e3-7aa2-4791-a3b7-8242b8c22b93 · outbound

This paper cites End-to-end learning of visual representations from uncurated instructional videos,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization End-to-end learning of visual representations from uncurated instructional videos,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:07.827552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:26:03.805158Z digest=sha256:95f95a731c129be3de2b8b7b736c415fab692c8dbf1831a1dc2a9e12f4b6be9c

Observation dd7dbc9d-7868-4a71-ae29-fec6e005ca2f · outbound

This paper cites Facenet: A unified embed- ding for face recognition and clustering,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Facenet: A unified embed- ding for face recognition and clustering,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:07.728889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:26:03.911202Z digest=sha256:da2952d8fa51ee81346f84a31fe62a1d4738f62a459746a0d91e6c8d44e98514

Observation 5858fc87-7a64-427c-8795-10f7672704a3 · outbound

This paper cites Circle loss: A unified perspective of pair similarity optimization,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Circle loss: A unified perspective of pair similarity optimization,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:07.605657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:26:03.992358Z digest=sha256:adddc2efe966c907614f2fbcd120813beed7270d5fb19d419969fdc77b525372

Observation 49ed26cd-60af-4674-93e2-0eb323b8581d · outbound

This paper cites Relevance-based margin for contrastively-trained video retrieval mod- els,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Relevance-based margin for contrastively-trained video retrieval mod- els,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:07.497252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:26:04.100339Z digest=sha256:7073bb6dded3e810467e086d52b1689089508c3d69788808a7baaa679f996f36

Observation 677b1418-3e49-410a-a9de-4e81b4db830a · outbound

This paper cites Masked video distillation: Rethinking masked feature modeling for self-supervised video representation learning,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Masked video distillation: Rethinking masked feature modeling for self-supervised video representation learning,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:07.350471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:26:04.178813Z digest=sha256:ef713a5dca30c825e2d81accd99c67a7e031f1f43776597b8d6db13d5f09d149

Observation 2f8337f0-1222-4470-8bbf-3b7e10e8222a · outbound

This paper cites Fine-grained action retrieval through multiple parts-of-speech embeddings,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Fine-grained action retrieval through multiple parts-of-speech embeddings,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:07.212680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:26:04.268357Z digest=sha256:42a603bda9af84fdc5cc5a54fe1163f854064b3eb2f78c14435df4eacd9e9731

Observation aa116c9e-61d8-4509-b642-0b598643aea5 · outbound

This paper cites On semantic similarity in video retrieval,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization On semantic similarity in video retrieval,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:07.075302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:26:04.332494Z digest=sha256:cf7c7ae98a6cb881fd6606938f8729702173d0303f9918e7818c2597f1f8b24c

Observation 431349d9-3713-4553-afbb-2d7961ed415d · outbound

This paper cites Egocentric Video-Language Pretraining @ EPIC-KITCHENS-100 Multi-Instance Retrieval Challenge 2022.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Egocentric Video-Language Pretraining @ EPIC-KITCHENS-100 Multi-Instance Retrieval Challenge 2022

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:26:05.630065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:26:04.405236Z digest=sha256:d71087ad9fd36f4ea97e95fe1242d1b951a94b2b44848e3c50ae5ddc379bbd02

Observation d653bc96-b238-49b9-bc71-1b6c966c223c · outbound

This paper cites Collecting highly parallel data for paraphrase evaluation,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Collecting highly parallel data for paraphrase evaluation,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:06.939788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:26:04.491269Z digest=sha256:55c3195aaac7d1bdead0c523dc9f957ca60056fddb82b1fb22a1a4c94ed4d4ba

Observation f784b997-ab31-4fbf-b6d6-9a2195580094 · outbound

This paper cites Epic-fusion: Audio-visual temporal binding for egocentric action recognition,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Epic-fusion: Audio-visual temporal binding for egocentric action recognition,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:06.834335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:26:04.577115Z digest=sha256:8f3d4a00daa73d5f474423317d70b388e6047b523c0206dcae3b17411d7f0af9

Observation 1b72421b-4411-4f93-bd51-689cb203ecfb · outbound

This paper cites Language models are unsupervised multitask learners,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Language models are unsupervised multitask learners,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:04.664578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:26:04.664578Z digest=sha256:748b0d364dd37bd3b23a76729d52ee3b930c19c5a7480140e661bd10664d3129

Observation 7a5978af-2eed-4efd-9175-fb775f39d593 · outbound

This paper cites HowTo100M: Learning a text-video embedding by watching hundred million narrated video clips,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization HowTo100M: Learning a text-video embedding by watching hundred million narrated video clips,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:06.732224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:26:04.750789Z digest=sha256:63c5bde0020026d5ac8ef0d903ec412cbe45f47dffb437d4b6c8e7c27890e733

Observation cc0c3d76-bd7f-44da-b907-3baf74be4d47 · outbound

This paper cites Rethinking spatiotem- poral feature learning: Speed-accuracy trade-offs in video classification,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Rethinking spatiotem- poral feature learning: Speed-accuracy trade-offs in video classification,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:06.594967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:26:04.815632Z digest=sha256:edff54601ed49e55b9fa2d225844752ef89b2beb24b6aef056e7638f7585d821

Observation b475f87a-9c8c-4317-b1b7-9243d14519ca · outbound

This paper cites Learning transferable visual models from natural language supervi- sion,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Learning transferable visual models from natural language supervi- sion,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:06.422516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:26:04.888141Z digest=sha256:9876a608b6c06f329a217a7a3dddccc84dbc91cfdbb32131984926320b59671a

Observation e9c8fe45-b40d-420a-b7cb-4a383d5660b0 · outbound

This paper cites Hiervl: Learning hierarchical video-language embeddings,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Hiervl: Learning hierarchical video-language embeddings,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:06.261608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:26:04.957092Z digest=sha256:49335c583ff22698ba286bb6b0c6cd8bca69f2f85da4dc760213f45164da4b82

Observation c85f5577-4f51-4526-b62d-8f395bb174e9 · outbound

This paper cites SViTT-Ego: A Sparse Video-Text Transformer for Egocentric Video.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization SViTT-Ego: A Sparse Video-Text Transformer for Egocentric Video

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:26:05.459079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:26:05.023665Z digest=sha256:5df99e4e26313368aea0854a924f5eb6cc0f204c4d78beb83f5f6c1451272c58

Observation 5ecc0798-5e90-4dfe-858a-40557e35fb32 · outbound

This paper cites Decoupled weight decay regularization.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Decoupled weight decay regularization

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:06.125884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T00:26:05.111484Z digest=sha256:58f25c6e88bfc0779b1af1bf5cc4ea6a2735cde2f912b46469f6848ca79d4743

Observation 4be5b5db-e73a-4a44-8640-2a6d06e242af · outbound

This paper cites DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:05.197274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:26:05.197274Z digest=sha256:741a6886e6ab43641674f7e32659f912ef6d0c873a0f81cb35898e64ca571994

Observation 784e0a32-bc55-4749-9325-564a43c2f4c0 · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:05.265658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:26:05.265658Z digest=sha256:16ef450c1537c082392a443bf43a3cb8910eb3e61879388c2fc5d93175964717

Pith citing papers

Observation f36432fc-d3bc-424a-92ef-4a6a6846fb18 · inbound

EARL: Towards a Unified Analysis-Guided Reinforcement Learning Framework for Egocentric Interaction Reasoning and Pixel Grounding cites this paper.

EARL: Towards a Unified Analysis-Guided Reinforcement Learning Framework for Egocentric Interaction Reasoning and Pixel Grounding EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-06-30T21:35:04.652174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T21:29:27.063028Z digest=sha256:cdf4fd9c8df6c78849c7886b7da78c3a30f338ada1795619fc74cec95d859f48