Pith. sign in

Paper Citation Record · LEDGER

Improving Keystep Recognition in Ego-Video via Dexterous Focus

As of 7 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 0 inbound Pith citation observations for arXiv:2506.00827.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.00827 v1

Coverage vector

measured 25 of 25 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:59:06.328656Z

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

25 of 25 outbound references displayed

  • verified exact0
  • verified fuzzy16
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3abed9b5-24a3-452f-9d78-d2c639b88a83 · outbound

This paper cites Introducing hot3d: An egocentric dataset for 3d hand and object tracking, 2024.

Improving Keystep Recognition in Ego-Video via Dexterous Focus Introducing hot3d: An egocentric dataset for 3d hand and object tracking, 2024

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:59:10.167863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:59:03.398227Z digest=sha256:ace26d54822a6d8a5e42cb56c573232a6ccaa83c7fe3b566705b8a6c95f36144

Observation 3e4ffd7f-5cf5-4159-a741-d0c9a61e8198 · outbound

This paper cites Is space-time attention all you need for video understanding? In Proceedings of the International Conference on Machine Learning (ICML), 2021.

Improving Keystep Recognition in Ego-Video via Dexterous Focus Is space-time attention all you need for video understanding? In Proceedings of the International Conference on Machine Learning (ICML), 2021

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:59:09.936909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:59:03.487431Z digest=sha256:9e06da3e088d46ac3fb29b60a92da6c9e5d3df0eb43818db2f9071c01141b627

Observation cd0d9544-8ffb-4cc7-b313-fee242e9d6d8 · outbound

This paper cites A dense-sparse complementary network for human ac- tion recognition based on rgb and skeleton modalities.Expert Systems with Applications, 244:123061, 2024.

Improving Keystep Recognition in Ego-Video via Dexterous Focus A dense-sparse complementary network for human ac- tion recognition based on rgb and skeleton modalities.Expert Systems with Applications, 244:123061, 2024

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:59:09.728003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:59:03.617735Z digest=sha256:3eeffa886ae343c6c0d3f3d767fcbd7b63d21610e519717398c6101012cbf320

Observation d01aa204-5528-421e-970c-266d7a7ec5b4 · outbound

This paper cites Rescaling egocentric vision: Collection, pipeline and challenges for epic-kitchens-100.

Improving Keystep Recognition in Ego-Video via Dexterous Focus Rescaling egocentric vision: Collection, pipeline and challenges for epic-kitchens-100

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:59:09.409102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:59:03.727388Z digest=sha256:6d810334759c0594dc8f991cc717f6a7bbfb03a09602e39dfb20fcef134bdb46

Observation 55c20920-77fb-49d5-9731-6115f10a76eb · outbound

This paper cites Activitynet: A large-scale video bench- mark for human activity understanding.

Improving Keystep Recognition in Ego-Video via Dexterous Focus Activitynet: A large-scale video bench- mark for human activity understanding

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:59:09.175037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:59:03.863529Z digest=sha256:dd3d9dfb7d2ad1738e6f9732b8ecebaa9c7e551a012e152da4da0ad321478bd7

Observation d53db83b-735f-422f-9d84-18dc87ed6150 · outbound

This paper cites The” something something” video database for learning and evaluating visual common sense.

Improving Keystep Recognition in Ego-Video via Dexterous Focus The” something something” video database for learning and evaluating visual common sense

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:03.974014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:03.974014Z digest=sha256:24a90af69b51c1e4985df943d0228fbe983766dcd2fb4cfb87057f6b4c493784

Observation 66932c3b-2b11-4283-aeec-d96c5049da21 · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video.

Improving Keystep Recognition in Ego-Video via Dexterous Focus Ego4d: Around the world in 3,000 hours of egocentric video

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:59:08.993075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:59:04.117291Z digest=sha256:015174179ccceeb590f394211c0315161e1d122108cbc595fa3689399661aa61

Observation 4c95153d-6583-4f16-9ede-e22832de8963 · outbound

This paper cites Ego-exo4d: Understanding skilled human activity from first- and third-person perspectives.

Improving Keystep Recognition in Ego-Video via Dexterous Focus Ego-exo4d: Understanding skilled human activity from first- and third-person perspectives

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:59:08.753227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:59:04.282107Z digest=sha256:77186e81527858d56aeb409ba180ae1161c9300cf7a79c0c36550da9abdd6d83

Observation 1efd576f-2e94-47ee-95b1-7a71c17383e7 · outbound

This paper cites Jiang, J.

Improving Keystep Recognition in Ego-Video via Dexterous Focus Jiang, J

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:59:08.535971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:59:04.419006Z digest=sha256:92ce281c51efb6ce6b888f3857795ddfed8fd2f3e6b49d087deba72debbb0998

Observation 48a33909-d9a7-4168-a6ad-a43e964f66c7 · outbound

This paper cites The Kinetics Human Action Video Dataset.

Improving Keystep Recognition in Ego-Video via Dexterous Focus The Kinetics Human Action Video Dataset

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:04.567274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:04.567274Z digest=sha256:bfb173bde07ddf88113fdb175f95467cb1a21c699a19a29b2ff293ebab9565fe

Observation 8451dedf-ce70-46ab-a700-ecaed6225350 · outbound

This paper cites Epic-fusion: Audio-visual temporal binding for egocentric action recognition.

Improving Keystep Recognition in Ego-Video via Dexterous Focus Epic-fusion: Audio-visual temporal binding for egocentric action recognition

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:59:08.297259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:59:04.723588Z digest=sha256:97a212b3a21683421f47c43c2415b5f2e79e95486953f3deb97425a2924eef77

Observation 40a7b2ac-a119-4eff-8dc0-629831ba756b · outbound

This paper cites Human action recognition and predic- tion: A survey.

Improving Keystep Recognition in Ego-Video via Dexterous Focus Human action recognition and predic- tion: A survey

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:04.874486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:04.874486Z digest=sha256:c76228347401ca8cd6ff31a6a5c9ac0b748b3baa4ade589d6c4c53f7a9691d09

Observation b0aeda98-7a95-4851-9142-0e617663e98c · outbound

This paper cites X-mic: Cross-modal instance conditioning for egocentric action gen- eralization.

Improving Keystep Recognition in Ego-Video via Dexterous Focus X-mic: Cross-modal instance conditioning for egocentric action gen- eralization

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:59:08.080916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:59:04.959598Z digest=sha256:5b0409fee442b1fbea320198767b1e11763d71ce3e9091aed4ca7ba5c79f0ae9

Observation fadc45a9-2ed8-401f-ba91-ae8cc6ed7a39 · outbound

This paper cites Ego-exo: Transferring visual representations from third-person to first-person videos.

Improving Keystep Recognition in Ego-Video via Dexterous Focus Ego-exo: Transferring visual representations from third-person to first-person videos

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:59:07.845496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:59:05.102608Z digest=sha256:ebd1658e865ce0fb6499e5f639a2205fb91000ef648d063b026026657a9df064

Observation 346dd980-3330-42c3-9f7d-44302af0302c · outbound

This paper cites Egocentric Video-Language Pretraining.

Improving Keystep Recognition in Ego-Video via Dexterous Focus Egocentric Video-Language Pretraining

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:05.218954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:05.218954Z digest=sha256:4a72f420ef68ea0039af0157cc98c15105312f8fcd182d4bca10435045cfd57f

Observation d544b563-dcdd-462c-892d-56b189b820b2 · outbound

This paper cites Where a strong backbone meets strong features – action- former for ego4d moment queries challenge.

Improving Keystep Recognition in Ego-Video via Dexterous Focus Where a strong backbone meets strong features – action- former for ego4d moment queries challenge

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:59:07.646633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:59:05.321906Z digest=sha256:545d2ade27418120941c37e07a5398e24625c484590dcd4ee22aaf543d905054

Observation 3d699295-a9ef-48c6-a957-15ee3edeb865 · outbound

This paper cites Egoenv: Human- centric environment representations from egocentric video.

Improving Keystep Recognition in Ego-Video via Dexterous Focus Egoenv: Human- centric environment representations from egocentric video

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:59:07.399174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:59:05.427141Z digest=sha256:665d671e54db32b7ceed6bb129a88880f81f7d3efb5f7ffe560b1d579c6d5054

Observation b028e65e-5af7-44d2-847d-cc2e36b35e00 · outbound

This paper cites Project aria: A new tool for ego- centric multi-modal ai research, 2023.

Improving Keystep Recognition in Ego-Video via Dexterous Focus Project aria: A new tool for ego- centric multi-modal ai research, 2023

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:59:07.179449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:59:05.530517Z digest=sha256:f7db6dc4e821417776eb03941595f38f5d98d46bd82d652be3cef86ff15109c7

Observation 947f55c1-1528-4471-91e6-07c289a7e6b9 · outbound

This paper cites EgoVideo: Exploring Egocentric Foundation Model and Downstream Adaptation.

Improving Keystep Recognition in Ego-Video via Dexterous Focus EgoVideo: Exploring Egocentric Foundation Model and Downstream Adaptation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:05.646267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:05.646267Z digest=sha256:f077d094dfec8e484360feabf2a8737ba90069a38bd5f623946fad4ae3c8b8dd

Observation fb824526-af44-4294-9793-660ddbe35d51 · outbound

This paper cites EgoVLPv2: Egocentric Video-Language Pre-training with Fusion in the Backbone.

Improving Keystep Recognition in Ego-Video via Dexterous Focus EgoVLPv2: Egocentric Video-Language Pre-training with Fusion in the Backbone

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:05.781867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:05.781867Z digest=sha256:f6069080a5decf5a102a88e4a6d4edb3c6e5fa5f9e80aadc6bf78d99cf9ecf46

Observation 79cdaa7a-0386-48ac-9509-6de7f4c6d0d8 · outbound

This paper cites Understanding human hands in contact at internet scale.

Improving Keystep Recognition in Ego-Video via Dexterous Focus Understanding human hands in contact at internet scale

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:59:06.899474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:59:05.899839Z digest=sha256:29d7c006f4faa9a0a24ab5d7d4b68e1cee66a1dc645c2c66df847dc978a2162a

Observation e00bff9a-e979-4536-b5e7-5081ae2ed929 · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

Improving Keystep Recognition in Ego-Video via Dexterous Focus Representation Learning with Contrastive Predictive Coding

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:06.003000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:06.003000Z digest=sha256:0f72416ceed227150fd781a2ca746010cce6e8ba865689b3bda9f5ba00e9b616

Observation 348e432f-cd03-416b-8af0-23bf91de4165 · outbound

This paper cites InternVideo: General Video Foundation Models via Generative and Discriminative Learning.

Improving Keystep Recognition in Ego-Video via Dexterous Focus InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:06.099409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:06.099409Z digest=sha256:a2f7d9bce957ffca8dae807a3b7a5dd58253a745aaa19dcdaf70b5f38e5e95f7

Observation adfd9535-5c8b-4070-a0fc-537de4b1c8a5 · outbound

This paper cites M&M Mix: A Multimodal Multiview Transformer Ensemble.

Improving Keystep Recognition in Ego-Video via Dexterous Focus M&M Mix: A Multimodal Multiview Transformer Ensemble

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:06.205021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:06.205021Z digest=sha256:93b8849fb3c2a809ab7b9e8e592a6204081f6c966f1b30362cafa5993e407e9c

Observation 25ad6a03-c343-4276-a302-0a57316e7302 · outbound

This paper cites Actionformer: Lo- calizing moments of actions with transformers.

Improving Keystep Recognition in Ego-Video via Dexterous Focus Actionformer: Lo- calizing moments of actions with transformers

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:59:06.664550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:59:06.328656Z digest=sha256:935b41a145738b97546bca926ead9b8c16d6f3f2c1e60f579381444bf5e8483a

Pith citing papers

No inbound Pith citation observations are available.