Pith. sign in

Paper Citation Record · LEDGER

Improving Keystep Recognition in Ego-Video via Dexterous Focus

As of 9 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 0 inbound Pith citation observations for arXiv:2506.00827.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.00827 v1

Coverage vector

measured 25 of 25 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:59:06.328656Z

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

25 of 25 outbound references displayed

  • verified exact0
  • verified fuzzy16
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3abed9b5-24a3-452f-9d78-d2c639b88a83 · outbound

This paper cites Introducing hot3d: An egocentric dataset for 3d hand and object tracking, 2024.

Improving Keystep Recognition in Ego-Video via Dexterous Focus Introducing hot3d: An egocentric dataset for 3d hand and object tracking, 2024

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:59:10.167863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:59:03.398227Z digest=sha256:495569b3cc0306680d5bcec2a9115c745f93d39eadd6cbe89eaaef4d950552d9

Observation 3e4ffd7f-5cf5-4159-a741-d0c9a61e8198 · outbound

This paper cites Is space-time attention all you need for video understanding? In Proceedings of the International Conference on Machine Learning (ICML), 2021.

Improving Keystep Recognition in Ego-Video via Dexterous Focus Is space-time attention all you need for video understanding? In Proceedings of the International Conference on Machine Learning (ICML), 2021

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:59:09.936909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:59:03.487431Z digest=sha256:a5b76f5d3ac80725e3654bf6ae7454536c2c83cd42bfc8acfa5bc6c2b69031bc

Observation cd0d9544-8ffb-4cc7-b313-fee242e9d6d8 · outbound

This paper cites A dense-sparse complementary network for human ac- tion recognition based on rgb and skeleton modalities.Expert Systems with Applications, 244:123061, 2024.

Improving Keystep Recognition in Ego-Video via Dexterous Focus A dense-sparse complementary network for human ac- tion recognition based on rgb and skeleton modalities.Expert Systems with Applications, 244:123061, 2024

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:59:09.728003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:59:03.617735Z digest=sha256:949651efdac039960769e8c3435a4818d2e5d3c7737f1dc34acd0cfc4f75025f

Observation d01aa204-5528-421e-970c-266d7a7ec5b4 · outbound

This paper cites Rescaling egocentric vision: Collection, pipeline and challenges for epic-kitchens-100.

Improving Keystep Recognition in Ego-Video via Dexterous Focus Rescaling egocentric vision: Collection, pipeline and challenges for epic-kitchens-100

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:59:09.409102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:59:03.727388Z digest=sha256:4ae91efd2e8a730a4a27e982558e7981027383d0b34506a9c7e0e671d42ea2b8

Observation 55c20920-77fb-49d5-9731-6115f10a76eb · outbound

This paper cites Activitynet: A large-scale video bench- mark for human activity understanding.

Improving Keystep Recognition in Ego-Video via Dexterous Focus Activitynet: A large-scale video bench- mark for human activity understanding

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:59:09.175037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:59:03.863529Z digest=sha256:96f8e414ad6aaaa669b0108dd8c302b6ceebe8edef904fe42e212bb057e1f0cd

Observation d53db83b-735f-422f-9d84-18dc87ed6150 · outbound

This paper cites The” something something” video database for learning and evaluating visual common sense.

Improving Keystep Recognition in Ego-Video via Dexterous Focus The” something something” video database for learning and evaluating visual common sense

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:03.974014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:03.974014Z digest=sha256:24a90af69b51c1e4985df943d0228fbe983766dcd2fb4cfb87057f6b4c493784

Observation 66932c3b-2b11-4283-aeec-d96c5049da21 · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video.

Improving Keystep Recognition in Ego-Video via Dexterous Focus Ego4d: Around the world in 3,000 hours of egocentric video

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:59:08.993075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:59:04.117291Z digest=sha256:45cb8b76dec5457ee30190c7c5502bf9afb0804e557f219755b7179ed081e72d

Observation 4c95153d-6583-4f16-9ede-e22832de8963 · outbound

This paper cites Ego-exo4d: Understanding skilled human activity from first- and third-person perspectives.

Improving Keystep Recognition in Ego-Video via Dexterous Focus Ego-exo4d: Understanding skilled human activity from first- and third-person perspectives

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:59:08.753227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:59:04.282107Z digest=sha256:d44aa50ba253a4255c5ca425149b2fe70905cefca31a2168b40a19d4aa2d4a65

Observation 1efd576f-2e94-47ee-95b1-7a71c17383e7 · outbound

This paper cites Jiang, J.

Improving Keystep Recognition in Ego-Video via Dexterous Focus Jiang, J

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:59:08.535971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:59:04.419006Z digest=sha256:ee223889f1f73d3a71c15d01e74a20b500433017eb7d780056b4f92d25edce04

Observation 48a33909-d9a7-4168-a6ad-a43e964f66c7 · outbound

This paper cites The Kinetics Human Action Video Dataset.

Improving Keystep Recognition in Ego-Video via Dexterous Focus The Kinetics Human Action Video Dataset

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:04.567274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:04.567274Z digest=sha256:6c6a1868360c83e3bb10187cd2acae612ffca715605dc18294906321083a9279

Observation 8451dedf-ce70-46ab-a700-ecaed6225350 · outbound

This paper cites Epic-fusion: Audio-visual temporal binding for egocentric action recognition.

Improving Keystep Recognition in Ego-Video via Dexterous Focus Epic-fusion: Audio-visual temporal binding for egocentric action recognition

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:59:08.297259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:59:04.723588Z digest=sha256:3d0eadf73208ffba6ac680ad8a12780d95a1f756ef64b45dc9a53e55586c0b84

Observation 40a7b2ac-a119-4eff-8dc0-629831ba756b · outbound

This paper cites Human action recognition and predic- tion: A survey.

Improving Keystep Recognition in Ego-Video via Dexterous Focus Human action recognition and predic- tion: A survey

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:04.874486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:04.874486Z digest=sha256:c76228347401ca8cd6ff31a6a5c9ac0b748b3baa4ade589d6c4c53f7a9691d09

Observation b0aeda98-7a95-4851-9142-0e617663e98c · outbound

This paper cites X-mic: Cross-modal instance conditioning for egocentric action gen- eralization.

Improving Keystep Recognition in Ego-Video via Dexterous Focus X-mic: Cross-modal instance conditioning for egocentric action gen- eralization

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:59:08.080916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:59:04.959598Z digest=sha256:9d0edcfb0e66f5a50eb68c8925edfcfe8a2567144e9ce3406826f09e6ddf37b9

Observation fadc45a9-2ed8-401f-ba91-ae8cc6ed7a39 · outbound

This paper cites Ego-exo: Transferring visual representations from third-person to first-person videos.

Improving Keystep Recognition in Ego-Video via Dexterous Focus Ego-exo: Transferring visual representations from third-person to first-person videos

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:59:07.845496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:59:05.102608Z digest=sha256:b4927e61053829f8483bd1b9da6b07d5bd7a53a8633983751a5ea16332865c6f

Observation 346dd980-3330-42c3-9f7d-44302af0302c · outbound

This paper cites Egocentric Video-Language Pretraining.

Improving Keystep Recognition in Ego-Video via Dexterous Focus Egocentric Video-Language Pretraining

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:05.218954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:05.218954Z digest=sha256:4a72f420ef68ea0039af0157cc98c15105312f8fcd182d4bca10435045cfd57f

Observation d544b563-dcdd-462c-892d-56b189b820b2 · outbound

This paper cites Where a strong backbone meets strong features – action- former for ego4d moment queries challenge.

Improving Keystep Recognition in Ego-Video via Dexterous Focus Where a strong backbone meets strong features – action- former for ego4d moment queries challenge

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:59:07.646633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:59:05.321906Z digest=sha256:7c93a9b6c16b18f8a39366ebabe60dc6008436f67069d966320399270a23c1c4

Observation 3d699295-a9ef-48c6-a957-15ee3edeb865 · outbound

This paper cites Egoenv: Human- centric environment representations from egocentric video.

Improving Keystep Recognition in Ego-Video via Dexterous Focus Egoenv: Human- centric environment representations from egocentric video

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:59:07.399174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:59:05.427141Z digest=sha256:a5c8431f1a532139181794f6cad43034ab2779d11e030745fc42b0bbb6b76649

Observation b028e65e-5af7-44d2-847d-cc2e36b35e00 · outbound

This paper cites Project aria: A new tool for ego- centric multi-modal ai research, 2023.

Improving Keystep Recognition in Ego-Video via Dexterous Focus Project aria: A new tool for ego- centric multi-modal ai research, 2023

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:59:07.179449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:59:05.530517Z digest=sha256:44041bf79899821d278451990cc54bb256752442f44af9e36ee5ccdd6f12b598

Observation 947f55c1-1528-4471-91e6-07c289a7e6b9 · outbound

This paper cites EgoVideo: Exploring Egocentric Foundation Model and Downstream Adaptation.

Improving Keystep Recognition in Ego-Video via Dexterous Focus EgoVideo: Exploring Egocentric Foundation Model and Downstream Adaptation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:05.646267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:05.646267Z digest=sha256:f077d094dfec8e484360feabf2a8737ba90069a38bd5f623946fad4ae3c8b8dd

Observation fb824526-af44-4294-9793-660ddbe35d51 · outbound

This paper cites EgoVLPv2: Egocentric Video-Language Pre-training with Fusion in the Backbone.

Improving Keystep Recognition in Ego-Video via Dexterous Focus EgoVLPv2: Egocentric Video-Language Pre-training with Fusion in the Backbone

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:05.781867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:05.781867Z digest=sha256:f6069080a5decf5a102a88e4a6d4edb3c6e5fa5f9e80aadc6bf78d99cf9ecf46

Observation 79cdaa7a-0386-48ac-9509-6de7f4c6d0d8 · outbound

This paper cites Understanding human hands in contact at internet scale.

Improving Keystep Recognition in Ego-Video via Dexterous Focus Understanding human hands in contact at internet scale

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:59:06.899474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:59:05.899839Z digest=sha256:c353ed8ec021bb416d51f8eeaeda0254ececa7d0e57b2dee7d83976448a5a3ff

Observation e00bff9a-e979-4536-b5e7-5081ae2ed929 · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

Improving Keystep Recognition in Ego-Video via Dexterous Focus Representation Learning with Contrastive Predictive Coding

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:06.003000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:06.003000Z digest=sha256:0f72416ceed227150fd781a2ca746010cce6e8ba865689b3bda9f5ba00e9b616

Observation 348e432f-cd03-416b-8af0-23bf91de4165 · outbound

This paper cites InternVideo: General Video Foundation Models via Generative and Discriminative Learning.

Improving Keystep Recognition in Ego-Video via Dexterous Focus InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:06.099409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:06.099409Z digest=sha256:a2f7d9bce957ffca8dae807a3b7a5dd58253a745aaa19dcdaf70b5f38e5e95f7

Observation adfd9535-5c8b-4070-a0fc-537de4b1c8a5 · outbound

This paper cites M&M Mix: A Multimodal Multiview Transformer Ensemble.

Improving Keystep Recognition in Ego-Video via Dexterous Focus M&M Mix: A Multimodal Multiview Transformer Ensemble

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:06.205021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:06.205021Z digest=sha256:a73c1289653c87b950850e40c06abc05e42a42db642403bdd581e20e0c9402e0

Observation 25ad6a03-c343-4276-a302-0a57316e7302 · outbound

This paper cites Actionformer: Lo- calizing moments of actions with transformers.

Improving Keystep Recognition in Ego-Video via Dexterous Focus Actionformer: Lo- calizing moments of actions with transformers

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:59:06.664550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:59:06.328656Z digest=sha256:2231f3065857e362d5178c0ef2b2495706bed95bac618cb15258f88f0cabf453

Pith citing papers

No inbound Pith citation observations are available.