Pith. sign in

Paper Citation Record · LEDGER

VidEgoThink: Assessing Egocentric Video Understanding Capabilities for Embodied AI

As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2410.11623.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.11623 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:28:39.723976Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T06:41:12.615349Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c3229e19-a5d5-4dad-a488-c6a4aacb021a · inbound

EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios cites this paper.

EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios VidEgoThink: Assessing Egocentric Video Understanding Capabilities for Embodied AI

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:00.262834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:00.262834Z digest=sha256:7c40d8dba623919ba2cea41cd84bf94efa48139ed056c940b09e2ccc87cd0be6

Observation d4e4993a-07d0-402b-ad54-f9c698aa5629 · inbound

EgoMe: A New Dataset and Challenge for Following Me via Egocentric View in Real World cites this paper.

EgoMe: A New Dataset and Challenge for Following Me via Egocentric View in Real World VidEgoThink: Assessing Egocentric Video Understanding Capabilities for Embodied AI

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-09T21:30:03.924813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T21:30:03.924813Z digest=sha256:01ba57a16d286b2825c59275b5c6e96ffc549cb0a25b7876a9bd8a19dd541be7

Observation 4b68f2c0-dcc7-4251-8541-cfa9be868bbd · inbound

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering cites this paper.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering VidEgoThink: Assessing Egocentric Video Understanding Capabilities for Embodied AI

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.151522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.151522Z digest=sha256:d4c7d7452ce6ed15e7e985657e3f0e80e8b7201ea0c67a64bda0f420e6bdc4f0

Observation 583c0f2b-dfff-4d05-9d3f-9bdc1775683d · inbound

Embodied-R: Collaborative Framework for Activating Embodied Spatial Reasoning in Foundation Models via Reinforcement Learning cites this paper.

Embodied-R: Collaborative Framework for Activating Embodied Spatial Reasoning in Foundation Models via Reinforcement Learning VidEgoThink: Assessing Egocentric Video Understanding Capabilities for Embodied AI

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T12:28:39.723976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:28:39.723976Z digest=sha256:6ab2fa63d72a0cf66a1e55f2af59147544880db81858251139eeb0208ba8b010

Observation 466cd78f-6508-472d-b095-c69c619c2e8e · inbound

EgoVLM: Policy Optimization for Egocentric Video Understanding cites this paper.

EgoVLM: Policy Optimization for Egocentric Video Understanding VidEgoThink: Assessing Egocentric Video Understanding Capabilities for Embodied AI

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:14:26.902583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:14:26.902583Z digest=sha256:390eaab4316986b23a7031c9bf6bc7283390e535163b0aa104a6ff8f7c75eed7

Observation 7cbc93df-fee0-4c19-be81-59e670680f03 · inbound

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models cites this paper.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models VidEgoThink: Assessing Egocentric Video Understanding Capabilities for Embodied AI

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T17:49:34.756138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:49:34.756138Z digest=sha256:443cf8ce06959f94f411b2c544a082e90d711e589f399c759c837ba16f2cc0c4

Observation c22ae530-88db-429e-a09c-eb3d51c257fe · inbound

Unified Embodied VLM Reasoning with Robotic Action via Autoregressive Discretized Pre-training cites this paper.

Unified Embodied VLM Reasoning with Robotic Action via Autoregressive Discretized Pre-training VidEgoThink: Assessing Egocentric Video Understanding Capabilities for Embodied AI

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-03T13:29:53.068893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:29:53.068893Z digest=sha256:d51cda952c74db017d83cccd3406acb78edaa6a4ee142867aa6a5d5e7b5ddd28

Observation 55295991-c419-4bdc-8535-c022ed56cf83 · inbound

EgoIntent: A Pre-Outcome Micro-Step Benchmark for Understanding What, Why, and Next cites this paper.

EgoIntent: A Pre-Outcome Micro-Step Benchmark for Understanding What, Why, and Next VidEgoThink: Assessing Egocentric Video Understanding Capabilities for Embodied AI

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T05:50:27.155757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:50:27.155757Z digest=sha256:f4791be6a7e6da68706ea658448084cec21191ac5b5692e4753a12ec9fb770c7

Observation 6651e0c6-564f-41b4-8796-26cfb2abd227 · inbound

Efficient Spatial-Temporal Focal Adapter with SSM for Temporal Action Detection cites this paper.

Efficient Spatial-Temporal Focal Adapter with SSM for Temporal Action Detection VidEgoThink: Assessing Egocentric Video Understanding Capabilities for Embodied AI

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:41:12.685790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T17:31:28.091617Z digest=sha256:0d4d197ac21d093357cfa7670201d39d48ddadf7cd1d75b5ed82dc2310e36910

Observation d7319970-128d-43ee-9474-4cdc11218371 · inbound

Watching Movies Like a Human: Egocentric Emotion Understanding for Embodied Companions cites this paper.

Watching Movies Like a Human: Egocentric Emotion Understanding for Embodied Companions VidEgoThink: Assessing Egocentric Video Understanding Capabilities for Embodied AI

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T08:53:04.284252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T08:49:33.107658Z digest=sha256:8d223637b6ffd5bc129fe76bbcbccec9459e35a3a8bc04dec611f9b61058f7ff

Observation edc05bd7-0390-45f0-82c1-7042fad95dfe · inbound

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation cites this paper.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation VidEgoThink: Assessing Egocentric Video Understanding Capabilities for Embodied AI

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:15.211777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:15.211777Z digest=sha256:53a5ad00a65146b306205ff51f67bffb7fbf47f41e1a2a8f443b4daa9cae9a53