Pith. sign in

Paper Citation Record · LEDGER

EgoVLM: Policy Optimization for Egocentric Video Understanding

As of 7 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 5 inbound Pith citation observations for arXiv:2506.03097.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.03097 v1

Coverage vector

measured 35 of 35 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:14:29.581142Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T05:50:27.250841Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-12T10:41:29.637290Z

Reference resolution

35 of 35 outbound references displayed

  • verified exact0
  • verified fuzzy12
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9cf0adb5-e990-427e-b9ba-fcc2aabeae8c · outbound

This paper cites an unresolved cited work.

EgoVLM: Policy Optimization for Egocentric Video Understanding Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:14:31.912954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:14:26.484755Z digest=sha256:b695d21730c32c76a59f1a4431144338c860a7e57f457768f5a9987a6a8656f0

Observation 7c62262a-f97b-4a42-a49e-e7c8c4d7f46d · outbound

This paper cites GPT-4 Technical Report.

EgoVLM: Policy Optimization for Egocentric Video Understanding GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:14:26.585399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:14:26.585399Z digest=sha256:1ed46fce726ccda21f08b3e229bd30be44e216e9ae328992f4c277c20e94279f

Observation 268c3634-4b6b-4f1d-8dc7-fdbe3453239c · outbound

This paper cites Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs.

EgoVLM: Policy Optimization for Egocentric Video Understanding Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:14:26.665985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:14:26.665985Z digest=sha256:de69fd7c72ac6d88d2f7afc2a2bd54ca7933242abb3f2474d107d51d5d453cd9

Observation 6e54bb2d-fa9d-4c34-aa24-d369516496b3 · outbound

This paper cites Qwen2.5-vl technical report, 2025.

EgoVLM: Policy Optimization for Egocentric Video Understanding Qwen2.5-vl technical report, 2025

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T11:14:26.747337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:14:26.747337Z digest=sha256:53148cfb8b44656faaf1658ac478246573dafa3f112ab4100961c9ccaac8a8c1

Observation 829fa18c-862f-4e6b-8a1a-caaeada02d0f · outbound

This paper cites EgoPlan-Bench: Benchmarking Multimodal Large Language Models for Human-Level Planning.

EgoVLM: Policy Optimization for Egocentric Video Understanding EgoPlan-Bench: Benchmarking Multimodal Large Language Models for Human-Level Planning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:14:26.803299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:14:26.803299Z digest=sha256:5fa365f04a9246bb52594b97dc2a6a603747343f6daf420c5f1d597467ec51f6

Observation 466cd78f-6508-472d-b095-c69c619c2e8e · outbound

This paper cites VidEgoThink: Assessing Egocentric Video Understanding Capabilities for Embodied AI.

EgoVLM: Policy Optimization for Egocentric Video Understanding VidEgoThink: Assessing Egocentric Video Understanding Capabilities for Embodied AI

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:14:26.902583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:14:26.902583Z digest=sha256:3402800d06fa736de7bc036052420ee94d34609829f06630d67572a30f337589

Observation 6531419f-4afd-4eed-8a25-94de004f008c · outbound

This paper cites an unresolved cited work.

EgoVLM: Policy Optimization for Egocentric Video Understanding Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:14:31.779038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:14:26.966208Z digest=sha256:3e90ab8ed324e70249a6c495f5aaefdebf5de23ac2e9207151203efd20c94b78

Observation 7281fecc-e61b-40e9-aba9-dc314e630fbf · outbound

This paper cites ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools.

EgoVLM: Policy Optimization for Egocentric Video Understanding ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:14:27.043868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:14:27.043868Z digest=sha256:8d3c1a614f0a02f37671ec8cee6261152207edb606546896eb14047e81da6557

Observation fd5018df-37a8-459b-bd86-54df58edf779 · outbound

This paper cites GPT-4o System Card.

EgoVLM: Policy Optimization for Egocentric Video Understanding GPT-4o System Card

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:14:27.136926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:14:27.136926Z digest=sha256:845064205811890910f3a628086d31fc0beb03242be020f23693eccebc1b2467

Observation 8e1f2ebc-5652-4ef5-9a09-f499df2d6a3d · outbound

This paper cites Mvbench: A comprehensive multi- modal video understanding benchmark, 2024.

EgoVLM: Policy Optimization for Egocentric Video Understanding Mvbench: A comprehensive multi- modal video understanding benchmark, 2024

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:31.650068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:14:27.251656Z digest=sha256:d02faaf912034f2a9e27fb99697f4f15b831f35afcc86b70795b2e03e794a9dc

Observation 9d089f33-2fa1-42a4-b067-705125e34387 · outbound

This paper cites Dual-Difficulty Curriculum Learning for Direct Preference Optimization.

EgoVLM: Policy Optimization for Egocentric Video Understanding Dual-Difficulty Curriculum Learning for Direct Preference Optimization

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:14:27.351094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:14:27.351094Z digest=sha256:4ffe6ecffd7de169ec5043748fb6868befd35130f5567de1bce0d58ed7f5012c

Observation 94afe2fa-f34c-4ffc-9b8f-dce191566858 · outbound

This paper cites ROUGE: A package for automatic evaluation of summaries.

EgoVLM: Policy Optimization for Egocentric Video Understanding ROUGE: A package for automatic evaluation of summaries

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:31.558170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:14:27.482076Z digest=sha256:6a0b8b610decf165151fff0c835f743391f6de21644085b20efa43314dbc83df

Observation 3094ddd2-5790-45dc-8488-ba0939c6f1c1 · outbound

This paper cites Microsoft coco: Common objects in context.

EgoVLM: Policy Optimization for Egocentric Video Understanding Microsoft coco: Common objects in context

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:14:27.604637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:14:27.604637Z digest=sha256:0357a11ab4b38aa6c09731f948b6241f837c41dc296d577f5f4d2fe3df23a431

Observation 1f733e8e-e71c-4a73-b8e3-a4d59c859591 · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

EgoVLM: Policy Optimization for Egocentric Video Understanding Understanding R1-Zero-Like Training: A Critical Perspective

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:14:27.701631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:14:27.701631Z digest=sha256:99a21b668f85a6e7a546a5b9bd152525bd41ab025fb70ee73b0b2fc7daaaf0f0

Observation 52c4be18-2d30-49e9-ba1d-c1880b371685 · outbound

This paper cites Visual-RFT: Visual Reinforcement Fine-Tuning.

EgoVLM: Policy Optimization for Egocentric Video Understanding Visual-RFT: Visual Reinforcement Fine-Tuning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:14:27.794624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:14:27.794624Z digest=sha256:f1fa54934fbc8bbe2729daccff9cbc34b8a3282ef852cf75737f6fec0cb7a6ed

Observation 32e78173-37f0-4384-9aa4-b0f083b520eb · outbound

This paper cites Reasoning models can be effective without thinking, 2025.

EgoVLM: Policy Optimization for Egocentric Video Understanding Reasoning models can be effective without thinking, 2025

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:31.418970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:14:27.897137Z digest=sha256:afe0e9cf5b5ddada0fb21331124c637ddd2e3ff610f1d9c8378fd7f702603944

Observation a91b1c91-5fe9-4fc6-8fb2-4495f31efc8e · outbound

This paper cites Egoschema: A diagnostic benchmark for very long- form video language understanding.

EgoVLM: Policy Optimization for Egocentric Video Understanding Egoschema: A diagnostic benchmark for very long- form video language understanding

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:31.299087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:14:27.989782Z digest=sha256:10c36ae43a6ac819303afd69b99076382173a372ddcdbb7a6a95ebe2926eac2b

Observation dd2855f9-5bca-45ea-969e-eb567c9996a8 · outbound

This paper cites Training language models to follow instructions with human feedback.

EgoVLM: Policy Optimization for Egocentric Video Understanding Training language models to follow instructions with human feedback

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:14:28.107807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:14:28.107807Z digest=sha256:9c920f2134ee6ed47897c2a94e73f79c57146efca3896c59cb70db4c32043f49

Observation 08c17e3c-3a5d-403c-b440-72646bd5320c · outbound

This paper cites Medvlm-r1: Incentivizing medical reasoning ca- pability of vision-language models (vlms) via reinforcement learning, 2025.

EgoVLM: Policy Optimization for Egocentric Video Understanding Medvlm-r1: Incentivizing medical reasoning ca- pability of vision-language models (vlms) via reinforcement learning, 2025

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:31.178831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:14:28.186843Z digest=sha256:570ccd179bb41f74b2addccc33adc52948696b6a34f44109494bc8c6e932ce79

Observation 298fb0c7-5478-40fc-ad63-b781ba56dcf8 · outbound

This paper cites Learning transferable visual models from natural language supervision.

EgoVLM: Policy Optimization for Egocentric Video Understanding Learning transferable visual models from natural language supervision

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:31.041234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:14:28.261259Z digest=sha256:d1a047ea05b2e7bfa6b1ff10f0449d66916ddde4fa4f2199ac0cbc0dc328145b

Observation 7742fe99-1aa1-48bf-9fde-6792242a6b3b · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

EgoVLM: Policy Optimization for Egocentric Video Understanding Direct preference optimization: Your language model is secretly a reward model

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:14:28.361137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:14:28.361137Z digest=sha256:e021d29856b0643062c9e1fe3a8571b1034c293e31017ab46c68648a5276addd

Observation f9c58009-197e-4dfe-8e46-e8a0dc3e55c6 · outbound

This paper cites Proximal Policy Optimization Algorithms.

EgoVLM: Policy Optimization for Egocentric Video Understanding Proximal Policy Optimization Algorithms

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:14:28.448038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:14:28.448038Z digest=sha256:94011e708f44312f481ef2620403eddc523d8a15c66056a5576e06c2e661f854

Observation e43a7e88-60ed-4972-a6b7-ff2fe5ebbd17 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

EgoVLM: Policy Optimization for Egocentric Video Understanding DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T11:14:28.572354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:14:28.572354Z digest=sha256:24c86f0bcabc8f052c24257a15230db0fffe9e86d2749d882768727320c80bd7

Observation 42623fe7-8624-4ba8-aad3-ed55fcf989cf · outbound

This paper cites an unresolved cited work.

EgoVLM: Policy Optimization for Egocentric Video Understanding Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:14:30.870663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:14:28.653266Z digest=sha256:b66fd7885f7cc4d4b07acd2e5cd243c1a6c4ac2e84231e8752ad623e0260d2a1

Observation 071969d7-a674-4a0c-ae25-d3cf448472a5 · outbound

This paper cites VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model.

EgoVLM: Policy Optimization for Egocentric Video Understanding VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T11:14:28.744187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:14:28.744187Z digest=sha256:b386f269545bab58fef432d9d88357051913a93895478fb1d7b7eff97a467d95

Observation 3badcadc-13ad-476b-9356-b973be507cc3 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

EgoVLM: Policy Optimization for Egocentric Video Understanding Gemini: A Family of Highly Capable Multimodal Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:14:28.861844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:14:28.861844Z digest=sha256:a1ebfed1380aeb25c0a4ae0d5a4a14b31263fe33fa0326206b85421a0557d000

Observation e7ae7185-273b-42c0-a295-50a794bcd589 · outbound

This paper cites Internvideo2.5: Empowering video mllms with long and rich context modeling, 2025.

EgoVLM: Policy Optimization for Egocentric Video Understanding Internvideo2.5: Empowering video mllms with long and rich context modeling, 2025

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T11:14:28.924627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:14:28.924627Z digest=sha256:e4168f513fe23c7c19d978ade66a7419505b126c7c5bb68f12099557083bca75

Observation 350801aa-6cc5-4d4c-a2e3-37346c54cf80 · outbound

This paper cites Wizardlm 2, 2024.

EgoVLM: Policy Optimization for Egocentric Video Understanding Wizardlm 2, 2024

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:30.711275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:14:29.032858Z digest=sha256:c55b341cf46f6f3d084701b6008e7f2e4497cbfc28953e3e03ad122b7e063d35

Observation 7825ea91-ba7e-42fb-a506-0c2806d66248 · outbound

This paper cites ST-Think: How Multimodal Large Language Models Reason About 4D Worlds from Ego-Centric Videos.

EgoVLM: Policy Optimization for Egocentric Video Understanding ST-Think: How Multimodal Large Language Models Reason About 4D Worlds from Ego-Centric Videos

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:14:29.112728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:14:29.112728Z digest=sha256:90db30d866cc9b7f9da81027a9636d568e90d908034af87339909e3541252cef

Observation 819dae82-f2af-4964-b142-c92a2946656e · outbound

This paper cites Egolife: Towards egocentric life assistant.

EgoVLM: Policy Optimization for Egocentric Video Understanding Egolife: Towards egocentric life assistant

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:30.571100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:14:29.229419Z digest=sha256:a9bef23945e8fb59cfc6f9e14533ab9d5e43b7a3e102d7c87e7127bee24f8f57

Observation f48211bc-4251-41c3-bd38-c7686456d4e3 · outbound

This paper cites Mm-ego: Towards build- ing egocentric multimodal llms.

EgoVLM: Policy Optimization for Egocentric Video Understanding Mm-ego: Towards build- ing egocentric multimodal llms

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:30.450637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:14:29.289694Z digest=sha256:f9bd5603530a678448a3c5f3827fe95d5e1b4dac896f0756986059fa7333ac10

Observation 2a23480f-3d85-4f22-964e-c9c8e56192d7 · outbound

This paper cites Video instruction tuning with synthetic data, 2024.

EgoVLM: Policy Optimization for Egocentric Video Understanding Video instruction tuning with synthetic data, 2024

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T11:14:29.397148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:14:29.397148Z digest=sha256:da6193eba418619f20dc9038692e3431a2f7c0ab1ef49d222508fbafe288c82d

Observation 6781ea2e-d1f7-4430-a462-1821dc6db063 · outbound

This paper cites Llamafac- tory: Unified efficient fine-tuning of 100+ language mod- els.

EgoVLM: Policy Optimization for Egocentric Video Understanding Llamafac- tory: Unified efficient fine-tuning of 100+ language mod- els

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:30.310972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:14:29.444276Z digest=sha256:717b6ddd794fe7accf054c59dd3bb8aef2e8d0edfe5105b035aeb2b18b8cf99b

Observation c3d95f10-d7c8-4aaa-b7c9-c88f519f0f44 · outbound

This paper cites R1-zero’s ”aha moment” in visual reasoning on a 2b non-sft model, 2025.

EgoVLM: Policy Optimization for Egocentric Video Understanding R1-zero’s ”aha moment” in visual reasoning on a 2b non-sft model, 2025

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:30.161229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:14:29.544152Z digest=sha256:349b79df3b9c04bec52f3fa148f5f3f1dffc9e53bd52439b5a39b70f6df6b09b

Observation da53d852-81dc-4541-ae7c-4a469c6a3fa0 · outbound

This paper cites Egotextvqa: Towards egocentric scene-text aware video question answering, 2025.

EgoVLM: Policy Optimization for Egocentric Video Understanding Egotextvqa: Towards egocentric scene-text aware video question answering, 2025

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:30.027639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:14:29.581142Z digest=sha256:99be13438382452558c29f74e7fafac856d2d1be87b44b408c94e252d88f8153

Pith citing papers

Observation 358e3010-f86b-4f49-a0c8-0dcd3bb9fe7c · inbound

EgoVITA: Learning to Plan and Verify for Egocentric Video Reasoning cites this paper.

EgoVITA: Learning to Plan and Verify for Egocentric Video Reasoning EgoVLM: Policy Optimization for Egocentric Video Understanding

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T20:52:51.049418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:52:51.049418Z digest=sha256:921ca8880c82848f66af8bdd0e8966e78c5d148dc3107aad16768994d2e6122d

Observation db8a76be-cd62-4481-9ca2-4f1014065a3b · inbound

EgoIntent: A Pre-Outcome Micro-Step Benchmark for Understanding What, Why, and Next cites this paper.

EgoIntent: A Pre-Outcome Micro-Step Benchmark for Understanding What, Why, and Next EgoVLM: Policy Optimization for Egocentric Video Understanding

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T05:50:27.250841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:50:27.250841Z digest=sha256:243fa30483b776175edb90e913b2b4961a498807f9eb93499bbb5b2dcc73a8b8

Observation 1fdb557d-46fb-422e-b3b0-ef98a0901133 · inbound

Robot Learning from Human Videos: A Survey cites this paper.

Robot Learning from Human Videos: A Survey EgoVLM: Policy Optimization for Egocentric Video Understanding

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:41:29.640044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-07T04:55:44.273643Z digest=sha256:8bf88a19fea868e84bdda6f4690e4657549dfc809e1faa9ed7d9de3e9670b79f

Observation 6dbe324b-7bf7-4845-adb5-069be76a7df9 · inbound

Pro$^2$Assist: Continuous Step-aware Proactive Assistance with Multi-modal Egocentric Perception for Long-horizon Procedural Tasks cites this paper.

Pro$^2$Assist: Continuous Step-aware Proactive Assistance with Multi-modal Egocentric Perception for Long-horizon Procedural Tasks EgoVLM: Policy Optimization for Egocentric Video Understanding

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:46:06.581439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T17:16:31.820718Z digest=sha256:58f24f188ffe552c5d48f98fb7b3ea40a3fd104d1cdab3412ee73fd4778ef118

Observation 2ed5ebf4-0079-4353-a6bd-f0edb17c0757 · inbound

Pro$^2$Assist: Continuous Step-aware Proactive Assistance with Multi-modal Egocentric Perception for Long-horizon Procedural Tasks cites this paper.

Pro$^2$Assist: Continuous Step-aware Proactive Assistance with Multi-modal Egocentric Perception for Long-horizon Procedural Tasks EgoVLM: Policy Optimization for Egocentric Video Understanding

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-04T05:19:56.255880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:19:56.255880Z digest=sha256:28f044482219823a556abc9a0ff174fcfc28330acae6553fb38a244174ae347b