Pith. sign in

Paper Citation Record · LEDGER

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models

As of 8 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 0 inbound Pith citation observations for arXiv:2606.09142.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.09142 v1

Coverage vector

measured 53 of 53 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-27T17:13:04.236686Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

53 of 53 outbound references displayed

  • verified exact10
  • verified fuzzy0
  • unresolved40
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b3b16424-1d11-4892-a331-214cc0b4fcae · outbound

This paper cites Pedestrian Behavior Prediction Using Deep Learning Methods for Urban Scenarios: A Review,.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models Pedestrian Behavior Prediction Using Deep Learning Methods for Urban Scenarios: A Review,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-27T17:13:04.236686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:023d94231a86550ab7d18d56eeff3a43f6edf9bc1e437be93db38841b0fb1d3d

Observation bed714ee-335a-474a-9ccf-df0bec14183d · outbound

This paper cites Predicting Pedestrian Crossing Intention in Autonomous Vehicles: A Review,.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models Predicting Pedestrian Crossing Intention in Autonomous Vehicles: A Review,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-06-27T17:13:04.236686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:0ca77096fd7de7d1e0705537929466cb96fb0bd72d4ab64275e66e0ab7f776d0

Observation 383f2029-8a44-4880-9512-a11756f0e28d · outbound

This paper cites EgoNav: Egocentric Scene-aware Human Trajectory Prediction.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models EgoNav: Egocentric Scene-aware Human Trajectory Prediction

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T00:27:29.693295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:d9235787ae0c14a4778a55bedb6883a8cccdb6eef2880ee5c1eb3cb92db8f590

Observation 07d7f1f7-3e63-4e8b-b169-94183c628fef · outbound

This paper cites Bridging Perspectives: A Survey on Cross-view Collaborative Intelligence with Egocentric-Exocentric Vision,.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models Bridging Perspectives: A Survey on Cross-view Collaborative Intelligence with Egocentric-Exocentric Vision,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-27T17:13:04.236686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:aa1ecbf566fd002bacdb9f21879ad29b9a0facc934dbabeae34e22eddf560a30

Observation ea260b1b-b7f3-4009-b230-f2259288b332 · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video,.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models Ego4d: Around the world in 3,000 hours of egocentric video,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-06-27T17:13:04.236686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:38db36daf6ea8f2831c95d7e0a13bab121421c82bf4e7234a6ff60208b6de20e

Observation 4726cbaa-423f-4dfa-bd47-047b4cc63bcb · outbound

This paper cites An outlook into the future of egocentric vision,.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models An outlook into the future of egocentric vision,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-27T17:13:04.236686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:37a963b214c4cfa9bee882dc8ae9135b0b56bf83ce639b4f8c29a0c507d79048

Observation e86667fc-7de0-4676-9c80-3a6280931162 · outbound

This paper cites Egolife: Towards egocentric life assistant,.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models Egolife: Towards egocentric life assistant,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-27T17:13:04.236686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:23717df244dd0bab6fd2cb4b0fcdf7137f7cb76e3b599f6fd29a477c62e5a2de

Observation 045b6d09-7e66-4972-bcb1-c3fa6acd7e14 · outbound

This paper cites EgoCogNav: Cognition-aware Human Egocentric Navigation.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models EgoCogNav: Cognition-aware Human Egocentric Navigation

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T00:27:29.660613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:2a991ea0936f702dfdfbe26786232f429a9117a342de3a14ac75bf7516c9cdf9

Observation b6facd0e-cec0-4beb-b701-c4b755743cf4 · outbound

This paper cites Lookout: Real-world humanoid egocentric navigation,.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models Lookout: Real-world humanoid egocentric navigation,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-27T17:13:04.236686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:3777637c4e6e5356f7ddd9a97029794c5c0699cb1db19bae46c04a557ff0d960

Observation a8e202bb-8cc7-4911-81f0-d38c8fe1147c · outbound

This paper cites HEADS-UP: Head-Mounted Egocentric Dataset for Trajectory Prediction in Blind Assistance Systems.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models HEADS-UP: Head-Mounted Egocentric Dataset for Trajectory Prediction in Blind Assistance Systems

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:27:29.675869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:0274417cfa120550e994cb929cb0ad615230227673ef6b57e7271ed131ce9d6a

Observation cb1ebd8b-00a6-4617-8d11-29e556b205ee · outbound

This paper cites Egocentric human trajectory forecasting with a wearable camera and multi-modal fusion,.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models Egocentric human trajectory forecasting with a wearable camera and multi-modal fusion,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-27T17:13:04.236686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:87915be753a25b8c73e19c0f22af2045310691cd79a2b7377c5fec7a7db5e380

Observation 85033b4b-6be4-41f6-99b7-c197bf8f19e4 · outbound

This paper cites KrishnaCam: Using a longitudinal, single-person, egocentric dataset for scene understanding tasks,.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models KrishnaCam: Using a longitudinal, single-person, egocentric dataset for scene understanding tasks,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-27T17:13:04.236686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:2878c1b33953fca69c5b047a85ada3d6f9f0458b3afc4c967f77c66743c3c572

Observation 833541b9-b19d-4a4c-8302-bd06132b7b72 · outbound

This paper cites Egocentric future localization,.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models Egocentric future localization,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-27T17:13:04.236686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:6e8f3ae3489268952b78c186c4f957c38e442cd10fe78e2393d181b91c0f5028

Observation 1078c9ee-9460-4c66-845d-30830139088d · outbound

This paper cites Pedestrian intention prediction for autonomous vehicles: A comprehensive survey,.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models Pedestrian intention prediction for autonomous vehicles: A comprehensive survey,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-27T17:13:04.236686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:00840fdca50d69db088e7d3447825599c49cc3ddc4bbf5325668525beb5d709f

Observation 0915f032-7cf4-4331-a100-210cd1ed0c95 · outbound

This paper cites GPT-4V Takes the Wheel: Promises and Challenges for Pedestrian Behavior Prediction,.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models GPT-4V Takes the Wheel: Promises and Challenges for Pedestrian Behavior Prediction,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-27T17:13:04.236686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:0e42cefe06f64f6c88fb421af43a452991f0264e25c8dfee0d6d8f670a470efa

Observation d64df1dd-9b5c-48ef-81d1-53719e388f7b · outbound

This paper cites OmniPredict: GPT-4o Enhanced Multi-modal Pedestrian Crossing Intention Prediction,.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models OmniPredict: GPT-4o Enhanced Multi-modal Pedestrian Crossing Intention Prediction,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-06-27T17:13:04.236686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:a77aaadde90effc66b9cb52c61e8c9c796e50fa0e2cf12a1a560a38a517fba38

Observation 451f2793-aea3-46ec-a694-3985bab79be0 · outbound

This paper cites Seeing beyond frames: Zero-shot pedestrian intention prediction with raw temporal video and multimodal cues,.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models Seeing beyond frames: Zero-shot pedestrian intention prediction with raw temporal video and multimodal cues,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-27T17:13:04.236686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:3f6f68f3f21ade0ef2fd72756542d383956cec4153bc1792558defcc0953d5f3

Observation f3a739d3-43fe-47ca-83a2-37f147ebb3be · outbound

This paper cites Pedestrian Intention Prediction via Vision-Language Foundation Models,.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models Pedestrian Intention Prediction via Vision-Language Foundation Models,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-27T17:13:04.236686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:63dd4a359f3fa630217d93d95d243088c6e5a66798413d20d6891b2870cfb3a8

Observation ba4a1df2-e453-46fe-a49b-631b7d2dba3b · outbound

This paper cites Optimizing Vision-Language Model for Road Crossing Intention Estimation,.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models Optimizing Vision-Language Model for Road Crossing Intention Estimation,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-27T17:13:04.236686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:36255b24d41737bf90568b3d046ab0de83a2280c233a040743dd10194d7af6ad

Observation a2c2e360-22a0-4e17-8e67-e9159ec078d0 · outbound

This paper cites Pedestrian Vision Language Model for Intentions Prediction,.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models Pedestrian Vision Language Model for Intentions Prediction,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-27T17:13:04.236686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:cc6f3485d580e49c86bb55afba4a3850050740c07cfb68cda059ef14bd799f2e

Observation aee234ee-929e-4b32-8ae2-c61d2a8c9c63 · outbound

This paper cites Application of Vision-Language Model to Pedestrians Behavior and Scene Understanding in Autonomous Driving.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models Application of Vision-Language Model to Pedestrians Behavior and Scene Understanding in Autonomous Driving

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:27:29.690221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:f96745cb29dd12b65a44d05df96411f5c0d5fcf48b85bb4cbf68421d8f30db19

Observation 30a7d0f3-9a98-4bf9-84c2-d60d19d40505 · outbound

This paper cites Vlmped-cot: A large vision-language model with chain-of-thought mechanism for pedestrian crossing intention prediction,.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models Vlmped-cot: A large vision-language model with chain-of-thought mechanism for pedestrian crossing intention prediction,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-06-27T17:13:04.236686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:11347f6e09b0f239b68fcf4258145c9c80a056babd338a7ddbc32405133a9e02

Observation 3f78746e-079d-4487-a15a-6bf6ffe3da56 · outbound

This paper cites Scaling Egocentric Vision: The EPIC-KITCHENS Dataset.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models Scaling Egocentric Vision: The EPIC-KITCHENS Dataset

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-07-03T00:27:29.665460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:707365282e7164357da1ef5901b5eed9bf8156e76630120d5caba4cc15bcdfea

Observation d07f28ea-020a-470b-a181-e3513eaab16b · outbound

This paper cites Actor and observer: Joint modeling of first and third-person videos,.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models Actor and observer: Joint modeling of first and third-person videos,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-06-27T17:13:04.236686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:b45e37d088714917b719dd37406f0ac36d37c09564f619dcf9663eb80abdf2d9

Observation d6876619-31bf-4cd8-9876-0942c10bed97 · outbound

This paper cites Egovlpv2: Egocentric video-language pre-training with fusion in the backbone,.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models Egovlpv2: Egocentric video-language pre-training with fusion in the backbone,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-06-27T17:13:04.236686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:3a5ae0c2977c5f1da58edefce1e4e00ff9260fa12fe9f8b397648bf5bbe4d9f3

Observation 238674b4-3214-4c28-90ac-257b3e17703c · outbound

This paper cites EgoVideo: Exploring Egocentric Foundation Model and Downstream Adaptation.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models EgoVideo: Exploring Egocentric Foundation Model and Downstream Adaptation

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:27:29.686182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:c8561e793fb7857b07e06795cadf1952c8956f963760fed4fe385e1add066230

Observation 08552ab6-46d6-4623-be20-038fdb3328b9 · outbound

This paper cites Video question answering: Datasets, algorithms and challenges,.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models Video question answering: Datasets, algorithms and challenges,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-06-27T17:13:04.236686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:663b2ba15ddf112d73d1d4ca1a65cf1854bbee66913618fac1058650af080c9a

Observation 84b83758-0e70-4074-8500-882c4cf276fa · outbound

This paper cites Video question answering via gradually refined attention over ap- pearance and motion,.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models Video question answering via gradually refined attention over ap- pearance and motion,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-06-27T17:13:04.236686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:4f46c5c24b41c884002fc3a9d89633653963993ec258a15cf57348c5157df129

Observation 526e1b70-464a-40e5-b994-342bd2f59207 · outbound

This paper cites Tgif-qa: Toward spatio- temporal reasoning in visual question answering,.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models Tgif-qa: Toward spatio- temporal reasoning in visual question answering,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-06-27T17:13:04.236686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:1cfd85c92b3daaadeeaf1c96d4f5927cc60a1e87f8fb2ae188a1ad410c7a962f

Observation 46884596-2fe6-4aca-bb42-45f976ef3b8e · outbound

This paper cites Activitynet-qa: A dataset for understanding complex web videos via question answering,.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models Activitynet-qa: A dataset for understanding complex web videos via question answering,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-06-27T17:13:04.236686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:8784fb936bfc9b6b859ea1b7307a0a513e4630597df6f1489ed562802fc7daa2

Observation d53d2105-5713-4960-a549-09c8bdf2b330 · outbound

This paper cites Next-qa: Next phase of question-answering to explaining temporal actions,.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models Next-qa: Next phase of question-answering to explaining temporal actions,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-06-27T17:13:04.236686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:b35496979f5d850fcc9594530fb3354155360a11e6e409d7c883fdc2ae7bd0bc

Observation 11f21bec-81d8-46b3-a156-6b2ddc5f1c30 · outbound

This paper cites Agqa: A benchmark for compositional spatio-temporal reasoning,.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models Agqa: A benchmark for compositional spatio-temporal reasoning,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-06-27T17:13:04.236686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:d84e9bf7d3ae889de1f38a39c217d6dff2b2805242ccc2d65ad9eb8e1d517206

Observation 0583cc37-0857-4cdb-9046-52b96999dc56 · outbound

This paper cites From representation to reasoning: To- wards both evidence and commonsense reasoning for video question- answering,.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models From representation to reasoning: To- wards both evidence and commonsense reasoning for video question- answering,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-06-27T17:13:04.236686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:b9623009301e7bb307ba5a6bc41ec9dce8dcb25d5a03b0447d41d29a5ac0b01f

Observation 8b5648aa-5505-41a4-b356-09ca8883ed37 · outbound

This paper cites Egotaskqa: Understanding human tasks in egocentric videos,.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models Egotaskqa: Understanding human tasks in egocentric videos,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-06-27T17:13:04.236686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:c91d7edf2600e27613aa849dcee0643fa120b1a0df5db1ebac5df53901d9cc77

Observation 394094fd-5615-476e-bf62-a9a7fce3d8fe · outbound

This paper cites Intentqa: Context-aware video intent reasoning,.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models Intentqa: Context-aware video intent reasoning,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-06-27T17:13:04.236686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:089de7ce076097460781a761dec4498b0c5a8c901b658dd56ebd9bd4823f3f8f

Observation 4dab975f-bb5c-49d6-8f2a-dbbb3f5f33b4 · outbound

This paper cites In the eye of mllm: Benchmarking egocentric video intent understanding with gaze-guided prompting.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models In the eye of mllm: Benchmarking egocentric video intent understanding with gaze-guided prompting

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:27:29.670686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:4bb31608b9dc6947236bd1dcccffd204177b12466e2597fd64b703732448e097

Observation 531b9dbd-e5fa-4fc7-9901-f87aa630568e · outbound

This paper cites Qwen3 Technical Report.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models Qwen3 Technical Report

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-07-03T00:27:29.667996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:21aa8d69a6184444405c771b204e426e65cd1034e49da6cae5bdb9aed2261342

Observation 73064eea-3446-4818-a45d-46df9badeba7 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-07-03T00:27:29.680667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:f4e9ae4c20a75610d9cb7d0bbb9c3a52a26bcdbdd492a8561dda6968e8302464

Observation 791ddcdc-faf3-4be6-b06e-8094bc62c4c1 · outbound

This paper cites Grounded question-answering in long egocentric videos,.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models Grounded question-answering in long egocentric videos,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-06-27T17:13:04.236686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:5987f56e455b5388f424855892320cd1285b62a2c7232fcdafb7aeab3e8653db

Observation 6007c9d7-e2e1-48fc-8187-659d98fae112 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models,.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models Chain-of-thought prompting elicits reasoning in large language models,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-06-27T17:13:04.236686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:e1d447af38ad9c3f37a6cb0d264863afd87abca2b44a7a6b9a583c3dda25f029

Observation 5875e811-4bf9-452d-91b9-2b0981eff465 · outbound

This paper cites Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-07-03T00:27:29.678342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:c3487cb56e49c1fd47906cf565c5902bfb70ebfb38da70d02bafd432dab6e744

Observation f58249e3-3ca7-4fa8-84bf-4296dfe282d1 · outbound

This paper cites Fine-grained visual prompting,.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models Fine-grained visual prompting,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-06-27T17:13:04.236686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:5a29eb50ea1bd4b1cdaaa18a78900c5cc36482d1f4af287b35d046783dd3fd7e

Observation 384b6566-a00f-402e-abaf-f675896a50aa · outbound

This paper cites Large language models are zero-shot reasoners,.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models Large language models are zero-shot reasoners,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-06-27T17:13:04.236686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:ba351755d2bb55d2b3f18511dbc70eb3725acb8bb02f790ac21ac41b6410d0df

Observation cb17fc77-4a29-4ce5-b5d5-a2ca876ba246 · outbound

This paper cites Grounding dino: Marrying dino with grounded pre-training for open-set object detection,.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models Grounding dino: Marrying dino with grounded pre-training for open-set object detection,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-06-27T17:13:04.236686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:0960cc1c5cc5d4896966c90136bf7999e4d93d093987a06fbec2f196037ebceb

Observation 665f8116-a965-402a-8e22-270f8ac8d387 · outbound

This paper cites Simple online and realtime tracking,.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models Simple online and realtime tracking,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-06-27T17:13:04.236686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:83a0c31f195fe8166648d6ea154d9a49d91a6fa595be3b92d70222e66a3dc4e2

Observation 39d58828-2788-459b-b4a4-2623688d79b3 · outbound

This paper cites Analyzing the behaviors of pedestrians and cyclists in interactions with autonomous systems using controlled experiments: A literature review,.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models Analyzing the behaviors of pedestrians and cyclists in interactions with autonomous systems using controlled experiments: A literature review,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-06-27T17:13:04.236686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:df48062e39aa458b99e0d2fef0e9ebba8b849318fdbe42f876bf9e3b63464fe0

Observation 0fb200aa-dc60-4f49-a0d5-c456b8f9001d · outbound

This paper cites Challenges and trends in egocentric vision: A survey,.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models Challenges and trends in egocentric vision: A survey,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-06-27T17:13:04.236686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:a69feb4df7de0734eedea0a282c0447797c8372654e1e52cf0b5bdf0065e8f73

Observation bbb3c8c5-db49-4ea4-98f1-cfa1e63911b3 · outbound

This paper cites Eye Gaze-Informed and Context-Aware Pedestrian Trajectory Prediction in Shared Spaces with Automated Shuttles: A Virtual Reality Study.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models Eye Gaze-Informed and Context-Aware Pedestrian Trajectory Prediction in Shared Spaces with Automated Shuttles: A Virtual Reality Study

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-07-03T00:27:29.662992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:4473c1f598ef341be74cfbb74e28b13cc74bf3f19fff60db6a19e0818d70d3ce

Observation 6a9dcb69-82a4-4c93-a649-fda35ff0141c · outbound

This paper cites Lora: Low-rank adaptation of large language models.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models Lora: Low-rank adaptation of large language models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-06-27T17:13:04.236686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:6657c7b4619ff0881402b581116e4d5ac378124036a253df54abf1229d04d282

Observation fd29ecfd-42f1-46db-84c6-7f9c799a6921 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Su- pervision,.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models Learning Transferable Visual Models From Natural Language Su- pervision,

Reference 51

Resolution
unresolved
no resolver link, observed 2026-06-27T17:13:04.236686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:9c7c1c1250ce4a906873ed0bcc24d36d7d760fa7f6653cbb4008190faf53e2b3

Observation a80980e7-0192-4524-a0ef-2b0480689392 · outbound

This paper cites Advancing Egocentric Video Question Answering with Multimodal Large Language Models.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models Advancing Egocentric Video Question Answering with Multimodal Large Language Models

Reference 52

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T00:27:29.683514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:05fe7449dfc672fa325b8c5ba96d9cc2fe6f7ac7f1396de8216699e3bfdbf25e

Observation 28a30740-3f42-43dd-bc95-8acb2038985a · outbound

This paper cites LatentOmni: Rethinking Omni-Modal Understanding via Unified Audio-Visual Latent Reasoning.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models LatentOmni: Rethinking Omni-Modal Understanding via Unified Audio-Visual Latent Reasoning

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-07-03T00:27:29.673053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:c3e7d3b2b21911a01c899e8216a43128f001904a93a09f2f4fe78a034e669561

Observation 9494379a-d940-41a3-9c29-07a85af5f911 · outbound

This paper cites Chrono: A simple blueprint for representing time in mllms,.

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models Chrono: A simple blueprint for representing time in mllms,

Reference 54

Resolution
unresolved
no resolver link, observed 2026-06-27T17:13:04.236686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T17:13:04.236686Z digest=sha256:bcc75205bbabf26371e58125a10e16dd55915aacaaf1519a2dd296164885f13d

Pith citing papers

No inbound Pith citation observations are available.