Pith. sign in

Paper Citation Record · LEDGER

Pedestrian Intention Prediction via Vision-Language Foundation Models

As of 10 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 1 inbound Pith citation observation for arXiv:2507.04141.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.04141 v1

Coverage vector

measured 36 of 36 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:59:56.791343Z

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-30T13:25:04.053283Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T13:34:40.932917Z

Reference resolution

36 of 36 outbound references displayed

  • verified exact2
  • verified fuzzy25
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 89755a4e-910f-4fd8-b046-ae3d7a248e5b · outbound

This paper cites Autonomous vehicles that interact with pedestrians: A survey of theory and practice,.

Pedestrian Intention Prediction via Vision-Language Foundation Models Autonomous vehicles that interact with pedestrians: A survey of theory and practice,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T19:59:53.815180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:59:53.815180Z digest=sha256:e9d1839bcb641d0b54aba6d10fbf34945385d5e03102e4821c9d602b81f0b3b0

Observation 6cc181ab-3fb6-4ce6-86b4-dfc0be51a1ff · outbound

This paper cites Pedestrian intention prediction: A convolutional bottom-up multi-task approach,.

Pedestrian Intention Prediction via Vision-Language Foundation Models Pedestrian intention prediction: A convolutional bottom-up multi-task approach,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:00:01.306110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:59:53.876159Z digest=sha256:f22eeec691bf08a4d8a6b51be736fbb9dc6254a6ffd50082335a4599871bb0d1

Observation 64c0a2e8-e139-4b94-9995-dd787f946b28 · outbound

This paper cites St crossingpose: A spatial- temporal graph convolutional network for skeleton-based pedestrian crossing intention prediction,.

Pedestrian Intention Prediction via Vision-Language Foundation Models St crossingpose: A spatial- temporal graph convolutional network for skeleton-based pedestrian crossing intention prediction,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:00:01.180728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:59:53.941914Z digest=sha256:c36812d6136538bf9ee4c3d99e2564d48b80c6ec1296bd04d1d2355eb973ece1

Observation 8a87e33a-a222-4467-b198-55377e59f960 · outbound

This paper cites PIP-Net: Pedestrian Intention Prediction in the Wild.

Pedestrian Intention Prediction via Vision-Language Foundation Models PIP-Net: Pedestrian Intention Prediction in the Wild

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:59:57.232162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:59:54.015290Z digest=sha256:891aa74effe206c50a03bb9d46c241d0d406c3f5e901c42925fb68c65f7506d9

Observation 0757d55c-1420-4cf7-b0c6-b743acbad52b · outbound

This paper cites Pedestrian action an- ticipation using contextual feature fusion in stacked rnns,.

Pedestrian Intention Prediction via Vision-Language Foundation Models Pedestrian action an- ticipation using contextual feature fusion in stacked rnns,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:00:01.062175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:59:54.088410Z digest=sha256:06cca3f5c3ea199862b7cf5f661acaeb049d15cc1807bd6e665aec53a15d4338

Observation 1fb433de-cfe6-4c8a-aea3-b8e46effe73e · outbound

This paper cites Do they want to cross? understanding pedestrian intention for behavior prediction,.

Pedestrian Intention Prediction via Vision-Language Foundation Models Do they want to cross? understanding pedestrian intention for behavior prediction,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:00:00.945196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:59:54.171753Z digest=sha256:ce0933243c60bbc5c3d926d6c16e7d89e6c22fbfe070d39f6c8812452423fa56

Observation 096beed4-6d41-4232-b594-91f5044818fc · outbound

This paper cites Multi-modal hybrid architecture for pedestrian action prediction,.

Pedestrian Intention Prediction via Vision-Language Foundation Models Multi-modal hybrid architecture for pedestrian action prediction,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:00:00.783093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:59:54.251824Z digest=sha256:ed699234ae10841a7d9160b0955e82a965baac44d3f629c0cb5d2b57cf27c356

Observation 57272dc5-5f20-416e-9830-2ae0172d7b8e · outbound

This paper cites Pedestrian graph+: A fast pedestrian crossing prediction model based on graph convo- lutional networks,.

Pedestrian Intention Prediction via Vision-Language Foundation Models Pedestrian graph+: A fast pedestrian crossing prediction model based on graph convo- lutional networks,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:00:00.594842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:59:54.354966Z digest=sha256:e633660fe4b9dc3f35f39bd1133860fe44822a526c989de6ce4410e74530d8a7

Observation 6d431e29-b0a2-484d-924a-97647a9e5bcc · outbound

This paper cites Visual reasoning using graph con- volutional networks for predicting pedestrian crossing intention,.

Pedestrian Intention Prediction via Vision-Language Foundation Models Visual reasoning using graph con- volutional networks for predicting pedestrian crossing intention,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:00:00.447152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:59:54.433966Z digest=sha256:0c4bd9683acc0e197e530a0c919fa9a64b7e3f7a52d6cc3768108a83b5a4074b

Observation e4e1f283-16ef-4438-84c3-4698cef4f6e0 · outbound

This paper cites CAPformer: Pedestrian crossing action prediction using transformer,.

Pedestrian Intention Prediction via Vision-Language Foundation Models CAPformer: Pedestrian crossing action prediction using transformer,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:00:00.297573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:59:54.497886Z digest=sha256:7b435ef0f644205a41a37d7a86d715c6a69ad1902c01ef35c1f2ac81db5b4016

Observation 7521fcd7-f346-42b0-82ff-37c9cd83ea1b · outbound

This paper cites Pit: Progressive interaction transformer for pedestrian crossing intention prediction,.

Pedestrian Intention Prediction via Vision-Language Foundation Models Pit: Progressive interaction transformer for pedestrian crossing intention prediction,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:00:00.122113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:59:54.565677Z digest=sha256:f93eaa711f093342937c6e9c04dc840622da1eeec22b162521869d06e5e4457a

Observation 841f712e-7dfd-4388-9ebb-c567f59da6c3 · outbound

This paper cites Predicting pedestrian inten- tions with multimodal intentformer: A co-learning approach,.

Pedestrian Intention Prediction via Vision-Language Foundation Models Predicting pedestrian inten- tions with multimodal intentformer: A co-learning approach,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:59:59.936485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:59:54.646030Z digest=sha256:31c9a37f3e3ba59de1c7cf47b70e219ca6f943d90aeb82b32e06bfb08c333d09

Observation cbc0cc2c-ff16-477a-aa65-247a29a4ad34 · outbound

This paper cites Pedestrian behavior inter- pretation from pose estimation,.

Pedestrian Intention Prediction via Vision-Language Foundation Models Pedestrian behavior inter- pretation from pose estimation,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:59:59.789304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:59:54.713024Z digest=sha256:fe9c191b4328fb3d3efe2794e1026e9014d2a4d6b1c9dda318f8eaab8fc8dc91

Observation e7e25164-28e8-4cf8-9723-5db6a1c92a99 · outbound

This paper cites Multi-scale pedestrian intent prediction using 3d joint information as spatio-temporal representation,.

Pedestrian Intention Prediction via Vision-Language Foundation Models Multi-scale pedestrian intent prediction using 3d joint information as spatio-temporal representation,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:59:59.600266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:59:54.773984Z digest=sha256:c0a87f93e07eb838fd55770145eb3f7cf67fe2a75089a083187631343795c31f

Observation d72d0405-b7da-412a-ab62-1046703f990a · outbound

This paper cites Spatiotemporal relationship reasoning for pedestrian intent prediction,.

Pedestrian Intention Prediction via Vision-Language Foundation Models Spatiotemporal relationship reasoning for pedestrian intent prediction,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T19:59:54.861620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:59:54.861620Z digest=sha256:68a45b89b08c2bca8e5fb18f98ba460a89aa003cecf1875b672b44006f9120af

Observation 42de927b-2ff4-4089-ad57-1ffa7ba398ac · outbound

This paper cites Real-time intent prediction of pedestrians for autonomous ground vehicles via spatio-temporal densenet,.

Pedestrian Intention Prediction via Vision-Language Foundation Models Real-time intent prediction of pedestrians for autonomous ground vehicles via spatio-temporal densenet,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:59:59.442338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:59:54.933786Z digest=sha256:f4861cfb994e3ca75843fec0ee563b92b925f87f388407dfe062e18591a98e06

Observation aa461dfd-b696-4dcc-9b73-896c835a4f53 · outbound

This paper cites Pedestrian-vehicle information modulation for pedestrian crossing intention prediction,.

Pedestrian Intention Prediction via Vision-Language Foundation Models Pedestrian-vehicle information modulation for pedestrian crossing intention prediction,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:59:59.265412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:59:54.998685Z digest=sha256:503ccb284db11c4a020567cb7ee67e871ff7af9f974d12e16c7a1c5ab9233f50

Observation 615b7547-8719-485d-bfb7-e2ac534abdcf · outbound

This paper cites Causal reasoning in typical computer vision tasks,.

Pedestrian Intention Prediction via Vision-Language Foundation Models Causal reasoning in typical computer vision tasks,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:59:59.071147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:59:55.083231Z digest=sha256:f969127a646e3a9df74396e1afff1addd237d2a9827d9c96e533e9b1ecef4b98

Observation 18069a05-b99a-4246-9453-288c0dc92147 · outbound

This paper cites Vision language models in autonomous driving: A survey and outlook,.

Pedestrian Intention Prediction via Vision-Language Foundation Models Vision language models in autonomous driving: A survey and outlook,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:59:58.920039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:59:55.171703Z digest=sha256:031eb9d8225343ef008c7512da9192d90691c9ad90a5b08ba41bee91f786b82f

Observation 35d3e16b-b52f-4d17-b32c-0190c0006e9a · outbound

This paper cites Gpt-4v takes the wheel: Promises and challenges for pedestrian behavior prediction,.

Pedestrian Intention Prediction via Vision-Language Foundation Models Gpt-4v takes the wheel: Promises and challenges for pedestrian behavior prediction,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:59:58.761794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:59:55.267162Z digest=sha256:37f836f784ee3e6da225ff7a36b55ca3e99da2c6322bd0fb54933607bc294ba4

Observation ef2988a4-73fc-4e49-8be5-1a12e03fc931 · outbound

This paper cites Omnipredict: Gpt-4o enhanced multi-modal pedestrian crossing intention prediction.

Pedestrian Intention Prediction via Vision-Language Foundation Models Omnipredict: Gpt-4o enhanced multi-modal pedestrian crossing intention prediction

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:59:58.575555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:59:55.367452Z digest=sha256:7bda0a2b585d48bc4be91b5d327bd065f4af241e202f3839019f9e65cb9e74bd

Observation a05142d2-747f-4bde-8387-19af28ae0283 · outbound

This paper cites Pedvlm: Pedestrian vision language model for intentions prediction,.

Pedestrian Intention Prediction via Vision-Language Foundation Models Pedvlm: Pedestrian vision language model for intentions prediction,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:59:58.398927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:59:55.459977Z digest=sha256:d7a161c677b708303644ba76207b63973aa9cda6cdd19dabb9860d4dc78e1929

Observation bf34fca6-b13d-4b6f-8afd-f1890ef41e1f · outbound

This paper cites Cross or wait? predicting pedestrian interaction outcomes at unsignalized crossings,.

Pedestrian Intention Prediction via Vision-Language Foundation Models Cross or wait? predicting pedestrian interaction outcomes at unsignalized crossings,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:59:58.222821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:59:55.569120Z digest=sha256:09ec58a3d7ae6df5b56d8be9acd58002ca713da29a7666814b0c027b51429f2f

Observation 5fdb116b-8981-422a-8255-343c2ee44d16 · outbound

This paper cites Feature Importance in Pedestrian Intention Prediction: A Context-Aware Review.

Pedestrian Intention Prediction via Vision-Language Foundation Models Feature Importance in Pedestrian Intention Prediction: A Context-Aware Review

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:59:57.055566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:59:55.663331Z digest=sha256:6418586d2a88ad4cba8f968d4868dc358f95b0c5e8f04191cefb56c07ee77aa0

Observation ca7ff6da-26b8-4be8-b4bd-6873e3b2a77b · outbound

This paper cites Hierarchical Prompting Taxonomy: A Universal Evaluation Framework for Large Language Models Aligned with Human Cognitive Principles.

Pedestrian Intention Prediction via Vision-Language Foundation Models Hierarchical Prompting Taxonomy: A Universal Evaluation Framework for Large Language Models Aligned with Human Cognitive Principles

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T19:59:55.734120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:59:55.734120Z digest=sha256:13629bb4a83ff377d4213672b0f5fde977fee681578319f226c4c36ecbe87fb8

Observation fc62114f-4913-4812-9b21-2bc71f52489a · outbound

This paper cites Large Language Models Are Human-Level Prompt Engineers.

Pedestrian Intention Prediction via Vision-Language Foundation Models Large Language Models Are Human-Level Prompt Engineers

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T19:59:55.815565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:59:55.815565Z digest=sha256:e8c24a7b8740b759078fa5be88be3902ae8ca368134f4ca7d737db824bbfaf89

Observation 9ceed48e-7f92-4f3b-9b9f-8c3bd66adac8 · outbound

This paper cites Benchmark for evaluating pedestrian action prediction,.

Pedestrian Intention Prediction via Vision-Language Foundation Models Benchmark for evaluating pedestrian action prediction,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:59:58.017275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:59:55.901830Z digest=sha256:e758bb118726e56f5d75d5885ab36a981eadb3a4c0a6029a8d6fd0de0fcefcfe

Observation 99d7b5b0-b85f-469a-8584-a08fa55707f7 · outbound

This paper cites Better Zero-Shot Reasoning with Role-Play Prompting.

Pedestrian Intention Prediction via Vision-Language Foundation Models Better Zero-Shot Reasoning with Role-Play Prompting

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T19:59:55.992136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:59:55.992136Z digest=sha256:013139cceba182c565537b1e2234089e570570e39c9a5d7b14b0f3ba5409d099

Observation 699aaa9e-581e-4a0c-afd9-1752e25544cb · outbound

This paper cites Good at captioning, bad at counting: Benchmarking GPT-4V on Earth observation data.

Pedestrian Intention Prediction via Vision-Language Foundation Models Good at captioning, bad at counting: Benchmarking GPT-4V on Earth observation data

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T19:59:56.092349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:59:56.092349Z digest=sha256:a8741f5dd680606f49ac4322a17b082e209df91519b0928650f99fa995f4695f

Observation 1cd8d6af-bea9-49c0-8646-1cf809825b86 · outbound

This paper cites Is the pedestrian going to cross? answering by 2d pose estimation,.

Pedestrian Intention Prediction via Vision-Language Foundation Models Is the pedestrian going to cross? answering by 2d pose estimation,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:59:57.884083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:59:56.186360Z digest=sha256:a1fb03a6f7f7dec4a87dae8332acdf9c8512b3b8541f3256ad543da081183fc0

Observation 4e747a18-0a1f-454e-8ab8-2bf16ea658da · outbound

This paper cites Chatgpt: Generative pre-trained transformer,.

Pedestrian Intention Prediction via Vision-Language Foundation Models Chatgpt: Generative pre-trained transformer,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:59:57.715335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:59:56.286479Z digest=sha256:84dfea999e521a2050a8fbc91bd83ee60b479b5ba1789912db43c3463dc70512

Observation 35420cca-34d7-4473-94d1-ff1a8173274b · outbound

This paper cites Agreeing to cross: How drivers and pedestrians communicate,.

Pedestrian Intention Prediction via Vision-Language Foundation Models Agreeing to cross: How drivers and pedestrians communicate,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:59:57.594088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:59:56.376470Z digest=sha256:52381422ba079a76fb6b4e8da295fb05fef02732a11f626680440dfcf49520ba

Observation c5cf1eb8-ead7-4137-abbd-bfce2e73541e · outbound

This paper cites PIE: A large-scale dataset and models for pedestrian intention estimation and trajectory prediction,.

Pedestrian Intention Prediction via Vision-Language Foundation Models PIE: A large-scale dataset and models for pedestrian intention estimation and trajectory prediction,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:59:57.417787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:59:56.479906Z digest=sha256:ffa4dc57fa077217f2cc7d9d8d75116803d14ede2661f1b98128cccfd8899409

Observation 398caa99-291c-4851-bb0a-8b63ce9b5728 · outbound

This paper cites GPT-4 Technical Report.

Pedestrian Intention Prediction via Vision-Language Foundation Models GPT-4 Technical Report

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T19:59:56.585090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:59:56.585090Z digest=sha256:3f12ba5c2fe10acbd6f6574e5666f3f6d9bc84f84e15c62dda9e16fb209c4648

Observation 768bfb91-1616-4cd3-82f5-8b553a37cf52 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Pedestrian Intention Prediction via Vision-Language Foundation Models MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T19:59:56.685176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:59:56.685176Z digest=sha256:2d8c709dabd1d5b77fc0a96b70ec88c95382485d9846bed9432d5b348dacbfeb

Observation f989038e-d8a4-4f79-a1a3-12bf90ba02b3 · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

Pedestrian Intention Prediction via Vision-Language Foundation Models LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T19:59:56.791343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:59:56.791343Z digest=sha256:b655b1d637de62a050d21b3a501210335b4a88f972313eb65d7a205c9c92feb9

Pith citing papers

Observation 0b661ac7-40da-44c8-b177-06d0579874c9 · inbound

PEDESTRIANQA: A Benchmark for Vision-Language Models on Pedestrian Intention and Trajectory Prediction cites this paper.

PEDESTRIANQA: A Benchmark for Vision-Language Models on Pedestrian Intention and Trajectory Prediction Pedestrian Intention Prediction via Vision-Language Foundation Models

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-06-30T13:34:40.934373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T13:25:04.053283Z digest=sha256:faa2dda2c7ace7c9571e1c6b48ab4d98499a36d284a0afcd676e971a76f92663