Pith. sign in

Paper Citation Record · LEDGER

Pedestrian Intention Prediction via Vision-Language Foundation Models

As of 7 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 1 inbound Pith citation observation for arXiv:2507.04141.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.04141 v1

Coverage vector

measured 36 of 36 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:59:56.791343Z

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-30T13:25:04.053283Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T13:34:40.932917Z

Reference resolution

36 of 36 outbound references displayed

  • verified exact2
  • verified fuzzy25
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 89755a4e-910f-4fd8-b046-ae3d7a248e5b · outbound

This paper cites Autonomous vehicles that interact with pedestrians: A survey of theory and practice,.

Pedestrian Intention Prediction via Vision-Language Foundation Models Autonomous vehicles that interact with pedestrians: A survey of theory and practice,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T19:59:53.815180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:59:53.815180Z digest=sha256:454c0526891d053914876144bc8ce49c58c0c92abeb3272a77367981de8ffd32

Observation 6cc181ab-3fb6-4ce6-86b4-dfc0be51a1ff · outbound

This paper cites Pedestrian intention prediction: A convolutional bottom-up multi-task approach,.

Pedestrian Intention Prediction via Vision-Language Foundation Models Pedestrian intention prediction: A convolutional bottom-up multi-task approach,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:00:01.306110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:59:53.876159Z digest=sha256:7d4f394ae7a121cbb3a5716869da1ab0ecf815f9ebba0d00e23b59b6a8ec4a95

Observation 64c0a2e8-e139-4b94-9995-dd787f946b28 · outbound

This paper cites St crossingpose: A spatial- temporal graph convolutional network for skeleton-based pedestrian crossing intention prediction,.

Pedestrian Intention Prediction via Vision-Language Foundation Models St crossingpose: A spatial- temporal graph convolutional network for skeleton-based pedestrian crossing intention prediction,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:00:01.180728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:59:53.941914Z digest=sha256:1ab32dbf6fc4970dcc3fde17d39ad1fd81aebaec61cad635312bd3e5da40d5e3

Observation 8a87e33a-a222-4467-b198-55377e59f960 · outbound

This paper cites PIP-Net: Pedestrian Intention Prediction in the Wild.

Pedestrian Intention Prediction via Vision-Language Foundation Models PIP-Net: Pedestrian Intention Prediction in the Wild

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:59:57.232162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:59:54.015290Z digest=sha256:5423372d3bae8fd80d0a26dd3817a92b60886f327a8308802d24d22bf72514dc

Observation 0757d55c-1420-4cf7-b0c6-b743acbad52b · outbound

This paper cites Pedestrian action an- ticipation using contextual feature fusion in stacked rnns,.

Pedestrian Intention Prediction via Vision-Language Foundation Models Pedestrian action an- ticipation using contextual feature fusion in stacked rnns,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:00:01.062175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:59:54.088410Z digest=sha256:a8df1f5453230d9cb426acb79340723766f8fb2707efccfc25c3c9394f672b83

Observation 1fb433de-cfe6-4c8a-aea3-b8e46effe73e · outbound

This paper cites Do they want to cross? understanding pedestrian intention for behavior prediction,.

Pedestrian Intention Prediction via Vision-Language Foundation Models Do they want to cross? understanding pedestrian intention for behavior prediction,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:00:00.945196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:59:54.171753Z digest=sha256:ac098f9e38a03f07d9ed4b030d9ddb3ebd3786a082922d5e6a27924c7284ee46

Observation 096beed4-6d41-4232-b594-91f5044818fc · outbound

This paper cites Multi-modal hybrid architecture for pedestrian action prediction,.

Pedestrian Intention Prediction via Vision-Language Foundation Models Multi-modal hybrid architecture for pedestrian action prediction,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:00:00.783093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:59:54.251824Z digest=sha256:4a61f334c6af32f542229cdf3f50aa229475987f77c9c4891cfc64bbea262020

Observation 57272dc5-5f20-416e-9830-2ae0172d7b8e · outbound

This paper cites Pedestrian graph+: A fast pedestrian crossing prediction model based on graph convo- lutional networks,.

Pedestrian Intention Prediction via Vision-Language Foundation Models Pedestrian graph+: A fast pedestrian crossing prediction model based on graph convo- lutional networks,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:00:00.594842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:59:54.354966Z digest=sha256:ad6c0b65ff917fdd8d43194cb3a1375fd2622dec04409eee4d8aced6cd30143c

Observation 6d431e29-b0a2-484d-924a-97647a9e5bcc · outbound

This paper cites Visual reasoning using graph con- volutional networks for predicting pedestrian crossing intention,.

Pedestrian Intention Prediction via Vision-Language Foundation Models Visual reasoning using graph con- volutional networks for predicting pedestrian crossing intention,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:00:00.447152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:59:54.433966Z digest=sha256:fe1d47cda183458643b99b323c12a9ccbfa8be6c2f04c873bf2ca1fad6704713

Observation e4e1f283-16ef-4438-84c3-4698cef4f6e0 · outbound

This paper cites CAPformer: Pedestrian crossing action prediction using transformer,.

Pedestrian Intention Prediction via Vision-Language Foundation Models CAPformer: Pedestrian crossing action prediction using transformer,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:00:00.297573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:59:54.497886Z digest=sha256:99446f1c55d13fffba7ac19cbc1254fe1bfa020afa1687c648ab0b4bd96e0461

Observation 7521fcd7-f346-42b0-82ff-37c9cd83ea1b · outbound

This paper cites Pit: Progressive interaction transformer for pedestrian crossing intention prediction,.

Pedestrian Intention Prediction via Vision-Language Foundation Models Pit: Progressive interaction transformer for pedestrian crossing intention prediction,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:00:00.122113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:59:54.565677Z digest=sha256:2f805b8f03cb72ff8fd6c10b092c3582721bf18ac2fbca416874124a1cc907fc

Observation 841f712e-7dfd-4388-9ebb-c567f59da6c3 · outbound

This paper cites Predicting pedestrian inten- tions with multimodal intentformer: A co-learning approach,.

Pedestrian Intention Prediction via Vision-Language Foundation Models Predicting pedestrian inten- tions with multimodal intentformer: A co-learning approach,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:59:59.936485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:59:54.646030Z digest=sha256:f8e149ae64aba4e5342edfd3f32de770d445491f5ba60bfad09fcd46c3c5fa5f

Observation cbc0cc2c-ff16-477a-aa65-247a29a4ad34 · outbound

This paper cites Pedestrian behavior inter- pretation from pose estimation,.

Pedestrian Intention Prediction via Vision-Language Foundation Models Pedestrian behavior inter- pretation from pose estimation,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:59:59.789304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:59:54.713024Z digest=sha256:efc7c618f8328c08f5f5a0fd0326326db86f0027eada8078185a99270e283f6b

Observation e7e25164-28e8-4cf8-9723-5db6a1c92a99 · outbound

This paper cites Multi-scale pedestrian intent prediction using 3d joint information as spatio-temporal representation,.

Pedestrian Intention Prediction via Vision-Language Foundation Models Multi-scale pedestrian intent prediction using 3d joint information as spatio-temporal representation,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:59:59.600266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:59:54.773984Z digest=sha256:1be8692d59267a5cb198951ad5ffcd8bcd9dd20b109e1d8b2cfc905ad08ab022

Observation d72d0405-b7da-412a-ab62-1046703f990a · outbound

This paper cites Spatiotemporal relationship reasoning for pedestrian intent prediction,.

Pedestrian Intention Prediction via Vision-Language Foundation Models Spatiotemporal relationship reasoning for pedestrian intent prediction,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T19:59:54.861620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:59:54.861620Z digest=sha256:7bb70f7d1cb9cf3aa65c232ef101c03888c585942699fbc5ed522a339b707691

Observation 42de927b-2ff4-4089-ad57-1ffa7ba398ac · outbound

This paper cites Real-time intent prediction of pedestrians for autonomous ground vehicles via spatio-temporal densenet,.

Pedestrian Intention Prediction via Vision-Language Foundation Models Real-time intent prediction of pedestrians for autonomous ground vehicles via spatio-temporal densenet,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:59:59.442338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:59:54.933786Z digest=sha256:bb05e4669f2d43b4e67d2c826a6b73ddaa1ebe6be5943e76b731b7597bf1ad89

Observation aa461dfd-b696-4dcc-9b73-896c835a4f53 · outbound

This paper cites Pedestrian-vehicle information modulation for pedestrian crossing intention prediction,.

Pedestrian Intention Prediction via Vision-Language Foundation Models Pedestrian-vehicle information modulation for pedestrian crossing intention prediction,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:59:59.265412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:59:54.998685Z digest=sha256:a1dbb0267cf2b82d0740fd7552f01a80fc679b2a22bde84a0207d211dfd08aad

Observation 615b7547-8719-485d-bfb7-e2ac534abdcf · outbound

This paper cites Causal reasoning in typical computer vision tasks,.

Pedestrian Intention Prediction via Vision-Language Foundation Models Causal reasoning in typical computer vision tasks,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:59:59.071147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:59:55.083231Z digest=sha256:a28e5214e2e32d93214faa274556f86b5444bc330c564258fa09d0a0461d9e27

Observation 18069a05-b99a-4246-9453-288c0dc92147 · outbound

This paper cites Vision language models in autonomous driving: A survey and outlook,.

Pedestrian Intention Prediction via Vision-Language Foundation Models Vision language models in autonomous driving: A survey and outlook,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:59:58.920039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:59:55.171703Z digest=sha256:a85855da926d2596a677dfeaf434fda7db7842f3828b1596d485fb3b71fca769

Observation 35d3e16b-b52f-4d17-b32c-0190c0006e9a · outbound

This paper cites Gpt-4v takes the wheel: Promises and challenges for pedestrian behavior prediction,.

Pedestrian Intention Prediction via Vision-Language Foundation Models Gpt-4v takes the wheel: Promises and challenges for pedestrian behavior prediction,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:59:58.761794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:59:55.267162Z digest=sha256:5b24ae9a7f75b5ca69dd3ccc8573142ea65129fc4f01c1f8900e14e3bf55e030

Observation ef2988a4-73fc-4e49-8be5-1a12e03fc931 · outbound

This paper cites Omnipredict: Gpt-4o enhanced multi-modal pedestrian crossing intention prediction.

Pedestrian Intention Prediction via Vision-Language Foundation Models Omnipredict: Gpt-4o enhanced multi-modal pedestrian crossing intention prediction

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:59:58.575555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:59:55.367452Z digest=sha256:3a0c7e9d766eb656c50531d90376f78b29430b761d7ae0a29e55d48697e6be22

Observation a05142d2-747f-4bde-8387-19af28ae0283 · outbound

This paper cites Pedvlm: Pedestrian vision language model for intentions prediction,.

Pedestrian Intention Prediction via Vision-Language Foundation Models Pedvlm: Pedestrian vision language model for intentions prediction,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:59:58.398927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:59:55.459977Z digest=sha256:edd14d1e206a745bbbc2b01710f7d6df484e3254c2971e599bdf11ade00a6fb1

Observation bf34fca6-b13d-4b6f-8afd-f1890ef41e1f · outbound

This paper cites Cross or wait? predicting pedestrian interaction outcomes at unsignalized crossings,.

Pedestrian Intention Prediction via Vision-Language Foundation Models Cross or wait? predicting pedestrian interaction outcomes at unsignalized crossings,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:59:58.222821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:59:55.569120Z digest=sha256:05dc9228da87d80c89e37cb1489619f1caecc0b3b77f7c44ba02cbf3765d4790

Observation 5fdb116b-8981-422a-8255-343c2ee44d16 · outbound

This paper cites Feature Importance in Pedestrian Intention Prediction: A Context-Aware Review.

Pedestrian Intention Prediction via Vision-Language Foundation Models Feature Importance in Pedestrian Intention Prediction: A Context-Aware Review

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:59:57.055566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:59:55.663331Z digest=sha256:4b747cde78ab0346f36d581364c80b41d298a5bf8767eb11a35564460a423397

Observation ca7ff6da-26b8-4be8-b4bd-6873e3b2a77b · outbound

This paper cites Hierarchical Prompting Taxonomy: A Universal Evaluation Framework for Large Language Models Aligned with Human Cognitive Principles.

Pedestrian Intention Prediction via Vision-Language Foundation Models Hierarchical Prompting Taxonomy: A Universal Evaluation Framework for Large Language Models Aligned with Human Cognitive Principles

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T19:59:55.734120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:59:55.734120Z digest=sha256:3ab629afb599cd023e706f58c1e18e894d2a0c347b1c9d741200d2ceec6d379d

Observation fc62114f-4913-4812-9b21-2bc71f52489a · outbound

This paper cites Large Language Models Are Human-Level Prompt Engineers.

Pedestrian Intention Prediction via Vision-Language Foundation Models Large Language Models Are Human-Level Prompt Engineers

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T19:59:55.815565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:59:55.815565Z digest=sha256:e1a21ea9e872be6187f527e9a221d477e631bbc37ed32818be9c72f207ea60e9

Observation 9ceed48e-7f92-4f3b-9b9f-8c3bd66adac8 · outbound

This paper cites Benchmark for evaluating pedestrian action prediction,.

Pedestrian Intention Prediction via Vision-Language Foundation Models Benchmark for evaluating pedestrian action prediction,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:59:58.017275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:59:55.901830Z digest=sha256:adb5b18a500393800bc2a4330b2b485edb6e4f308570ba5d3453d03d7cdaccc0

Observation 99d7b5b0-b85f-469a-8584-a08fa55707f7 · outbound

This paper cites Better Zero-Shot Reasoning with Role-Play Prompting.

Pedestrian Intention Prediction via Vision-Language Foundation Models Better Zero-Shot Reasoning with Role-Play Prompting

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T19:59:55.992136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:59:55.992136Z digest=sha256:800500a5f16778d96f89d2fe4bd00d52baea1edc1f55b91b36237ff905fd5e16

Observation 699aaa9e-581e-4a0c-afd9-1752e25544cb · outbound

This paper cites Good at captioning, bad at counting: Benchmarking GPT-4V on Earth observation data.

Pedestrian Intention Prediction via Vision-Language Foundation Models Good at captioning, bad at counting: Benchmarking GPT-4V on Earth observation data

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T19:59:56.092349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:59:56.092349Z digest=sha256:326dc564f7e43393972d5ab93851bac2b61d24996b7e9653dae2f68bd437a33d

Observation 1cd8d6af-bea9-49c0-8646-1cf809825b86 · outbound

This paper cites Is the pedestrian going to cross? answering by 2d pose estimation,.

Pedestrian Intention Prediction via Vision-Language Foundation Models Is the pedestrian going to cross? answering by 2d pose estimation,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:59:57.884083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:59:56.186360Z digest=sha256:7eb5bd59c4efc52195aa77a773231317e4a35bf1b48af5b0e474fa8396413342

Observation 4e747a18-0a1f-454e-8ab8-2bf16ea658da · outbound

This paper cites Chatgpt: Generative pre-trained transformer,.

Pedestrian Intention Prediction via Vision-Language Foundation Models Chatgpt: Generative pre-trained transformer,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:59:57.715335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:59:56.286479Z digest=sha256:d0ab3155a5bdf7b6eab51033a1a5713a44d69bb3bfca8d0d572ff2b6471d5e34

Observation 35420cca-34d7-4473-94d1-ff1a8173274b · outbound

This paper cites Agreeing to cross: How drivers and pedestrians communicate,.

Pedestrian Intention Prediction via Vision-Language Foundation Models Agreeing to cross: How drivers and pedestrians communicate,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:59:57.594088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:59:56.376470Z digest=sha256:18a9de3c4498391c2f55cdb97a3a39ade32571737d3d202679b348280b2227da

Observation c5cf1eb8-ead7-4137-abbd-bfce2e73541e · outbound

This paper cites PIE: A large-scale dataset and models for pedestrian intention estimation and trajectory prediction,.

Pedestrian Intention Prediction via Vision-Language Foundation Models PIE: A large-scale dataset and models for pedestrian intention estimation and trajectory prediction,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:59:57.417787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:59:56.479906Z digest=sha256:5e7db44ea015d1cbfa4ae475e5e34d3fe86a7a539c8615cff9ea73d13c0be913

Observation 398caa99-291c-4851-bb0a-8b63ce9b5728 · outbound

This paper cites GPT-4 Technical Report.

Pedestrian Intention Prediction via Vision-Language Foundation Models GPT-4 Technical Report

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T19:59:56.585090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:59:56.585090Z digest=sha256:3df70e6c2398aab4d0e820eff38e65238cbda2aba462676b5505ba21e83258ba

Observation 768bfb91-1616-4cd3-82f5-8b553a37cf52 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Pedestrian Intention Prediction via Vision-Language Foundation Models MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T19:59:56.685176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:59:56.685176Z digest=sha256:a6162baefb8204567b32bbcc688775b90e7a58239a77a78e1dbf6207add20a36

Observation f989038e-d8a4-4f79-a1a3-12bf90ba02b3 · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

Pedestrian Intention Prediction via Vision-Language Foundation Models LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T19:59:56.791343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:59:56.791343Z digest=sha256:df81fd238f8152c5b43e32f840dc90d5564a4cbe33e7ab7d8c92bc2a6f847609

Pith citing papers

Observation 0b661ac7-40da-44c8-b177-06d0579874c9 · inbound

PEDESTRIANQA: A Benchmark for Vision-Language Models on Pedestrian Intention and Trajectory Prediction cites this paper.

PEDESTRIANQA: A Benchmark for Vision-Language Models on Pedestrian Intention and Trajectory Prediction Pedestrian Intention Prediction via Vision-Language Foundation Models

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-06-30T13:34:40.934373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T13:25:04.053283Z digest=sha256:d893697eb9f1f7ddc072b97d66e4c1e6eed4a561e3fada1ddbee6a8f169683ac