Pith. sign in

Paper Citation Record · LEDGER

CrowdVLA: Embodied Vision-Language-Action Agents for Context-Aware Crowd Simulation

As of 5 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 1 inbound Pith citation observation for arXiv:2604.05525.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.05525 v1

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T18:53:56.269363Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-31T00:25:50.472392Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

26 of 26 outbound references displayed

  • verified exact2
  • verified fuzzy15
  • unresolved2
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch7

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8b0758d0-77d9-4530-90e8-ac90a88f73f2 · outbound

This paper cites Qwen2.5-VL Technical Report.

CrowdVLA: Embodied Vision-Language-Action Agents for Context-Aware Crowd Simulation Qwen2.5-VL Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-10T23:45:52.148362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:53:56.269363Z digest=sha256:0a499ab1d1a5f02334a4c5f2b2b3d41f1010e9363c6a8f8a4483f8ad08c85f0f

Observation 1bf56a5a-cba9-4fef-b2a1-a1b1477f5677 · outbound

This paper cites Peng Chen, Pi Bu, Yingyao Wang, Xinyi Wang, Ziming Wang, Jie Guo, Yingxiu Zhao, Qi Zhu, Jun Song, Siran Yang, et al.

CrowdVLA: Embodied Vision-Language-Action Agents for Context-Aware Crowd Simulation Peng Chen, Pi Bu, Yingyao Wang, Xinyi Wang, Ziming Wang, Jie Guo, Yingxiu Zhao, Qi Zhu, Jun Song, Siran Yang, et al

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:12:55.261995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:53:56.269363Z digest=sha256:19632deeb6f972a8ffa7ec29bf96353ee2c4ef225261e24b143fedd7fdc2bd0e

Observation 9420eae4-948e-4aa0-8a6e-01b82769fb84 · outbound

This paper cites Combatvla: An efficient vision-language-action model for combat tasks in 3d action role-playing games.arXiv preprint arXiv:2503.09527, 2025a.

CrowdVLA: Embodied Vision-Language-Action Agents for Context-Aware Crowd Simulation Combatvla: An efficient vision-language-action model for combat tasks in 3d action role-playing games.arXiv preprint arXiv:2503.09527, 2025a

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:45:52.072465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:53:56.269363Z digest=sha256:853d84e17fc075392b860cc1bf8863094580024c6f9eb0d081bfd500c0986587

Observation 8e9868df-4e1e-4bbc-8d25-93d2f8af14a0 · outbound

This paper cites VADv2: End-to-End Vectorized Autonomous Driving via Probabilistic Planning.

CrowdVLA: Embodied Vision-Language-Action Agents for Context-Aware Crowd Simulation VADv2: End-to-End Vectorized Autonomous Driving via Probabilistic Planning

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-05-10T23:45:52.097393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:53:56.269363Z digest=sha256:71c153c454ad732b18deb9921cb9fab42109a0d52fe7a9732d874982cab17ae9

Observation 133ce87f-d7f4-45a8-9116-e16f998d9133 · outbound

This paper cites InProceedings of the 2009 ACM SIGGRAPH/Eurographics Symposium on Computer Animation.

CrowdVLA: Embodied Vision-Language-Action Agents for Context-Aware Crowd Simulation InProceedings of the 2009 ACM SIGGRAPH/Eurographics Symposium on Computer Animation

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:12:55.272271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:53:56.269363Z digest=sha256:9084369e78dcd57de6ec81dcc2f0624c3e08a591f4f99facd88ed0ce29f44b1c

Observation 78636f50-b63e-47fd-90d9-1f63cc94990b · outbound

This paper cites InProceedings of the 2011 ACM SIGGRAPH/Eurographics symposium on computer animation.

CrowdVLA: Embodied Vision-Language-Action Agents for Context-Aware Crowd Simulation InProceedings of the 2011 ACM SIGGRAPH/Eurographics symposium on computer animation

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:12:55.264908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:53:56.269363Z digest=sha256:d9a73b4e6ba2df03f160fa6c2a2a2a2082eda613968017261b86a230579a6eee

Observation 753d8ad8-2021-441b-9487-eed1130b110c · outbound

This paper cites Physical review E51, 5 (1995).

CrowdVLA: Embodied Vision-Language-Action Agents for Context-Aware Crowd Simulation Physical review E51, 5 (1995)

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:12:55.270036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:53:56.269363Z digest=sha256:7e52c16c02957a832167de559efc2b680a871e3855e5e9743cd16b78591bb319

Observation f9f526bd-ae45-494f-8404-7959e8bda668 · outbound

This paper cites an unresolved cited work.

CrowdVLA: Embodied Vision-Language-Action Agents for Context-Aware Crowd Simulation Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-05-16T13:12:55.267530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:53:56.269363Z digest=sha256:0f4fe840215ee6774c4edb834792fd66df12af42b40fc1c5779d4bdca9fdda21

Observation 0ab72f39-2341-413e-84a7-bc66e9061ba5 · outbound

This paper cites Xuebo Ji, Zherong Pan, Xifeng Gao, and Jia Pan.

CrowdVLA: Embodied Vision-Language-Action Agents for Context-Aware Crowd Simulation Xuebo Ji, Zherong Pan, Xifeng Gao, and Jia Pan

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:12:55.259618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:53:56.269363Z digest=sha256:6784e6ee8b18efc382faf77cc3d358a14ac52908da7b5a22c1c5c8ad5ee933dd

Observation 09fa6925-eb75-4f87-89fd-3257ac945c6b · outbound

This paper cites InACM SIGGRAPH 2024 Conference Papers.

CrowdVLA: Embodied Vision-Language-Action Agents for Context-Aware Crowd Simulation InACM SIGGRAPH 2024 Conference Papers

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:12:55.226490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:53:56.269363Z digest=sha256:ebb730b4e9d9acac20cd17b6bcd7f4978dbe8cb4c80244b862920d7d7a77a715

Observation f908a077-21db-4516-a107-6ea65820188c · outbound

This paper cites Mubbasir Kapadia, Alejandro Beacco, Francisco Garcia, Vivek Reddy, Nuria Pelechano, and Norman I Badler.

CrowdVLA: Embodied Vision-Language-Action Agents for Context-Aware Crowd Simulation Mubbasir Kapadia, Alejandro Beacco, Francisco Garcia, Vivek Reddy, Nuria Pelechano, and Norman I Badler

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:12:55.251075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:53:56.269363Z digest=sha256:ae2627997556b7d705e67f47789a65bccc4cad6473e53f4a6ed368564539df05

Observation ddcac3d9-d801-47c3-90c3-314532b7988a · outbound

This paper cites Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success.

CrowdVLA: Embodied Vision-Language-Action Agents for Context-Aware Crowd Simulation Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:35:33.294794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:53:56.269363Z digest=sha256:3ec0ca9cc8384e4db78faf7c4df42f536cbb3c78b744af358f05d45abe8c6c55

Observation 82d1df69-a857-4da9-9a3d-3a8387320ad2 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

CrowdVLA: Embodied Vision-Language-Action Agents for Context-Aware Crowd Simulation OpenVLA: An Open-Source Vision-Language-Action Model

Reference 13

Resolution
metadata mismatch
local_arxiv, observed 2026-05-10T23:45:52.042498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:53:56.269363Z digest=sha256:12f23a65058c84d4c36f2e6c59b567152f34ff45f86c2cc3e0ecba6fa3216617

Observation e7c655bf-f00a-45ba-9401-7a06f4526cc2 · outbound

This paper cites InProceedings of the 2007 ACM SIGGRAPH/Eurographics symposium on Computer animation.

CrowdVLA: Embodied Vision-Language-Action Agents for Context-Aware Crowd Simulation InProceedings of the 2007 ACM SIGGRAPH/Eurographics symposium on Computer animation

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:12:55.248223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:53:56.269363Z digest=sha256:ce411084fb7d47d8567986f8473e4c9d072eb246fc20c7a149c7dc41c1f77b1a

Observation 1f5f9000-ea78-43b2-b4fc-55436cda8c1b · outbound

This paper cites End-to-End Driving with Online Trajectory Evaluation via BEV World Model.

CrowdVLA: Embodied Vision-Language-Action Agents for Context-Aware Crowd Simulation End-to-End Driving with Online Trajectory Evaluation via BEV World Model

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:45:52.131692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:53:56.269363Z digest=sha256:361ec855c8cd2131cfae62d1bce01c89176e27314e3a17e7fe4e0a86bcbe7045

Observation 77444caf-1518-46da-a822-559696a857fe · outbound

This paper cites Hydra-MDP: End-to-end Multimodal Planning with Multi-target Hydra-Distillation.

CrowdVLA: Embodied Vision-Language-Action Agents for Context-Aware Crowd Simulation Hydra-MDP: End-to-end Multimodal Planning with Multi-target Hydra-Distillation

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T23:13:55.802367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:53:56.269363Z digest=sha256:e2a77719b45869a7dcaa626c39c249c0735101896f1596e93fb8ec473631040f

Observation 2f7d31c6-b764-4571-83dd-dd4b43937f37 · outbound

This paper cites Andreas Panayiotou, Theodoros Kyriakou, Marilena Lemonari, Yiorgos Chrysanthou, and Panayiotis Charalambous.

CrowdVLA: Embodied Vision-Language-Action Agents for Context-Aware Crowd Simulation Andreas Panayiotou, Theodoros Kyriakou, Marilena Lemonari, Yiorgos Chrysanthou, and Panayiotis Charalambous

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:12:55.256608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:53:56.269363Z digest=sha256:d865ca440630c26b8e91de4b4c8f25855f5e3bce8ce06551886ea52c27ab967d

Observation 0df1422b-7454-41ed-91f9-82b730160382 · outbound

This paper cites InACM SIGGRAPH 2022 conference proceedings.

CrowdVLA: Embodied Vision-Language-Action Agents for Context-Aware Crowd Simulation InACM SIGGRAPH 2022 conference proceedings

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:12:55.242176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:53:56.269363Z digest=sha256:2433faaaef441e0b5e4c3e4c6c6c2bcf223f3cd55c81a4d5895a6c0acc478f4f

Observation 02b3638e-bf40-4c09-8a81-084307a1f989 · outbound

This paper cites Alexandre Robicquet, Amir Sadeghian, Alexandre Alahi, and Silvio Savarese.

CrowdVLA: Embodied Vision-Language-Action Agents for Context-Aware Crowd Simulation Alexandre Robicquet, Amir Sadeghian, Alexandre Alahi, and Silvio Savarese

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:12:55.239129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:53:56.269363Z digest=sha256:c90326170ecab9868a4032148ddfd102485f71d6ca0fb5f76746165b7c527144

Observation c645131f-79e6-44b3-a2fc-3a13a0221f20 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

CrowdVLA: Embodied Vision-Language-Action Agents for Context-Aware Crowd Simulation Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-10T23:45:52.139304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:53:56.269363Z digest=sha256:e4ce1d63557a9d372a36f61b2eb10e48f5d435eb05d269089cfa22b7cd18fd0a

Observation 94226142-bfce-48ee-ab90-2fe6d95c5f45 · outbound

This paper cites In2008 IEEE international conference on robotics and automation.

CrowdVLA: Embodied Vision-Language-Action Agents for Context-Aware Crowd Simulation In2008 IEEE international conference on robotics and automation

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:12:55.245023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:53:56.269363Z digest=sha256:af47221347b4437a38d018a014d4711f43c045a72c22d35675b7c0fb953a2aad

Observation 3c2dcaac-b23e-4d48-98c5-32df28d06a3c · outbound

This paper cites an unresolved cited work.

CrowdVLA: Embodied Vision-Language-Action Agents for Context-Aware Crowd Simulation Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-05-16T13:12:55.229610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:53:56.269363Z digest=sha256:3ba9743cf805e126983678c3e2456adb7643f046f8285756052f92c3d5554823

Observation 520c66c2-aa88-437f-9eb9-7774378805fa · outbound

This paper cites Zewei Zhou, Tianhui Cai, Seth Z Zhao, Yun Zhang, Zhiyu Huang, Bolei Zhou, and Jiaqi Ma.

CrowdVLA: Embodied Vision-Language-Action Agents for Context-Aware Crowd Simulation Zewei Zhou, Tianhui Cai, Seth Z Zhao, Yun Zhang, Zhiyu Huang, Bolei Zhou, and Jiaqi Ma

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:12:55.253747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:53:56.269363Z digest=sha256:de6355c47cc43aac76d139f04db1a6bc14c58983d85b60606e84d3cbf10b3e1c

Observation a00a75dc-a1a0-49ad-ac97-598749d5c958 · outbound

This paper cites AutoVLA: A Vision-Language-Action Model for End-to-End Autonomous Driving with Adaptive Reasoning and Reinforcement Fine-Tuning.

CrowdVLA: Embodied Vision-Language-Action Agents for Context-Aware Crowd Simulation AutoVLA: A Vision-Language-Action Model for End-to-End Autonomous Driving with Adaptive Reasoning and Reinforcement Fine-Tuning

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T21:46:44.401546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:53:56.269363Z digest=sha256:73480b5a3cfd21a8e6dcd2fc8a7aa53485dd9bd2bd516da8bfc1dd67034de63c

Observation 438946f1-7e29-4430-88d4-187b56ee472c · outbound

This paper cites We compare pedestrian trajectories when nine agents cross in each scene.

CrowdVLA: Embodied Vision-Language-Action Agents for Context-Aware Crowd Simulation We compare pedestrian trajectories when nine agents cross in each scene

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:12:55.232857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:53:56.269363Z digest=sha256:539853037f50ea4e91da347dcfcdf3ea3a398dfd6e87c094b61764fab9726eea

Observation f06f05ef-49c5-4ac7-b0d4-51aed9b60bba · outbound

This paper cites Each column shows the agent’s trajectory (top) and corresponding third-person observations at selected timesteps (bottom).

CrowdVLA: Embodied Vision-Language-Action Agents for Context-Aware Crowd Simulation Each column shows the agent’s trajectory (top) and corresponding third-person observations at selected timesteps (bottom)

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:12:55.236363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:53:56.269363Z digest=sha256:eaa592ef1f91ea8acb8c346b4271eab2decbb8f7efc963466aa54ac08140a457

Pith citing papers

Observation 5919e810-075f-4faf-8fa4-4dfe790f7e52 · inbound

How Affect Propagates among LLM Agents: Emergent Emotional Contagion in Crowd Simulation cites this paper.

How Affect Propagates among LLM Agents: Emergent Emotional Contagion in Crowd Simulation CrowdVLA: Embodied Vision-Language-Action Agents for Context-Aware Crowd Simulation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-31T00:25:50.472392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:25:50.472392Z digest=sha256:cac322511d0fc291904c279a0a3a4789f0ae296545475910725064a609d52580