Pith. sign in

Paper Citation Record · LEDGER

HRIBench: Benchmarking Vision-Language Models for Real-Time Human Perception in Human-Robot Interaction

As of 8 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 2 inbound Pith citation observations for arXiv:2506.20566.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.20566 v1

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:49:10.881232Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T08:43:33.391528Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact0
  • verified fuzzy18
  • unresolved8
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1a90adff-d994-4835-914a-c1ef00a2389d · outbound

This paper cites Evaluating Vision-Language Models as Evaluators in Path Planning.

HRIBench: Benchmarking Vision-Language Models for Real-Time Human Perception in Human-Robot Interaction Evaluating Vision-Language Models as Evaluators in Path Planning

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T22:49:11.200446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:49:07.509682Z digest=sha256:97bb626bb5023e2824e2d2315815fc17ae0ec391f550cefe5959c91f18efd0b1

Observation 962eeeb6-509b-491b-a217-70702f5a12f5 · outbound

This paper cites In: Proceedings of the 19th ACM International Conference on Mul- timodal Interaction, pp.

HRIBench: Benchmarking Vision-Language Models for Real-Time Human Perception in Human-Robot Interaction In: Proceedings of the 19th ACM International Conference on Mul- timodal Interaction, pp

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:17.057335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:49:07.590463Z digest=sha256:179efd7e8bdf201dc2d6eb1f5acff8039422f854a7c4ea73ce2196f4a44c5415

Observation 71e0cde9-4a6e-48b7-bf08-c46222e1b4e5 · outbound

This paper cites Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets.

HRIBench: Benchmarking Vision-Language Models for Real-Time Human Perception in Human-Robot Interaction Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:07.721472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:07.721472Z digest=sha256:f1bd6e979a2633caea6ab75f617a4dcb682b4a50550a74991166a55fd8063a31

Observation 6138fe75-e541-450c-a7b0-2affee204481 · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp.

HRIBench: Benchmarking Vision-Language Models for Real-Time Human Perception in Human-Robot Interaction In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:16.795102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:49:07.813078Z digest=sha256:6780d58b80e35820659c43eb8254c9b56521ee6aaba27fe018ec660e1285b2e6

Observation a50c31f4-58d4-4301-bb48-80366c6c1157 · outbound

This paper cites In: International Conference on Learning Representations (ICLR) (2025).

HRIBench: Benchmarking Vision-Language Models for Real-Time Human Perception in Human-Robot Interaction In: International Conference on Learning Representations (ICLR) (2025)

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:16.485703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:49:07.915292Z digest=sha256:8dcd848c5069c927fa0971fd84bf2f271c5847e10f500a21c8d26351788e5bc6

Observation 500ca7e8-8220-4aff-b648-776631922f9f · outbound

This paper cites In: 2008 8th Ieee International Conference on Automatic Face & Gesture Recognition, pp.

HRIBench: Benchmarking Vision-Language Models for Real-Time Human Perception in Human-Robot Interaction In: 2008 8th Ieee International Conference on Automatic Face & Gesture Recognition, pp

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:16.193859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:49:08.077537Z digest=sha256:474b775cad3118726fe6f21aa11a63c854e6e03ea8ea2cc70a5b720b4aaab871

Observation 43fee879-e3ea-4fe5-9578-b03d930b3efa · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

HRIBench: Benchmarking Vision-Language Models for Real-Time Human Perception in Human-Robot Interaction Gemini: A Family of Highly Capable Multimodal Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:08.200579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:08.200579Z digest=sha256:dad62ec569eb12ba4e17140864fa095ec36c95f4aae2c25bec578d5b576f3e93

Observation 04f16770-e207-4dac-b03e-0b76b6c81956 · outbound

This paper cites The Llama 3 Herd of Models.

HRIBench: Benchmarking Vision-Language Models for Real-Time Human Perception in Human-Robot Interaction The Llama 3 Herd of Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:08.371492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:08.371492Z digest=sha256:adb66dd52196bcbe90b726bbc71a1d9ab149cb6c3dfdbf71047f4260351fa107

Observation 7cfc5b71-6772-4b10-abd6-4f8f340cb62e · outbound

This paper cites In: Neural Information Processing Systems (NeurIPS) (2018).

HRIBench: Benchmarking Vision-Language Models for Real-Time Human Perception in Human-Robot Interaction In: Neural Information Processing Systems (NeurIPS) (2018)

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:15.866164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:49:08.475881Z digest=sha256:a0f6e0a0f663a6f4a09b1f3304de42653a21c857817af07fd15fb97daa428a78

Observation 84759529-e6ae-4686-a506-72c9db046e09 · outbound

This paper cites Benchmarking Vision, Language, & Action Models on Robotic Learning Tasks.

HRIBench: Benchmarking Vision-Language Models for Real-Time Human Perception in Human-Robot Interaction Benchmarking Vision, Language, & Action Models on Robotic Learning Tasks

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:08.555776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:08.555776Z digest=sha256:564efccd4816d182c9c00508468359523399ff081a8a6bed57de46a138aac7ef

Observation eaf7af65-57fd-49c9-b145-d09867dd3e2f · outbound

This paper cites Human factors53(5), 517–527 (2011).

HRIBench: Benchmarking Vision-Language Models for Real-Time Human Perception in Human-Robot Interaction Human factors53(5), 517–527 (2011)

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:15.496895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:49:08.658326Z digest=sha256:9ecb25bd83c4567696d18528b28e625331ad50702bf17e72548ba5182240925c

Observation 59c4f372-2edd-4317-aaea-6ba8c91afd8c · outbound

This paper cites $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization.

HRIBench: Benchmarking Vision-Language Models for Real-Time Human Perception in Human-Robot Interaction $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:08.758255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:08.758255Z digest=sha256:fbf831a4bd6b715b1c46f19d34fdcf053426e247db5dbbb601fef914c688f077

Observation 3440f35e-8fb1-4ba9-9e23-b74c19f90f4d · outbound

This paper cites In: Conference on Robot Learning (2023).

HRIBench: Benchmarking Vision-Language Models for Real-Time Human Perception in Human-Robot Interaction In: Conference on Robot Learning (2023)

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:15.105483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:49:08.876306Z digest=sha256:e12391e4a8a043d42b0706b3ccdbce420bc4ef5c28837b242b3dea25ca6dc74f

Observation 85b25a58-aa63-465c-9dac-fc4084039e6b · outbound

This paper cites Discourse Processes52(4), 255–289 (2015).

HRIBench: Benchmarking Vision-Language Models for Real-Time Human Perception in Human-Robot Interaction Discourse Processes52(4), 255–289 (2015)

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:14.752505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:49:08.993107Z digest=sha256:14abf621c65a3dadb658e3a1b8f3af8da79c7b153a3e7a284eca91b61511edf9

Observation 2a620a9f-dda5-4e84-9be3-0afa2b7e443a · outbound

This paper cites A Survey of State of the Art Large Vision Language Models: Alignment, Benchmark, Evaluations and Challenges.

HRIBench: Benchmarking Vision-Language Models for Real-Time Human Perception in Human-Robot Interaction A Survey of State of the Art Large Vision Language Models: Alignment, Benchmark, Evaluations and Challenges

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:09.088375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:09.088375Z digest=sha256:d2a28e9d09f38f0ddab5d14d65214213701bcb6786c2fca29300b789f02640b4

Observation f7deb7e5-620d-4d07-afaa-baa9253c1041 · outbound

This paper cites In: Robotics: Science and Systems (RSS) (2024).

HRIBench: Benchmarking Vision-Language Models for Real-Time Human Perception in Human-Robot Interaction In: Robotics: Science and Systems (RSS) (2024)

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:14.371159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:49:09.199127Z digest=sha256:1e2a8215ae7eb5ac278074b511135ead9c439de840a8ea89fd1069bef36bba7c

Observation 771afcae-c807-47e1-bc8d-dfd15e5b7eca · outbound

This paper cites ACM Transactions on Human-Robot Interaction12(3), 1–39 (2023).

HRIBench: Benchmarking Vision-Language Models for Real-Time Human Perception in Human-Robot Interaction ACM Transactions on Human-Robot Interaction12(3), 1–39 (2023)

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:14.066289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:49:09.350710Z digest=sha256:3682e482e64bc8b44bddfbff6938132a68cf25a0a6fc3e1ebc19af352f680f39

Observation f17af41b-8e60-4a27-bc26-6d2157040c84 · outbound

This paper cites https://robotsguide.com/robots/kuri (n.d.).

HRIBench: Benchmarking Vision-Language Models for Real-Time Human Perception in Human-Robot Interaction https://robotsguide.com/robots/kuri (n.d.)

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:13.769324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:49:09.505490Z digest=sha256:10c59f45b3f1a659837301cdfefb9ef797bd31d2099d7ea1de655d83ce6f5b70

Observation c8b0254b-7d9e-4bb3-8206-9beb73ad23f3 · outbound

This paper cites In: 2023 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pp.

HRIBench: Benchmarking Vision-Language Models for Real-Time Human Perception in Human-Robot Interaction In: 2023 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pp

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:13.346604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:49:09.666785Z digest=sha256:5c3a76531d2cd0122d61af8669656ce97fe7eee48477ec23335a936b14c1cf68

Observation 7fc6ee28-5da2-48ca-b3e4-943eb4eaed6d · outbound

This paper cites Image and Vision Computing25(12), 1875–1884 (2007).

HRIBench: Benchmarking Vision-Language Models for Real-Time Human Perception in Human-Robot Interaction Image and Vision Computing25(12), 1875–1884 (2007)

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:12.972703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:49:09.869220Z digest=sha256:f53bc5386d534df1d1ffc4b9c08d869b22f5b1ba4a3dcededd8a6098611db8d2

Observation da941580-0f76-4fb9-a786-311d074a0bac · outbound

This paper cites GPT-4o System Card.

HRIBench: Benchmarking Vision-Language Models for Real-Time Human Perception in Human-Robot Interaction GPT-4o System Card

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:10.058113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:10.058113Z digest=sha256:9915a6ea629ea400a46d2616608d059a758e037f5346b7968413e23f9d5250f3

Observation e9872746-cccb-424c-8a30-bae22d03fa1b · outbound

This paper cites ACM Transactions on Human-Robot Interaction12(1), 1–66 (2023).

HRIBench: Benchmarking Vision-Language Models for Real-Time Human Perception in Human-Robot Interaction ACM Transactions on Human-Robot Interaction12(1), 1–66 (2023)

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:12.673737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:49:10.170501Z digest=sha256:aec972fef8e7be754b00dce9249f8f55ba8a871e318aa19a885aadf203aa2331

Observation ac043cf2-4523-4df1-a8a3-7a424b53de3f · outbound

This paper cites IEEE Transactions on Robotics38(3), 1755–1772 (2021).

HRIBench: Benchmarking Vision-Language Models for Real-Time Human Perception in Human-Robot Interaction IEEE Transactions on Robotics38(3), 1755–1772 (2021)

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:12.374698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:49:10.326554Z digest=sha256:770067b1dbeef619e313cc3fd6a1d8f2da55034a0078984eb75d43e2d5736512

Observation d24db9c4-7666-4fdf-9c03-eba2412b1b6d · outbound

This paper cites ACM Transactions on Human-Robot Interaction 12(2), 1–21 (2023).

HRIBench: Benchmarking Vision-Language Models for Real-Time Human Perception in Human-Robot Interaction ACM Transactions on Human-Robot Interaction 12(2), 1–21 (2023)

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:12.095926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:49:10.477606Z digest=sha256:a6c10a44565ee6e457da58496b0d7234b625b54c934420192b508fd670415c86

Observation e218a7f9-66ac-44ac-9c42-c00fa81747a2 · outbound

This paper cites Advances in Neural Information Processing Systems35, 12,014–12,026 (2022) 10 Zhonghao Shi et al.

HRIBench: Benchmarking Vision-Language Models for Real-Time Human Perception in Human-Robot Interaction Advances in Neural Information Processing Systems35, 12,014–12,026 (2022) 10 Zhonghao Shi et al

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:11.776077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:49:10.621525Z digest=sha256:5831183e4b7be9656063d2777873f1ee294dda5df757a312daa31414be500e70

Observation f2ad57d4-4681-4293-8b70-f9451405badc · outbound

This paper cites ManipBench: Benchmarking Vision-Language Models for Low-Level Robot Manipulation.

HRIBench: Benchmarking Vision-Language Models for Real-Time Human Perception in Human-Robot Interaction ManipBench: Benchmarking Vision-Language Models for Low-Level Robot Manipulation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:10.769599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:10.769599Z digest=sha256:de5c3f4df11408607b53edd1c87d6aaa8157c7a1752ca8362317cafa3566ab5e

Observation d6210fe0-1cc8-4a6a-9776-53739c97e593 · outbound

This paper cites In: Artificial In- telligence and Machine Learning in Defense Applications II (2020).

HRIBench: Benchmarking Vision-Language Models for Real-Time Human Perception in Human-Robot Interaction In: Artificial In- telligence and Machine Learning in Defense Applications II (2020)

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:11.460335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:49:10.881232Z digest=sha256:95984f395894024e5fecc8d710cb07bed7fced444a8606c095922d309a07e509

Pith citing papers

Observation e94ac13b-5b75-40fb-92fe-c1c5ee6012dc · inbound

ERR@HRI 3.0 Challenge: Multimodal Detection of Errors and Anticipation in Human-Robot Interactions cites this paper.

ERR@HRI 3.0 Challenge: Multimodal Detection of Errors and Anticipation in Human-Robot Interactions HRIBench: Benchmarking Vision-Language Models for Real-Time Human Perception in Human-Robot Interaction

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-14T04:42:38.419346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:42:38.419346Z digest=sha256:7ea130af5e29da774632c88a622e26113a2b39e3dcb4de2d0e58a4b1f348634f

Observation f3cf56be-7b34-4131-812e-7362bbde9172 · inbound

HRIBench: Benchmarking Interaction-Centric Human-Robot Collaboration cites this paper.

HRIBench: Benchmarking Interaction-Centric Human-Robot Collaboration HRIBench: Benchmarking Vision-Language Models for Real-Time Human Perception in Human-Robot Interaction

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T08:43:33.391528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:43:33.391528Z digest=sha256:3304f1f8cd384880b08dd5de84df88959adb6f03f98d5f2bf8f6c441f9c3fbc4