Pith. sign in

Paper Citation Record · LEDGER

Assessing the Value of Visual Input: A Benchmark of Multimodal Large Language Models for Robotic Path Planning

As of 7 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 1 inbound Pith citation observation for arXiv:2507.12391.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.12391 v1

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T16:50:09.623638Z

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T11:40:21.327031Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T11:40:25.831314Z

Reference resolution

23 of 23 outbound references displayed

  • verified exact3
  • verified fuzzy10
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 08cbf44a-4b00-4086-b09c-742698c10219 · outbound

This paper cites A survey on large language model based autonomous agents.

Assessing the Value of Visual Input: A Benchmark of Multimodal Large Language Models for Robotic Path Planning A survey on large language model based autonomous agents

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:50:12.672882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:50:07.398980Z digest=sha256:a3153faa0eccd4b032e84c23bb9d7aef8b98064bfb5fdb2feabecd48fbfdceab

Observation 9986f59e-4f4c-475a-965f-7ee8a08b25c0 · outbound

This paper cites Large language models for robotics: Opportunities, challenges, and perspectives.

Assessing the Value of Visual Input: A Benchmark of Multimodal Large Language Models for Robotic Path Planning Large language models for robotics: Opportunities, challenges, and perspectives

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:50:12.514886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:50:07.474604Z digest=sha256:66243761bbf0abb9848f8f51387eb81e5a184b19654bd632dcbbc47e654df242

Observation c65b190b-166d-4fcc-a5f3-a3722221d776 · outbound

This paper cites Human-Robot collaboration in surgery: Advances and challenges towards autonomous surgical assis- tants.

Assessing the Value of Visual Input: A Benchmark of Multimodal Large Language Models for Robotic Path Planning Human-Robot collaboration in surgery: Advances and challenges towards autonomous surgical assis- tants

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:50:12.354229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:50:07.565616Z digest=sha256:3781902c22a1df1b51a76e2cf68153a9a3a5cd9a1c6673fc2bb9113fbb3a59a2

Observation 57623a7e-2b68-4061-abe0-f2ee19db3928 · outbound

This paper cites Language models as zero-shot planners: Extracting actionable knowledge for embodied agents.

Assessing the Value of Visual Input: A Benchmark of Multimodal Large Language Models for Robotic Path Planning Language models as zero-shot planners: Extracting actionable knowledge for embodied agents

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:50:12.147985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:50:07.700942Z digest=sha256:0f5834f46bb248904b9a33d744d9e195dd16d8a479503b2e8c8703860d8e09b2

Observation cbc1c722-8cdf-4b41-b90a-9a502e50e7a8 · outbound

This paper cites Do As I Can, Not As I Say: Grounding Language in Robotic Affordances.

Assessing the Value of Visual Input: A Benchmark of Multimodal Large Language Models for Robotic Path Planning Do As I Can, Not As I Say: Grounding Language in Robotic Affordances

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:50:11.952899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:50:07.858503Z digest=sha256:0d64a415e1fd60f6acee56334ca851ffde46997fa1be7175dc7553a53c7fa3e6

Observation 3f2e3fba-f0c5-4623-88d5-2bbb2c9d94e3 · outbound

This paper cites Progprompt: Generating situated robot task plans using large language models.

Assessing the Value of Visual Input: A Benchmark of Multimodal Large Language Models for Robotic Path Planning Progprompt: Generating situated robot task plans using large language models

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:50:11.717772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:50:07.985416Z digest=sha256:36edf74cec64a49884c1eb8f60a06b177e133f3b29f5d4b0731865bbc2c49234

Observation 19a8e19c-e42c-4bdf-ab88-6750cce49d7a · outbound

This paper cites V oice con- trol interface for surgical robot assistants.

Assessing the Value of Visual Input: A Benchmark of Multimodal Large Language Models for Robotic Path Planning V oice con- trol interface for surgical robot assistants

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:50:11.412453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:50:08.092121Z digest=sha256:223c3a158aca588bfc769172683b191f27488d16609a1f2759b4b5c1927add7d

Observation a94424e7-7caa-4cee-a8cf-f67ecc0d3653 · outbound

This paper cites LLM-based ambiguity detection in natural language instructions for collaborative surgical robots.

Assessing the Value of Visual Input: A Benchmark of Multimodal Large Language Models for Robotic Path Planning LLM-based ambiguity detection in natural language instructions for collaborative surgical robots

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:50:11.129833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:50:08.202134Z digest=sha256:629f4553088e5bd09e612db1ad1510b726ddf82fa59a5709dd883c438ad68a90

Observation 1a42df65-51b0-4524-ba32-23ca0dd7a425 · outbound

This paper cites Toward autonomous robotic minimally invasive surgery: A hybrid framework combining task-motion planning and dynamic behavior trees.

Assessing the Value of Visual Input: A Benchmark of Multimodal Large Language Models for Robotic Path Planning Toward autonomous robotic minimally invasive surgery: A hybrid framework combining task-motion planning and dynamic behavior trees

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:50:10.835604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:50:08.392134Z digest=sha256:29db2aa69ccb09124ff15a17b701cebce7a9c85c839ce5ce61d8495cb5b763a2

Observation 4a39828d-d7e5-453d-88f5-03b56020e6eb · outbound

This paper cites Exploring embodied mul- timodal large models: Development, datasets, and future directions.

Assessing the Value of Visual Input: A Benchmark of Multimodal Large Language Models for Robotic Path Planning Exploring embodied mul- timodal large models: Development, datasets, and future directions

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:50:10.620752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:50:08.518679Z digest=sha256:dc24aa0e6e75b3ff069e1f9fdfa1c784dfcb1256dc1a3728d4cc58435b8cf855

Observation 47faf2df-fb07-430d-ad0a-1957377fb85c · outbound

This paper cites Multimodal Fusion and Vision-Language Models: A Survey for Robot Vision.

Assessing the Value of Visual Input: A Benchmark of Multimodal Large Language Models for Robotic Path Planning Multimodal Fusion and Vision-Language Models: A Survey for Robot Vision

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:08.597731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:08.597731Z digest=sha256:0ef48f2b079947b336257c6c1af0b84bde497fae2b73e0fb59e5d242e5d9ea6c

Observation 16c1735e-0042-4033-abad-8427410ddd74 · outbound

This paper cites PaLM-E: An Embodied Multimodal Language Model.

Assessing the Value of Visual Input: A Benchmark of Multimodal Large Language Models for Robotic Path Planning PaLM-E: An Embodied Multimodal Language Model

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:08.663316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:08.663316Z digest=sha256:191e0ea9313d997f339db01d9a24987215d2a813521217e6dbdc761fce6446b9

Observation 07c40737-2ab5-4210-8b4e-34d859072622 · outbound

This paper cites RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control.

Assessing the Value of Visual Input: A Benchmark of Multimodal Large Language Models for Robotic Path Planning RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:08.750449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:08.750449Z digest=sha256:2b8e6afa795b0f119989a7f8da95dad5ebc17ab9f3481a74f512ab16553670af

Observation 3c1d1cae-4389-435f-a7d6-0e4ba6a3e83e · outbound

This paper cites VIMA: General Robot Manipulation with Multimodal Prompts.

Assessing the Value of Visual Input: A Benchmark of Multimodal Large Language Models for Robotic Path Planning VIMA: General Robot Manipulation with Multimodal Prompts

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:08.815246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:08.815246Z digest=sha256:67aab607b60516be74d91aea425445299b60f3dfc6dbfc1f645b18af2148eb3c

Observation f1d0268a-a77f-4a04-9f93-2964c0b7155d · outbound

This paper cites LLM-A*: Large Language Model Enhanced Incremental Heuristic Search on Path Planning.

Assessing the Value of Visual Input: A Benchmark of Multimodal Large Language Models for Robotic Path Planning LLM-A*: Large Language Model Enhanced Incremental Heuristic Search on Path Planning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:08.926716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:08.926716Z digest=sha256:cf6ac8496c728f5f1f5dea3acaa390d844836c14881859df72059ceaea0b9bce

Observation 3ccd6026-280f-407f-892c-740167a13455 · outbound

This paper cites LLM-Advisor: An LLM Benchmark for Cost-efficient Path Planning across Multiple Terrains.

Assessing the Value of Visual Input: A Benchmark of Multimodal Large Language Models for Robotic Path Planning LLM-Advisor: An LLM Benchmark for Cost-efficient Path Planning across Multiple Terrains

Reference 16

Resolution
verified exact
raw_fallback, observed 2026-08-06T16:50:10.352632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:50:09.019979Z digest=sha256:045befbe73a8a38710456532f837a19839dfc68465304892ca7df095adfaea65

Observation 622f095f-1e64-4971-a277-1bd0feddc52a · outbound

This paper cites LLM-Enhanced Path Planning: Safe and Efficient Autonomous Navigation with Instructional Inputs.

Assessing the Value of Visual Input: A Benchmark of Multimodal Large Language Models for Robotic Path Planning LLM-Enhanced Path Planning: Safe and Efficient Autonomous Navigation with Instructional Inputs

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:09.102714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:09.102714Z digest=sha256:e01d3e3b7fe7ec729f625d17a74a391d8b6b9175d5ee394fbd1d56f54782c538

Observation 235df9a0-fdba-4ab2-a74f-909b7a894a4e · outbound

This paper cites Guide-LLM: An Embodied LLM Agent and Text-Based Topological Map for Robotic Guidance of People with Visual Impairments.

Assessing the Value of Visual Input: A Benchmark of Multimodal Large Language Models for Robotic Path Planning Guide-LLM: An Embodied LLM Agent and Text-Based Topological Map for Robotic Guidance of People with Visual Impairments

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-06T16:50:10.109881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:50:09.172796Z digest=sha256:2bbf74a06e309f955ecdcde5005af54efe451fec4e84524e111831826ea4000e

Observation acc50030-12b0-4839-a227-482856a712ca · outbound

This paper cites From Text to Space: Mapping Abstract Spatial Models in LLMs during a Grid-World Navigation Task.

Assessing the Value of Visual Input: A Benchmark of Multimodal Large Language Models for Robotic Path Planning From Text to Space: Mapping Abstract Spatial Models in LLMs during a Grid-World Navigation Task

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:09.264012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:09.264012Z digest=sha256:4e5b9cf1f98bcbc7681fd8aa5ad98946a90d0374701fdd8e2b5b6a7c5947393c

Observation eeccf834-057c-4c69-be2d-ae4783029d6c · outbound

This paper cites Scaling and Beyond: Advancing Spatial Reasoning in MLLMs Requires New Recipes.

Assessing the Value of Visual Input: A Benchmark of Multimodal Large Language Models for Robotic Path Planning Scaling and Beyond: Advancing Spatial Reasoning in MLLMs Requires New Recipes

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:09.379020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:09.379020Z digest=sha256:8b71bdca233d07a3f05d9ac953dae91dd444d7120dc98bc3756af2054b8c8cdb

Observation 6367c078-f26e-4924-bb48-ed3c51739195 · outbound

This paper cites Mitigating Cross-Modal Distraction and Ensuring Geometric Feasibility via Affordance- Guided, Self-Consistent MLLMs for Food Prepara- tion Task Planning.

Assessing the Value of Visual Input: A Benchmark of Multimodal Large Language Models for Robotic Path Planning Mitigating Cross-Modal Distraction and Ensuring Geometric Feasibility via Affordance- Guided, Self-Consistent MLLMs for Food Prepara- tion Task Planning

Reference 21

Resolution
verified exact
raw_fallback, observed 2026-08-06T16:50:09.900096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:50:09.470153Z digest=sha256:7ae3642ecca2e044c00b4a21a7fd5bd075f7fef8afd9a2fee8aa1fb1aacf5c49

Observation 7f6edbf4-224d-4902-ae90-0d0f32c2d5d7 · outbound

This paper cites Can Large Language Models be Good Path Planners? A Benchmark and Investigation on Spatial-temporal Reasoning.

Assessing the Value of Visual Input: A Benchmark of Multimodal Large Language Models for Robotic Path Planning Can Large Language Models be Good Path Planners? A Benchmark and Investigation on Spatial-temporal Reasoning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:09.535109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:09.535109Z digest=sha256:b113ec0657063aca043de281767e1b780b4511e7c77adf9babafdabf160e7242

Observation 20a633f0-d5a2-4c3f-93ba-942cbab5df6d · outbound

This paper cites EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents.

Assessing the Value of Visual Input: A Benchmark of Multimodal Large Language Models for Robotic Path Planning EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:09.623638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:09.623638Z digest=sha256:6e8c77ed23b5cf04a5203f99d61878d9169dfde5bc2267968de7275a8cb5ef81

Pith citing papers

Observation ff00f254-9fce-4018-87cd-dad413f06d7e · inbound

Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture cites this paper.

Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture Assessing the Value of Visual Input: A Benchmark of Multimodal Large Language Models for Robotic Path Planning

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-05T11:40:25.837858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-05T11:40:21.327031Z digest=sha256:5f02e4a1b1873ee0daecc00ade3d72697c0c05bba2396c7659541f48793b0fd6