Pith. sign in

Paper Citation Record · LEDGER

Towards End-to-End Embodied Decision Making via Multi-modal Large Language Model: Explorations with GPT4-Vision and Beyond

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2310.02071.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.02071 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:42:16.234764Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T19:30:07.232112Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 7de0c3b4-2a71-4710-9273-fd710dfd9103 · inbound

Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations cites this paper.

Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations Towards End-to-End Embodied Decision Making via Multi-modal Large Language Model: Explorations with GPT4-Vision and Beyond

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:34:15.938776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-14T22:34:15.638114Z digest=sha256:20a2d1000ab141e39461e43be146dbca48b1bbc661cd23ae729b8526c63625f0

Observation 17ea7b7e-b8fd-4e11-8470-9104513acf82 · inbound

VisualSphinx: Large-Scale Synthetic Vision Logic Puzzles for RL cites this paper.

VisualSphinx: Large-Scale Synthetic Vision Logic Puzzles for RL Towards End-to-End Embodied Decision Making via Multi-modal Large Language Model: Explorations with GPT4-Vision and Beyond

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:16.234764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:16.234764Z digest=sha256:a9f7f729289645fb6c9842e9e792cc5b0f9a1ccffffba02d5507ea6ce896d68a

Observation 2d722460-63ac-4ef0-bdcd-672c1891413e · inbound

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents cites this paper.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents Towards End-to-End Embodied Decision Making via Multi-modal Large Language Model: Explorations with GPT4-Vision and Beyond

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:42.313802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:36:42.313802Z digest=sha256:a6752398fdbfee84d1e775b77fb69bd3e52b9618c9a7e0c99d03b2f3e54b1dde

Observation 391f5fc7-8137-4a86-810f-9b9ef62626dc · inbound

VeriOS: Query-Driven Proactive Human-Agent-GUI Interaction for Trustworthy OS Agents cites this paper.

VeriOS: Query-Driven Proactive Human-Agent-GUI Interaction for Trustworthy OS Agents Towards End-to-End Embodied Decision Making via Multi-modal Large Language Model: Explorations with GPT4-Vision and Beyond

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-18T18:06:42.715464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T18:06:12.349285Z digest=sha256:e46fd7461b1b5d4bd84413fabcddb2ffe90a95402de60e6a7b7146add1704a42

Observation 891377d2-ed42-4360-a8f0-76f793a5c286 · inbound

ASSCG: Just-Right Gating over Chattering for Fast-Slow LLM Planning in Autonomous Driving cites this paper.

ASSCG: Just-Right Gating over Chattering for Fast-Slow LLM Planning in Autonomous Driving Towards End-to-End Embodied Decision Making via Multi-modal Large Language Model: Explorations with GPT4-Vision and Beyond

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T19:30:07.233474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-25T21:20:14.654132Z digest=sha256:d8ec0569fdf1df7d203de44ff94f04e04904f8f5f98684f1be6fa234277bf60b

Observation f2a2672f-19d1-4f01-943d-2249ddd162fe · inbound

MARS-RA: Rank Aggregation for Credit Assignment via Multimodal Comparisons in Embodied Multi-Agent Cooperation cites this paper.

MARS-RA: Rank Aggregation for Credit Assignment via Multimodal Comparisons in Embodied Multi-Agent Cooperation Towards End-to-End Embodied Decision Making via Multi-modal Large Language Model: Explorations with GPT4-Vision and Beyond

Reference 74

Resolution
unresolved
no resolver link, observed 2026-07-31T21:55:17.407982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T21:55:17.407982Z digest=sha256:e6d2e02d37856697d488e312c8333fa6cfee108e9952596527b71106b5ce47ad