Pith. sign in

Paper Citation Record · LEDGER

EmbodiedGPT: Vision-Language Pre-Training via Embodied Chain of Thought

As of 17 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 14 inbound Pith citation observations for arXiv:2305.15021.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2305.15021 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T19:06:38.499008Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-24T01:25:54.580634Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e22b3f59-15d9-4265-8545-7ebfc909ed16 · inbound

A Survey on Multimodal Large Language Models cites this paper.

A Survey on Multimodal Large Language Models EmbodiedGPT: Vision-Language Pre-Training via Embodied Chain of Thought

Reference 157

Resolution
verified exact
arxiv_id, observed 2026-05-16T02:56:42.009453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-16T02:56:41.658658Z digest=sha256:21e9ac7c3f9edf155d5a79daf48375332930b148dbbacc6ceceda277368877b2

Observation f4bdaa72-505a-480a-96eb-0578164cde0f · inbound

The Rise and Potential of Large Language Model Based Agents: A Survey cites this paper.

The Rise and Potential of Large Language Model Based Agents: A Survey EmbodiedGPT: Vision-Language Pre-Training via Embodied Chain of Thought

Reference 122

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:47:47.811843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-11T10:47:44.152066Z digest=sha256:0b7bc9c82eac3768d3d43e480726ffc66d0fa1e94402e5ca7b4a9e110b0188fc

Observation fb89be34-c3d6-4b79-979b-9854cd441244 · inbound

Open X-Embodiment: Robotic Learning Datasets and RT-X Models cites this paper.

Open X-Embodiment: Robotic Learning Datasets and RT-X Models EmbodiedGPT: Vision-Language Pre-Training via Embodied Chain of Thought

Reference 116

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:23:24.360032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-11T17:23:24.255829Z digest=sha256:53ee21634c5c2036206d265b205c0fbe667f562a1a13d4e29ff9fada5ef2e26c

Observation 6d22969d-fb1d-4bad-9f85-4875ad9ebaa3 · inbound

InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks cites this paper.

InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks EmbodiedGPT: Vision-Language Pre-Training via Embodied Chain of Thought

Reference 109

Resolution
verified exact
arxiv_id, observed 2026-05-13T22:46:10.072892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-13T22:46:09.693156Z digest=sha256:0450a7dbbfdcba984f80f5a58741a2c913a92500162426a19385beb2f9b61bc1

Observation e7427b10-da64-4515-bb17-faef5e9396a3 · inbound

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices cites this paper.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices EmbodiedGPT: Vision-Language Pre-Training via Embodied Chain of Thought

Reference 86

Resolution
verified exact
arxiv_id, observed 2026-05-16T16:35:38.024396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:8a878a609060cd9133601f1aaab4939afb08e5d4a38ae731b18c4ba66d1b77c6

Observation b32a5873-1895-4450-af0d-22abda013faf · inbound

A Survey on Vision-Language-Action Models for Embodied AI cites this paper.

A Survey on Vision-Language-Action Models for Embodied AI EmbodiedGPT: Vision-Language Pre-Training via Embodied Chain of Thought

Reference 142

Resolution
verified exact
arxiv_id, observed 2026-05-24T01:25:54.583569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-24T01:25:10.150459Z digest=sha256:432766745215fbca62e2d55f7b0b860452f0a6423e3bc2eeac0590572eab8b55

Observation bcef1a3b-470f-4f3b-b800-448028075787 · inbound

Aggregated Structural Representation with Large Language Models for Human-Centric Layout Generation cites this paper.

Aggregated Structural Representation with Large Language Models for Human-Centric Layout Generation EmbodiedGPT: Vision-Language Pre-Training via Embodied Chain of Thought

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:06.337735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:18:06.337735Z digest=sha256:917efa978bdfff1dc8a93ab7d423a769c0c85456da4ef54f1211a8e29db8506d

Observation a379916d-3750-496e-bd5e-d1b9bb0b6812 · inbound

Reinforced Reasoning for Embodied Planning cites this paper.

Reinforced Reasoning for Embodied Planning EmbodiedGPT: Vision-Language Pre-Training via Embodied Chain of Thought

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:23.750410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:23.750410Z digest=sha256:3d47cf483b8fe2bdd963620fb21d051f023c0820c89fadc608d04c487f6bd27f

Observation 7b35c8ee-390c-4f4a-9fa7-3ffb0547b308 · inbound

RoboMonkey: Scaling Test-Time Sampling and Verification for Vision-Language-Action Models cites this paper.

RoboMonkey: Scaling Test-Time Sampling and Verification for Vision-Language-Action Models EmbodiedGPT: Vision-Language Pre-Training via Embodied Chain of Thought

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T19:06:38.499008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:06:38.499008Z digest=sha256:015d21e14a4e8b3273a5cffe2f60679edee787236420c8d0e60e368dd51b671c

Observation 36e3c82d-b90b-4dab-a148-3541925de2a7 · inbound

TCPO: Thought-Centric Preference Optimization for Effective Embodied Decision-making cites this paper.

TCPO: Thought-Centric Preference Optimization for Effective Embodied Decision-making EmbodiedGPT: Vision-Language Pre-Training via Embodied Chain of Thought

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:00.908527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:00.908527Z digest=sha256:6c5adc2fc228ac9f9b21c120345e09d7dcab60ec7e929010e8a875d859e1854e

Observation b922f013-e444-4eac-bea0-709163224b4c · inbound

Think in Strokes, Not Pixels: Process-Driven Image Generation via Interleaved Reasoning cites this paper.

Think in Strokes, Not Pixels: Process-Driven Image Generation via Interleaved Reasoning EmbodiedGPT: Vision-Language Pre-Training via Embodied Chain of Thought

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:10:53.736351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T19:16:58.323955Z digest=sha256:108b8ac04c1edcd55d417311df44bcabac5e1b053ea52806ef1e1f2654980812

Observation c188561e-1a89-4e68-a47d-d5946e36366c · inbound

From Static Analysis to Audience Dissemination: A Training-Free Multimodal Controversy Detection Multi-Agent Framework cites this paper.

From Static Analysis to Audience Dissemination: A Training-Free Multimodal Controversy Detection Multi-Agent Framework EmbodiedGPT: Vision-Language Pre-Training via Embodied Chain of Thought

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:16:06.736427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-09T20:26:09.112189Z digest=sha256:c19bd885c0ff21463e159a143495865d44f66920a17f098097115688e714077e

Observation 52b4e99e-4ad0-4d0f-8ed1-e360a9b950f0 · inbound

RT-SHCUA: Real-Time Self-Hosted Computer-Use Agent for UAV Control cites this paper.

RT-SHCUA: Real-Time Self-Hosted Computer-Use Agent for UAV Control EmbodiedGPT: Vision-Language Pre-Training via Embodied Chain of Thought

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T16:35:14.298657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:35:14.298657Z digest=sha256:ca7165705546564383bd21c43b48ebfc95fafc6bd7ccd1c7417356b31dc5ff6d

Observation 51fa3e4b-0791-4803-8bc9-8570789a1135 · inbound

Data Pyramid for Embodied Manipulation: A Survey cites this paper.

Data Pyramid for Embodied Manipulation: A Survey EmbodiedGPT: Vision-Language Pre-Training via Embodied Chain of Thought

Reference 258

Resolution
unresolved
no resolver link, observed 2026-07-31T06:18:55.914327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:18:55.914327Z digest=sha256:64149aa583292c3a56f5638206b30cf0146a7e5272eaa1c23bf2891520c9134f