Pith. sign in

Paper Citation Record · LEDGER

EmbodiedGPT: Vision-Language Pre-Training via Embodied Chain of Thought

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2305.15021.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2305.15021 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:18:06.337735Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-24T01:25:54.580634Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e22b3f59-15d9-4265-8545-7ebfc909ed16 · inbound

A Survey on Multimodal Large Language Models cites this paper.

A Survey on Multimodal Large Language Models EmbodiedGPT: Vision-Language Pre-Training via Embodied Chain of Thought

Reference 157

Resolution
verified exact
arxiv_id, observed 2026-05-16T02:56:42.009453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T02:56:41.658658Z digest=sha256:49040a50af0a15a47e6629da424cdda322d92fde50604286925dd589dd3233fa

Observation f4bdaa72-505a-480a-96eb-0578164cde0f · inbound

The Rise and Potential of Large Language Model Based Agents: A Survey cites this paper.

The Rise and Potential of Large Language Model Based Agents: A Survey EmbodiedGPT: Vision-Language Pre-Training via Embodied Chain of Thought

Reference 122

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:47:47.811843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T10:47:44.152066Z digest=sha256:ff6dc6e26a8341ecbb93d035359ef43e40bc8ce4b7c907db998422159e2288f9

Observation fb89be34-c3d6-4b79-979b-9854cd441244 · inbound

Open X-Embodiment: Robotic Learning Datasets and RT-X Models cites this paper.

Open X-Embodiment: Robotic Learning Datasets and RT-X Models EmbodiedGPT: Vision-Language Pre-Training via Embodied Chain of Thought

Reference 116

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:23:24.360032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T17:23:24.255829Z digest=sha256:88e342b64b2444cb5847c3d45ef61367498dffff5fa6a24fe8308fa23e3dd0aa

Observation 6d22969d-fb1d-4bad-9f85-4875ad9ebaa3 · inbound

InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks cites this paper.

InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks EmbodiedGPT: Vision-Language Pre-Training via Embodied Chain of Thought

Reference 109

Resolution
verified exact
arxiv_id, observed 2026-05-13T22:46:10.072892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T22:46:09.693156Z digest=sha256:7894bd89ed9d2e6d113bbeb09e6c19de1efb104b2be5817d655e442f17649947

Observation e7427b10-da64-4515-bb17-faef5e9396a3 · inbound

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices cites this paper.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices EmbodiedGPT: Vision-Language Pre-Training via Embodied Chain of Thought

Reference 86

Resolution
verified exact
arxiv_id, observed 2026-05-16T16:35:38.024396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:9aa84539b27b5046b58c88fe2f13c1adadaf9073eb4c477e873ff0a80a9109a9

Observation b32a5873-1895-4450-af0d-22abda013faf · inbound

A Survey on Vision-Language-Action Models for Embodied AI cites this paper.

A Survey on Vision-Language-Action Models for Embodied AI EmbodiedGPT: Vision-Language Pre-Training via Embodied Chain of Thought

Reference 142

Resolution
verified exact
arxiv_id, observed 2026-05-24T01:25:54.583569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-24T01:25:10.150459Z digest=sha256:a9ff90a643fec748e368f373b2b95de0c0127da575f6a61ffd301878fbf3ac84

Observation bcef1a3b-470f-4f3b-b800-448028075787 · inbound

Aggregated Structural Representation with Large Language Models for Human-Centric Layout Generation cites this paper.

Aggregated Structural Representation with Large Language Models for Human-Centric Layout Generation EmbodiedGPT: Vision-Language Pre-Training via Embodied Chain of Thought

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:06.337735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:18:06.337735Z digest=sha256:fccfb490bf5525c4ac1a7e47e4bdf2dae5cadc05a649e7b61864449d0ab755b6

Observation a379916d-3750-496e-bd5e-d1b9bb0b6812 · inbound

Reinforced Reasoning for Embodied Planning cites this paper.

Reinforced Reasoning for Embodied Planning EmbodiedGPT: Vision-Language Pre-Training via Embodied Chain of Thought

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:23.750410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:23.750410Z digest=sha256:2b0badcdac98b3182233933dce450cc6ca6a93d9daaeb8c4112a5fab60900119

Observation b922f013-e444-4eac-bea0-709163224b4c · inbound

Think in Strokes, Not Pixels: Process-Driven Image Generation via Interleaved Reasoning cites this paper.

Think in Strokes, Not Pixels: Process-Driven Image Generation via Interleaved Reasoning EmbodiedGPT: Vision-Language Pre-Training via Embodied Chain of Thought

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:10:53.736351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T19:16:58.323955Z digest=sha256:7ba33eb9fe84c5a9be9e88b0cb7952666a9d82e0876d5f8aac4466f7b9abea3f

Observation c188561e-1a89-4e68-a47d-d5946e36366c · inbound

From Static Analysis to Audience Dissemination: A Training-Free Multimodal Controversy Detection Multi-Agent Framework cites this paper.

From Static Analysis to Audience Dissemination: A Training-Free Multimodal Controversy Detection Multi-Agent Framework EmbodiedGPT: Vision-Language Pre-Training via Embodied Chain of Thought

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:16:06.736427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T20:26:09.112189Z digest=sha256:dadad79f0b25378a3378f85d3e71336289d1fb43011e63f89e341c8f85c3cb17

Observation 52b4e99e-4ad0-4d0f-8ed1-e360a9b950f0 · inbound

RT-SHCUA: Real-Time Self-Hosted Computer-Use Agent for UAV Control cites this paper.

RT-SHCUA: Real-Time Self-Hosted Computer-Use Agent for UAV Control EmbodiedGPT: Vision-Language Pre-Training via Embodied Chain of Thought

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T16:35:14.298657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:35:14.298657Z digest=sha256:d810b5b3b4967a29610c1c568cdf491e51d5f4cd73f67e5bbae033832a107817

Observation 51fa3e4b-0791-4803-8bc9-8570789a1135 · inbound

Data Pyramid for Embodied Manipulation cites this paper.

Data Pyramid for Embodied Manipulation EmbodiedGPT: Vision-Language Pre-Training via Embodied Chain of Thought

Reference 258

Resolution
unresolved
no resolver link, observed 2026-07-31T06:18:55.914327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:18:55.914327Z digest=sha256:9c68375b1f7670ae3bdc8ff016b84cbc92cc3742b3ec935963f17b56c6827085