Pith. sign in

Paper Citation Record · LEDGER

Reinforced Reasoning for Embodied Planning

As of 10 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 6 inbound Pith citation observations for arXiv:2505.22050.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.22050 v2

Coverage vector

measured 65 of 65 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:22:26.851204Z

measured 71 of 71 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T04:50:32.142169Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-17T20:28:15.942007Z

Reference resolution

65 of 65 outbound references displayed

  • verified exact2
  • verified fuzzy15
  • unresolved47
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 24e9227a-ffd2-414c-8a7c-fc9cb3a023e7 · outbound

This paper cites URL: https://openai.com/index/ gpt-4o-mini-advancing-cost-efficient-intelligence/.

Reinforced Reasoning for Embodied Planning URL: https://openai.com/index/ gpt-4o-mini-advancing-cost-efficient-intelligence/

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:22:31.068071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:22:20.697606Z digest=sha256:ad55730567ea5f4979f51c75e5dd52e811c391a7a7a99aa3908e1ed6b5374963

Observation 65f69c35-873e-4102-b65b-9e0a69ef96ee · outbound

This paper cites URL: https://openai.com/index/hello-gpt-4o/.

Reinforced Reasoning for Embodied Planning URL: https://openai.com/index/hello-gpt-4o/

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:22:30.915929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:22:20.785954Z digest=sha256:f60c604ced7072521548934d1fa363fe37e1de944c14f54c099fc2b00cbadc78

Observation 3b59e599-9cd7-4bf1-9ca2-cdcc395a60a3 · outbound

This paper cites URL: https://www.anthropic.com/news/ claude-3-5-sonnet.

Reinforced Reasoning for Embodied Planning URL: https://www.anthropic.com/news/ claude-3-5-sonnet

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:22:30.704191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:22:20.861921Z digest=sha256:0ea1f0e38d82755770d7d2492f129e0dc9379f0c3cd91a7fbbcdd877b7601536

Observation 3eb58784-d7ba-4e28-ad95-8d9006b35e7f · outbound

This paper cites URL: https://blog.

Reinforced Reasoning for Embodied Planning URL: https://blog

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:22:30.537405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:22:20.971449Z digest=sha256:9b3fd8a890f7a20180b76bc71b69e8e97fd442c6d0c2f276d9997f82c1be2785

Observation 94ffbd43-4381-4b69-acce-c470cc9f74ba · outbound

This paper cites URL: https://ai.meta.com/blog/ llama-3-2-connect-2024-vision-edge-mobile-devices/.

Reinforced Reasoning for Embodied Planning URL: https://ai.meta.com/blog/ llama-3-2-connect-2024-vision-edge-mobile-devices/

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:22:30.286481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:22:21.081823Z digest=sha256:a17b334649c64247f704ead8bf470127f5857489704a2e329966cbb43ed29d3e

Observation 6d72c2fe-222c-4599-935a-6740ea66b12a · outbound

This paper cites Do As I Can, Not As I Say: Grounding Language in Robotic Affordances.

Reinforced Reasoning for Embodied Planning Do As I Can, Not As I Say: Grounding Language in Robotic Affordances

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:21.227527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:21.227527Z digest=sha256:ff5d3551899cbf82b3eeb1c57a8e9741c602c85b1c45d179fface3e92d9f67c4

Observation 1466b7fc-2907-4e6d-a50a-285ceff2fbb6 · outbound

This paper cites Qwen2.5-VL Technical Report.

Reinforced Reasoning for Embodied Planning Qwen2.5-VL Technical Report

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:21.376588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:21.376588Z digest=sha256:c6e04f2a108fa400e42fceeac1495bb94fcb19a3ab11158ded2835ecdd0d7703

Observation 15d6ded9-411f-4a3c-8807-c8451aa847fe · outbound

This paper cites RoboGPT: an intelligent agent of making embodied long-term decisions for daily instruction tasks.

Reinforced Reasoning for Embodied Planning RoboGPT: an intelligent agent of making embodied long-term decisions for daily instruction tasks

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:21.473781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:21.473781Z digest=sha256:413eabc089a437d822636e1a6cab51c9e6b0703a352554ee3f06a1a0eeec3bc2

Observation ab1fe195-2f7d-4b5c-bb0e-a44cd8a1818e · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

Reinforced Reasoning for Embodied Planning Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:21.610900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:21.610900Z digest=sha256:4873f80de52bbb3eb3ce16de6ababe48b727054e0287fefb449b7517305566ce

Observation 7eaaf131-0b02-4d53-bea1-345e82ad6be2 · outbound

This paper cites Process Reinforcement through Implicit Rewards.

Reinforced Reasoning for Embodied Planning Process Reinforcement through Implicit Rewards

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:21.739893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:21.739893Z digest=sha256:a599ded32182554079439b95db8be3f828258abd419e316bbe7601df0d747590

Observation ba721db9-318d-48eb-bfe2-050a9a589979 · outbound

This paper cites A survey of embodied ai: From simulators to research tasks.

Reinforced Reasoning for Embodied Planning A survey of embodied ai: From simulators to research tasks

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:22:30.065537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:22:21.894396Z digest=sha256:f76b147b5de2dc7393a0031826211477e4ea0731b06979ee94331901b39a2c83

Observation 144e999e-c563-48ce-84b8-313f1e88c37d · outbound

This paper cites Open r1: A fully open reproduction of deepseek-r1, January 2025.

Reinforced Reasoning for Embodied Planning Open r1: A fully open reproduction of deepseek-r1, January 2025

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:22:29.814108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:22:22.008668Z digest=sha256:42cb55e9831a500cac6bcb0674459d3b64a59448ed1638ee9760bb09c3b330f4

Observation 521e9aaa-7842-462c-aa30-ced4c358632e · outbound

This paper cites What can vlms do for zero-shot embodied task planning? In ICML 2024 Workshop on LLMs and Cognition, 2024.

Reinforced Reasoning for Embodied Planning What can vlms do for zero-shot embodied task planning? In ICML 2024 Workshop on LLMs and Cognition, 2024

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:22:29.618245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:22:22.096292Z digest=sha256:0b6f8238494f3a86c6ce86eba236d3f85d69af42ab7e4219d0e7faa99eba289f

Observation 486f722f-2e98-4710-994c-fcb43486f642 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Reinforced Reasoning for Embodied Planning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:22.190695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:22.190695Z digest=sha256:3499dcd40d5f4a5bf575d22ae0833ac169f12395f2411900041d04a3aacab9e1

Observation b35a1a89-f812-4f18-8a4e-ba084039e568 · outbound

This paper cites Lora: Low-rank adaptation of large language models.

Reinforced Reasoning for Embodied Planning Lora: Low-rank adaptation of large language models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:22.314432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:22.314432Z digest=sha256:3826ef806a613232550952eac0bdbfb4f2e523015a68f9d27df4be82625b8057

Observation f5814911-feb8-457b-8a69-998ae1b6c810 · outbound

This paper cites OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework.

Reinforced Reasoning for Embodied Planning OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:22.439915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:22.439915Z digest=sha256:e3ae6ef57bf90dc7fa7b47cf0ac3f69a9041de75792ff3aefb9414b555e3242c

Observation 306cb825-6fc5-49d0-b7ac-c2cf156d7404 · outbound

This paper cites Look Before You Leap: Unveiling the Power of GPT-4V in Robotic Vision-Language Planning.

Reinforced Reasoning for Embodied Planning Look Before You Leap: Unveiling the Power of GPT-4V in Robotic Vision-Language Planning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:22.543422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:22.543422Z digest=sha256:0c26a510b448edcddb54b4d0059067531598d7f619805f623e16ac5b72c2818c

Observation 953bad54-af75-45d8-a8f4-18ea783b6597 · outbound

This paper cites Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models.

Reinforced Reasoning for Embodied Planning Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:22.647159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:22.647159Z digest=sha256:8fa0428f590051809508ed0a092575fcb11558fd23a318832dbf31144f7fc05c

Observation 0e8cbed3-1426-4168-89c8-6ca6e30aecaa · outbound

This paper cites RoboBrain: A Unified Brain Model for Robotic Manipulation from Abstract to Concrete.

Reinforced Reasoning for Embodied Planning RoboBrain: A Unified Brain Model for Robotic Manipulation from Abstract to Concrete

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:22.795873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:22.795873Z digest=sha256:bd1dd3d0c19acd07c77697420fd4c5cb0ddb539a9e3972fb690dcdde3eeb810c

Observation fb0633e3-f0b5-4c7c-8c62-12fad8a46a5a · outbound

This paper cites Context-Aware Planning and Environment-Aware Memory for Instruction Following Embodied Agents.

Reinforced Reasoning for Embodied Planning Context-Aware Planning and Environment-Aware Memory for Instruction Following Embodied Agents

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:22:27.092321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:22:22.902090Z digest=sha256:54f860351f9056d22eea5dda4975cc98a435ccc81b1d04516f95ba508ae1ffe7

Observation 015d696c-6ac0-4202-80c8-9d98520e52c0 · outbound

This paper cites Openvla: An open-source vision-language-action model.

Reinforced Reasoning for Embodied Planning Openvla: An open-source vision-language-action model

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:22:29.432336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:22:23.000783Z digest=sha256:3e126c40874f3bdc89fe1aacf2a33568244b67f9ec9b21caf5f0f72c8f98728d

Observation 59e6988c-2cf4-4ed4-b91d-b3763ddf060e · outbound

This paper cites AI2-THOR: An Interactive 3D Environment for Visual AI.

Reinforced Reasoning for Embodied Planning AI2-THOR: An Interactive 3D Environment for Visual AI

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:23.067075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:23.067075Z digest=sha256:4cfdbb354cc6de93b822c11ab10722437d68f8577bff88634748eb657cd2d783

Observation 0a9a00ae-3fd0-46c2-be4e-a04520f0240e · outbound

This paper cites VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning.

Reinforced Reasoning for Embodied Planning VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:23.131565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:23.131565Z digest=sha256:7472dcf8b7270df96a7652bff1294e2a5b188431ba2c10af07f4c3e32d89aa0d

Observation cd364a8c-639e-4932-8fb5-0d90a571eece · outbound

This paper cites Let’s verify step by step.

Reinforced Reasoning for Embodied Planning Let’s verify step by step

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:23.205810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:23.205810Z digest=sha256:020c9ac738de68b9f4a3eefb65bb1f9209109d67d53f33745a2fb369e0adf5d2

Observation 068979f8-27f0-45ce-abe9-af07d4cc13f9 · outbound

This paper cites Visual-RFT: Visual Reinforcement Fine-Tuning.

Reinforced Reasoning for Embodied Planning Visual-RFT: Visual Reinforcement Fine-Tuning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:23.320295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:23.320295Z digest=sha256:81a54f707d43b46a27820976ed928c7e695ea8f30470d5b10db4e673f340615a

Observation 75ed4314-1218-4b26-9220-0e431c3a42f3 · outbound

This paper cites A Survey on Vision-Language-Action Models for Embodied AI.

Reinforced Reasoning for Embodied Planning A Survey on Vision-Language-Action Models for Embodied AI

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:23.425995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:23.425995Z digest=sha256:1a01ce8609f04f2acdef4a6f5c7f54a0b8ef1e4f52b2ffb858bebf6bf11b68d8

Observation 594ba930-4cde-4303-9a9a-16a9b0917d39 · outbound

This paper cites MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning.

Reinforced Reasoning for Embodied Planning MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:23.503690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:23.503690Z digest=sha256:2e7cf58b854d58b8e5440499e8e721b48eb06d2f797d62c0ffad9fe4ab4ef04e

Observation c4064ed7-2bcd-4381-970d-e4a9ac66cc24 · outbound

This paper cites Composi- tional chain-of-thought prompting for large multimodal models.

Reinforced Reasoning for Embodied Planning Composi- tional chain-of-thought prompting for large multimodal models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:22:29.284038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:22:23.564040Z digest=sha256:ba41b76655aeb3ab5016f73e27af33fcbedc4ef55b66837df16027c948546ea6

Observation e6233108-f2b8-4000-acc6-16e29d635747 · outbound

This paper cites Kam-cot: Knowledge augmented multimodal chain-of-thoughts reasoning.

Reinforced Reasoning for Embodied Planning Kam-cot: Knowledge augmented multimodal chain-of-thoughts reasoning

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:22:29.088642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:22:23.666632Z digest=sha256:ac31f68df1b7f4f250efed21bcc3643119d9358e73921a096dbf05ee438c76bd

Observation a379916d-3750-496e-bd5e-d1b9bb0b6812 · outbound

This paper cites EmbodiedGPT: Vision-Language Pre-Training via Embodied Chain of Thought.

Reinforced Reasoning for Embodied Planning EmbodiedGPT: Vision-Language Pre-Training via Embodied Chain of Thought

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:23.750410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:23.750410Z digest=sha256:2b0badcdac98b3182233933dce450cc6ca6a93d9daaeb8c4112a5fab60900119

Observation 2893fc7d-e8dd-4666-bbd9-bb90b0675797 · outbound

This paper cites Training language models to follow instructions with human feedback.

Reinforced Reasoning for Embodied Planning Training language models to follow instructions with human feedback

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:23.820276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:23.820276Z digest=sha256:4c1bec33c06e5c116e78dd42f728af9bfd47a560c267cada842b78ad824cad16

Observation b6b71b1c-ea92-4cd9-9494-f9eafb0d7505 · outbound

This paper cites Reasoning with large language models, a survey.

Reinforced Reasoning for Embodied Planning Reasoning with large language models, a survey

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:23.927663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:23.927663Z digest=sha256:ea714e7186119820a19255b69a8d6c17900788e20df75b1ea35bac7406110b20

Observation d138363b-8b9b-4f20-9d31-e6b264d3761b · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Reinforced Reasoning for Embodied Planning Direct preference optimization: Your language model is secretly a reward model

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:24.070429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:24.070429Z digest=sha256:177cb673059e5ab8df1ec55fa080fcb6d25ff25fc977995301721c63ca02e474

Observation 2ae79c2a-828f-4a90-8707-f0c4835ee5d8 · outbound

This paper cites Say- Plan: Grounding large language models using 3d scene graphs for scalable robot task planning.

Reinforced Reasoning for Embodied Planning Say- Plan: Grounding large language models using 3d scene graphs for scalable robot task planning

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:22:28.955807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:22:24.159911Z digest=sha256:dda77afd9e546e25510a3f00ab9f01ca4f373d57be5fd91b6b731d8ee51f7703

Observation 2673d062-04d3-4f5a-b57a-7fce5fcb0cac · outbound

This paper cites Habitat: A platform for embodied ai research.

Reinforced Reasoning for Embodied Planning Habitat: A platform for embodied ai research

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:24.234147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:24.234147Z digest=sha256:48aa44e3c6f781f87f3fca4d967889ef931329ae1ad868d5e4b651fe09cf3bdf

Observation d065253d-040a-4b08-a27e-ec1e2567bcfe · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Reinforced Reasoning for Embodied Planning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:24.317114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:24.317114Z digest=sha256:aadbf594a0b2c4c73d40cbea428f5087fe5eb9b4cadca52dd300d007f74e9d5d

Observation 82715e8d-d392-48f3-ab1c-027a5694d301 · outbound

This paper cites VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model.

Reinforced Reasoning for Embodied Planning VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:24.441874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:24.441874Z digest=sha256:13fb4132d9c5c23430b0342b89ce3be11d05b4106d0413c88eefe44acabaaf40

Observation 695f2968-25c7-4f31-8cea-ac6239232391 · outbound

This paper cites Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models.

Reinforced Reasoning for Embodied Planning Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:24.515252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:24.515252Z digest=sha256:d4181dde283421e74b57eca6b9746e2eb66eca9d02a2bd4b8c7e29c1648a17e8

Observation ccb339cc-995d-455f-be9d-1b86e9eb32e4 · outbound

This paper cites Socratic Planner: Self-QA-Based Zero-Shot Planning for Embodied Instruction Following.

Reinforced Reasoning for Embodied Planning Socratic Planner: Self-QA-Based Zero-Shot Planning for Embodied Instruction Following

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:22:27.751720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:22:24.564479Z digest=sha256:a99a881e51ec7d3e3fa11acc32701f5c9e24129710b0f5c15e955f24a535de2f

Observation e46ce3b6-0470-4a02-a8e7-14900943e03c · outbound

This paper cites Alfred: A benchmark for interpreting grounded instructions for everyday tasks.

Reinforced Reasoning for Embodied Planning Alfred: A benchmark for interpreting grounded instructions for everyday tasks

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:22:28.729821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:22:24.632852Z digest=sha256:f70639fb05b63388229bb8a1419a2eccb961cf89a3af2f51a57b818066f3777e

Observation 9082b7a0-ca47-426e-8226-73f21312d533 · outbound

This paper cites Tenenbaum, Leslie Kaelbling, and Michael Katz.

Reinforced Reasoning for Embodied Planning Tenenbaum, Leslie Kaelbling, and Michael Katz

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:24.721896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:24.721896Z digest=sha256:c31ec9a9ed9ef663c576b880149579a721545401158bd6d241be37f2794d19a1

Observation 6c45aed3-41bc-4ada-bbe9-090799f28ba2 · outbound

This paper cites ProgPrompt: Generating Situated Robot Task Plans using Large Language Models.

Reinforced Reasoning for Embodied Planning ProgPrompt: Generating Situated Robot Task Plans using Large Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:24.847360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:24.847360Z digest=sha256:9cdec1fa33530cc500517ebd245250d87df5cecf986b8bfb03e4352654678b0a

Observation 5083a2e4-4759-4062-869c-bc31d46e0826 · outbound

This paper cites Trial and Error: Exploration-Based Trajectory Optimization for LLM Agents.

Reinforced Reasoning for Embodied Planning Trial and Error: Exploration-Based Trajectory Optimization for LLM Agents

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:24.912504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:24.912504Z digest=sha256:2379b1e98f4d7ea5c9b7f1439d974ef451c1594e7f9fcfd83ece57628fe85c45

Observation 276bf8bd-f071-4bbf-b57f-362978007a35 · outbound

This paper cites Reason-rft: Reinforcement fine-tuning for visual reasoning.

Reinforced Reasoning for Embodied Planning Reason-rft: Reinforcement fine-tuning for visual reasoning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:25.024255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:25.024255Z digest=sha256:902a20e19c455e42906d25c581bda8dbdb871cc4d3f160b328f405705a750cb9

Observation 5a330e41-731c-4acf-98b7-a7e1e58d8eb7 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

Reinforced Reasoning for Embodied Planning Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:25.118215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:25.118215Z digest=sha256:6ef5fef573f023dabbbab36645bfd1f5298569cbc22b5c8e9bd38b2a3b3de418

Observation 1ef661ec-05e2-4e3e-be61-65c46d171ec2 · outbound

This paper cites World Modeling Makes a Better Planner: Dual Preference Optimization for Embodied Task Planning.

Reinforced Reasoning for Embodied Planning World Modeling Makes a Better Planner: Dual Preference Optimization for Embodied Task Planning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:25.221513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:25.221513Z digest=sha256:5e59a78ed13a4ad4258565029957d42b675e71a0de42d3b3f2dc44a343164983

Observation ea7a803f-cf6c-450e-8f42-ca62b9577c92 · outbound

This paper cites Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey.

Reinforced Reasoning for Embodied Planning Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:25.297429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:25.297429Z digest=sha256:5c1c3d6208785c4b53844e58da8fcff8011616fb319339b84184ff8331b255f3

Observation 41ebaff6-6b3f-4a3f-b640-c6982e3f7ce5 · outbound

This paper cites Exploring the Reasoning Abilities of Multimodal Large Language Models (MLLMs): A Comprehensive Survey on Emerging Trends in Multimodal Reasoning.

Reinforced Reasoning for Embodied Planning Exploring the Reasoning Abilities of Multimodal Large Language Models (MLLMs): A Comprehensive Survey on Emerging Trends in Multimodal Reasoning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:25.358034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:25.358034Z digest=sha256:78964720e0465fe8d1afde6d72de14ed249bdc04c023308e535385d6cc193a5d

Observation fc5ff3e5-e4fe-45b7-9a14-d9e94a21e443 · outbound

This paper cites RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning.

Reinforced Reasoning for Embodied Planning RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:25.429370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:25.429370Z digest=sha256:a16358f53ea3555088160d44f9200bdc665acf095e26b9bc752efed9e3c7577b

Observation 15fc3c6d-8b35-4c25-836b-2832a0b31c7f · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

Reinforced Reasoning for Embodied Planning Chain-of-thought prompting elicits reasoning in large language models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:25.565706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:25.565706Z digest=sha256:64d8233bade711e33f06a7fbdad0a0da5f29b4b3886534f3fd5b7319a1e32d14

Observation 5a3083d7-dbd6-4f84-9679-685ec544d6eb · outbound

This paper cites Embodied Task Planning with Large Language Models.

Reinforced Reasoning for Embodied Planning Embodied Task Planning with Large Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:25.663807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:25.663807Z digest=sha256:45c8efc469d2190373e111156b815ddf3b0f501aa1d461c5dc751c7515e3405d

Observation 4b9d0635-0905-4b8f-a6d7-e53cb7270b75 · outbound

This paper cites The rise and potential of large language model based agents: A survey.

Reinforced Reasoning for Embodied Planning The rise and potential of large language model based agents: A survey

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:25.727513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:25.727513Z digest=sha256:d8a457902acf9803dfead28c566f1eedaa26266725a44cace4038833809eeeb3

Observation cf71a2cf-9922-41d1-8c85-ac16f9273565 · outbound

This paper cites A Survey on Robotics with Foundation Models: toward Embodied AI.

Reinforced Reasoning for Embodied Planning A Survey on Robotics with Foundation Models: toward Embodied AI

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:25.853615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:25.853615Z digest=sha256:39a11820b5f588dfdf005fb978a4d5cb2d2ed2995c0b56bf0ddfeda14e2dad92

Observation 33ed233e-2815-4edb-9741-68178e0e3971 · outbound

This paper cites EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents.

Reinforced Reasoning for Embodied Planning EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:25.974006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:25.974006Z digest=sha256:0d135e4698f827d9a8c3568ced07b213a468066105e8d9e44416cd918720b285

Observation b0ed911a-61c7-49ab-87f3-87547d71b64f · outbound

This paper cites LIMO: Less is More for Reasoning.

Reinforced Reasoning for Embodied Planning LIMO: Less is More for Reasoning

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:26.046698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:26.046698Z digest=sha256:b768c6d633ba41f1cc859bc1731b3a9a0f305e9d04fd95f5fd33521bc534c350

Observation d3d90c58-7546-4690-99e6-d95fd96565d2 · outbound

This paper cites Robotic control via embodied chain-of-thought reasoning.

Reinforced Reasoning for Embodied Planning Robotic control via embodied chain-of-thought reasoning

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:22:28.569347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:22:26.135876Z digest=sha256:ea946eaf1a3e4d43fcb20516922e09d1239117bdbedbcac20e91fecf6394c49b

Observation 7b573dbf-8950-48cd-911b-78c2857b79dd · outbound

This paper cites HiRT: Enhancing Robotic Control with Hierarchical Robot Transformers.

Reinforced Reasoning for Embodied Planning HiRT: Enhancing Robotic Control with Hierarchical Robot Transformers

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:26.186961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:26.186961Z digest=sha256:254e10b99e39480617388b6199c717a2ef1649b853e5858d502b68541f1cbdb4

Observation e896157e-17c9-47f5-bca2-da733487fbde · outbound

This paper cites Vision-language models for vision tasks: A survey.

Reinforced Reasoning for Embodied Planning Vision-language models for vision tasks: A survey

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:26.234405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:26.234405Z digest=sha256:4fd19a9522e4979ad5bdc77f32839fc1a2615470b4b2a827be3e78521e5955a8

Observation 2401c5a6-982c-49ee-81e2-e35b5de658ad · outbound

This paper cites R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization.

Reinforced Reasoning for Embodied Planning R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:26.330503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:26.330503Z digest=sha256:bd4cfab64713b5337bb8a5371729390dc63ba8ebb126c77cbbd0716638b4b3a1

Observation 3365ba74-cedf-4a8a-bb6a-1d6134f17680 · outbound

This paper cites Embodied-Reasoner: Synergizing Visual Search, Reasoning, and Action for Embodied Interactive Tasks.

Reinforced Reasoning for Embodied Planning Embodied-Reasoner: Synergizing Visual Search, Reasoning, and Action for Embodied Interactive Tasks

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:26.448609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:26.448609Z digest=sha256:3aff4bf8af95b6b95a9932e305b23ea1aa3c67c243ebeba44368c2271b74f6c4

Observation ba02ac77-4754-4ff3-8aeb-e459fef570f7 · outbound

This paper cites Multimodal Chain-of-Thought Reasoning in Language Models.

Reinforced Reasoning for Embodied Planning Multimodal Chain-of-Thought Reasoning in Language Models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:26.520134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:26.520134Z digest=sha256:8e5c33df35983e43aa4105c9aa6de756363cbbcd4dfc378ec3be4c6598791d30

Observation 72e92544-ffec-44c3-805d-e83ce60e3698 · outbound

This paper cites Embodied-R: Collaborative Framework for Activating Embodied Spatial Reasoning in Foundation Models via Reinforcement Learning.

Reinforced Reasoning for Embodied Planning Embodied-R: Collaborative Framework for Activating Embodied Spatial Reasoning in Foundation Models via Reinforcement Learning

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:26.571212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:26.571212Z digest=sha256:3c20c8d83a70e879af6f9121f65d24620d9b8ba843d17b72e0409dc0f49d3099

Observation d2b29aa4-dbbc-4fdb-8da3-7385faca2cff · outbound

This paper cites LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models.

Reinforced Reasoning for Embodied Planning LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models

Reference 63

Resolution
malformed identifier
no resolver link, observed 2026-08-07T13:22:26.635776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:26.635776Z digest=sha256:eec4d0549d922804f389e4bee30eb6f7eaa497b791759c37a0d07d0d31a9c3de

Observation 1f57eff8-5fb6-4f00-974e-911afd275ee8 · outbound

This paper cites reasoning_and_reflection\.

Reinforced Reasoning for Embodied Planning reasoning_and_reflection\

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:22:28.353645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:22:26.700111Z digest=sha256:11bd1ab649bda20b007f780f9bc002217ad9b50879cb9ed31a91a9dc6b39dc80

Observation 9bede566-deae-4e2d-a0ba-ca8adacd5eec · outbound

This paper cites an unresolved cited work.

Reinforced Reasoning for Embodied Planning Unresolved cited work

Reference 224

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:22:28.207679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:22:26.851204Z digest=sha256:2a508874e034d3074dfd54c50d1cc5c9841a9d01dc65efdeb8fee0e1556ee11d

Pith citing papers

Observation dc9d4a05-26d6-4f64-b7e6-b01053167dd8 · inbound

Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey cites this paper.

Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey Reinforced Reasoning for Embodied Planning

Reference 150

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:28:15.947442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T20:28:15.818016Z digest=sha256:fd5cbc67fe12979cb49cca5a81af5af17a0645d6ccc1736373a4a80035515f08

Observation 8355abd1-3ad3-4a41-b036-c0dc8ba2d1d3 · inbound

Large Language Models for Next-Generation Wireless Network Management: A Survey and Tutorial cites this paper.

Large Language Models for Next-Generation Wireless Network Management: A Survey and Tutorial Reinforced Reasoning for Embodied Planning

Reference 168

Resolution
unresolved
no resolver link, observed 2026-08-05T04:50:32.142169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:50:32.142169Z digest=sha256:76fb0358ea92a48911080b37910a9b54025cf5a6491f3564ae607f0010445a35

Observation 228d765b-1279-480a-b96e-82f25ef54e93 · inbound

RoboGPT-R1: Enhancing Robot Task Planning with Reinforcement Learning cites this paper.

RoboGPT-R1: Enhancing Robot Task Planning with Reinforcement Learning Reinforced Reasoning for Embodied Planning

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-04T09:33:41.474386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:33:41.474386Z digest=sha256:79a362a0cda640956835d644f1ff049d308bce646d63f70274646603a9aac450

Observation d0aefcf2-b090-463b-a156-431113c66b6c · inbound

RoboAgent: Chaining Basic Capabilities for Embodied Task Planning cites this paper.

RoboAgent: Chaining Basic Capabilities for Embodied Task Planning Reinforced Reasoning for Embodied Planning

Reference 115

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:15:57.247031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T18:15:08.727921Z digest=sha256:f2b9a8855c6e1d390f6e48e106c9892438124f18a3fcb6e42429bf47c7c2d223

Observation 29bc1a9f-64c8-42e9-b5c6-1f94a5eab6b9 · inbound

Odysseus: Scaling VLMs to 100+ Turn Decision-Making in Games via Reinforcement Learning cites this paper.

Odysseus: Scaling VLMs to 100+ Turn Decision-Making in Games via Reinforcement Learning Reinforced Reasoning for Embodied Planning

Reference 66

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:16:09.217964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-09T20:22:58.061772Z digest=sha256:34300b15e1c3ef0fde9ae6f8dc78e6633f120e245de870230d9cc99661aed333

Observation d7ee1b53-19af-41f1-91ac-da99f4e9d0b8 · inbound

RoboEvolve: Co-Evolving Planner-Simulator for Robotic Manipulation with Limited Data cites this paper.

RoboEvolve: Co-Evolving Planner-Simulator for Robotic Manipulation with Limited Data Reinforced Reasoning for Embodied Planning

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T17:57:33.219109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-14T17:54:50.325820Z digest=sha256:26ff9f2a5260389a62e88f8c084b77b98c45d24d4befdb2c32c51a4374b1845f