Pith. sign in

Paper Citation Record · LEDGER

3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning

As of 8 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 0 inbound Pith citation observations for arXiv:2502.08903.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.08903 v1

Coverage vector

measured 44 of 44 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T23:21:24.869616Z

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

44 of 44 outbound references displayed

  • verified exact1
  • verified fuzzy21
  • unresolved21
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 99e568e2-f83f-431d-ace4-4ad52797d845 · outbound

This paper cites Embodied intelligence toward future smart manufacturing in the era of ai foundation model,.

3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning Embodied intelligence toward future smart manufacturing in the era of ai foundation model,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:21:25.759129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:21:24.633975Z digest=sha256:840b5754e2e53b61d69cf34269c2803abee9a511317939c4fabca93f91084268

Observation 51c0a5fc-055d-4859-a414-5c75c9b916ca · outbound

This paper cites Navigating industry 5.0: A survey of key enabling technologies, trends, challenges, and opportunities,.

3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning Navigating industry 5.0: A survey of key enabling technologies, trends, challenges, and opportunities,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T23:21:24.640204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:21:24.640204Z digest=sha256:d34cc86f69bd86662164cf08b4fd58e4e698fab4e85a347aa19faa0da3774d67

Observation 3465b0df-345b-40e3-aa67-d38d2d3b8a89 · outbound

This paper cites Advanced manufacturing in industry 5.0: A survey of key enabling technologies and future trends,.

3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning Advanced manufacturing in industry 5.0: A survey of key enabling technologies and future trends,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:21:25.733182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:21:24.645694Z digest=sha256:e1bd95f4e1686a9acfe437f43e35ad1bb3bda4e9455754107373d04d818245cc

Observation 3a4efae1-1a48-41ea-977e-ed6aa77142a7 · outbound

This paper cites Large language models for human-robot interaction: A review,.

3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning Large language models for human-robot interaction: A review,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:21:25.716684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:21:24.651712Z digest=sha256:0b9081ad6b02e4672257998795e2f21937382f3554bfa3609f0827827bb8d04e

Observation de8fb097-6b20-4673-a04d-64a6d8db1463 · outbound

This paper cites A survey of optimization-based task and motion planning: From classical to learning approaches,.

3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning A survey of optimization-based task and motion planning: From classical to learning approaches,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T23:21:24.657401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:21:24.657401Z digest=sha256:506fca74e08491a0beac915db23294cf006adbc5ec6b8f1a0776e5cd65a8c8d4

Observation 53084b09-2c62-47f4-9372-1669815f2edb · outbound

This paper cites Computer vision techniques in manufacturing,.

3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning Computer vision techniques in manufacturing,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:21:25.674499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:21:24.669227Z digest=sha256:031bb5119ee0924c4a24d7655667b428759fd20ffdee4c3814597bd7af717970

Observation 203fdb2d-1d6c-4b6d-9f3c-88b732fd39f0 · outbound

This paper cites A comprehen- sive study of 3-d vision-based robot manipulation,.

3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning A comprehen- sive study of 3-d vision-based robot manipulation,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:21:25.658425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:21:24.675049Z digest=sha256:17a5b885bc7a53d015b42ee432df8de5c27e0c78f0ebfc32ad31ae38305b2854

Observation e7c61138-a666-4cd9-a763-4f3da896d222 · outbound

This paper cites Multi-modal 3d object detection in autonomous driving: A survey and taxonomy,.

3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning Multi-modal 3d object detection in autonomous driving: A survey and taxonomy,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:21:25.690542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:21:24.680327Z digest=sha256:971ad5fe6cbd003cb31e255c69ebffe0bcc83c7aaff792da3cb66b5b4fd6f2e4

Observation 05b98e09-6909-461a-9135-c48742e77be7 · outbound

This paper cites Human–robot object handover: Recent progress and future direction,.

3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning Human–robot object handover: Recent progress and future direction,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:21:25.642300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:21:24.685653Z digest=sha256:e99c66716b3aebc63e2182b94a24897183f9782ce4ff2ce71abb4d1e744394fa

Observation c4234f14-6199-41d4-82f7-9f2161be8ee2 · outbound

This paper cites Cmx: Cross-modal fusion for rgb-x semantic segmentation with transformers,.

3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning Cmx: Cross-modal fusion for rgb-x semantic segmentation with transformers,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T23:21:24.690968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:21:24.690968Z digest=sha256:4f3165f3ab5df9935ecc50e6f351edb71b80627941faaf00e63abd2ce26dfc65

Observation 7f89dc9b-5cce-4541-9531-f593264bf14a · outbound

This paper cites Multi-modal feature constraint based tightly coupled monocular visual-lidar odometry and mapping,.

3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning Multi-modal feature constraint based tightly coupled monocular visual-lidar odometry and mapping,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:21:25.615825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:21:24.696367Z digest=sha256:ce6f64478a2aa71d697e541046ac64e81d58b04e6953c650bddfed33dd44d9c2

Observation b6a33347-4d35-4141-b73d-4c17fe331785 · outbound

This paper cites Llm-bt: Performing robotic adaptive tasks based on large language models and behavior trees,.

3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning Llm-bt: Performing robotic adaptive tasks based on large language models and behavior trees,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:21:25.598148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:21:24.701794Z digest=sha256:a2404ee28dc89c2c17640a22102bdff5859a48dd9284c1e833e6504aab29bc37

Observation 2506b0b2-1379-4d9c-8689-affbe82a38dc · outbound

This paper cites Ll3da: Visual interactive instruction tuning for omni-3d understanding, reasoning, and planning,.

3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning Ll3da: Visual interactive instruction tuning for omni-3d understanding, reasoning, and planning,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:21:25.582169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:21:24.706864Z digest=sha256:46fc37ffdc9f20272bc8982d309a7eb3ed400484fc1e7c427c666bad2973c728

Observation be13511d-bc3e-45a3-b473-7bd3ed7a2c0f · outbound

This paper cites Drivegpt4: Interpretable end-to-end autonomous driving via large language model,.

3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning Drivegpt4: Interpretable end-to-end autonomous driving via large language model,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T23:21:24.716904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:21:24.716904Z digest=sha256:657325aa88d82e456f77a254f9b06b5a7752c5a08e38cb46dc3e05dbc14a7da5

Observation c23206f2-6fca-46f8-9667-86428a6b30aa · outbound

This paper cites Progprompt: Generating situated robot task plans using large language models,.

3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning Progprompt: Generating situated robot task plans using large language models,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T23:21:24.721994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:21:24.721994Z digest=sha256:570e4726f22182c6f10651cdaa48294de64e7879ad06ec73a86aed9cb2756412

Observation 8b45d3cb-8d91-4180-a3d5-ed4b7341ddf9 · outbound

This paper cites 3d- llm: injecting the 3d world into large language models,.

3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning 3d- llm: injecting the 3d world into large language models,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:21:25.530677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:21:24.727041Z digest=sha256:164cfaa1945bdf993ecd065a535a929fe25e2f0e4783dfb55d4832b958ede072

Observation 9f79243e-13d3-4c04-bda7-0faf4516e5f4 · outbound

This paper cites LLMI3D: MLLM-based 3D Perception from a Single 2D Image.

3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning LLMI3D: MLLM-based 3D Perception from a Single 2D Image

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T23:21:24.732040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:21:24.732040Z digest=sha256:caee74db38acad347437c72ae893363fd6d51b301b4383e2511ced70f7c6f706

Observation e4303813-9f78-4651-87a1-57c8dd7caa7e · outbound

This paper cites Rdt-1b: a diffusion foundation model for bimanual manipulation,.

3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning Rdt-1b: a diffusion foundation model for bimanual manipulation,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:21:25.513861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:21:24.737749Z digest=sha256:d155fc8eb6d0ef3170dc70f48cca7ef999f566de324722c00f88a7afc6fce5c5

Observation db939c55-5a15-45c2-bea6-2a2e39185b03 · outbound

This paper cites Efficient prompting for llm-based generative internet of things,.

3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning Efficient prompting for llm-based generative internet of things,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T23:21:24.742557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:21:24.742557Z digest=sha256:35242d64069e7291ab5dae0f26f91d874c13657d02572c92c6545211659a0a28

Observation 63d58fa0-3574-4969-ac1b-568fd8927efc · outbound

This paper cites Semantic Abstraction: Open-World 3D Scene Understanding from 2D Vision-Language Models.

3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning Semantic Abstraction: Open-World 3D Scene Understanding from 2D Vision-Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T23:21:24.747622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:21:24.747622Z digest=sha256:3cd4f5761ae82d1bd40612c1b471d25a0b0fb0062befec257c1b8d130f4a6d7b

Observation b3dd1daa-2245-406c-9a63-245a6f9dc0e8 · outbound

This paper cites MoPE: Mixture of Prompt Experts for Parameter-Efficient and Scalable Multimodal Fusion.

3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning MoPE: Mixture of Prompt Experts for Parameter-Efficient and Scalable Multimodal Fusion

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T23:21:24.753064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:21:24.753064Z digest=sha256:757ebc526d31249b291fbd90491009e2ad9186cf5be1751d0481f84c6826d628

Observation 2b21e68d-9dcf-41fb-9b49-debcb5b69290 · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning RT-1: Robotics Transformer for Real-World Control at Scale

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T23:21:24.758518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:21:24.758518Z digest=sha256:20dd02332c32250a5cf6aaa06bb223a2fdc7ca0a9b112d93a763897a4c5ed14c

Observation 2edb81ef-8720-445f-ac1b-a00f0160b2a8 · outbound

This paper cites Rt-2: Vision-language-action models transfer web knowledge to robotic control,.

3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning Rt-2: Vision-language-action models transfer web knowledge to robotic control,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:21:25.487551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:21:24.763904Z digest=sha256:fd7a335d9b11c3e1879c49f71f4447a1daeffbd57213def6af373b69451bd436

Observation 5dac971e-7165-498f-bab5-2a201487cda7 · outbound

This paper cites Vision-Language Foundation Models as Effective Robot Imitators.

3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning Vision-Language Foundation Models as Effective Robot Imitators

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T23:21:24.768834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:21:24.768834Z digest=sha256:698a99f8eec04fd0852be33e1d63d54c2408596d74fc943c7b6cb9db272d65f5

Observation 1be82fd3-1d8a-455e-a8f3-e46619735dc8 · outbound

This paper cites Language-guided Robot Grasping: CLIP-based Referring Grasp Synthesis in Clutter.

3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning Language-guided Robot Grasping: CLIP-based Referring Grasp Synthesis in Clutter

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T23:21:24.774141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:21:24.774141Z digest=sha256:ead58469eec2b6d851d9dfc5438ff2f1e088009b2a9f20b33788892764c3ddbc

Observation 5672150c-e946-4f3c-b69c-2e4663618561 · outbound

This paper cites Clip-fo3d: Learning free open-world 3d scene representations from 2d dense clip,.

3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning Clip-fo3d: Learning free open-world 3d scene representations from 2d dense clip,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:21:25.566103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:21:24.779551Z digest=sha256:7a5dd96be361008e4ace6e376e2144e46971759a82f5d900c8d0aac29f4390ca

Observation d6f29f30-3c78-4bd7-8c59-bf4e3af90faf · outbound

This paper cites Survey on large language model-enhanced reinforce- ment learning: Concept, taxonomy, and methods,.

3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning Survey on large language model-enhanced reinforce- ment learning: Concept, taxonomy, and methods,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:21:25.471120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:21:24.783804Z digest=sha256:a2e608ad5ba5777f11459c5bc8edc67092d68b32bc217ecc066ac16ada8fa349

Observation d0ee76e6-55ff-4f75-bc82-6464aa67469a · outbound

This paper cites To boost zero- shot generalization for embodied reasoning with vision-language pre- training,.

3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning To boost zero- shot generalization for embodied reasoning with vision-language pre- training,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:21:25.453133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:21:24.788247Z digest=sha256:23741c03f7170a8a7bceca30387eeb1a3312a92816b1c31b68a3a8b1bb7087e4

Observation 6bdcfdb4-4f9c-43a8-9975-ecc8eeaabb98 · outbound

This paper cites A survey of visual navigation: From geometry to embodied ai,.

3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning A survey of visual navigation: From geometry to embodied ai,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:21:25.436777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:21:24.792831Z digest=sha256:a3c189b7400b4e9106351b63a144e35b0fcaf857dcab0b459b0f61eae23280c3

Observation c56990b9-2c38-4115-b99a-27b781f9bc51 · outbound

This paper cites Delving into multi-modal multi-task foundation models for road scene understanding: From learning paradigm perspectives,.

3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning Delving into multi-modal multi-task foundation models for road scene understanding: From learning paradigm perspectives,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:21:25.421188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:21:24.797181Z digest=sha256:59a6c754ed2e5ab505fd2c9b82a1116771a905565c55c6f633a3be99f6debbcc

Observation 8e710c2c-f061-4eec-bc6b-dd7432edafe1 · outbound

This paper cites Do As I Can, Not As I Say: Grounding Language in Robotic Affordances.

3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning Do As I Can, Not As I Say: Grounding Language in Robotic Affordances

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T23:21:24.801646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:21:24.801646Z digest=sha256:74ec15d3994319485f98ddedb31543359a636c4c86a6652edf927316c64b9140

Observation ce9e92b8-da09-4c38-9e66-14355fae9f8e · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T23:21:24.806266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:21:24.806266Z digest=sha256:ae2171aa0a468d12f5b46dc8f9ff040916be1223fdcf70b59fbb557f1e0a08bf

Observation 5eef3c2e-3460-4598-bd12-47a07501e139 · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T23:21:24.811203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:21:24.811203Z digest=sha256:6119f1b6d14656f21d8ff75e6da8c8694ebdd359ae603197b5a915688c746610

Observation 31967dc6-4202-4999-9f0f-f97c8f0ba81a · outbound

This paper cites Interactive planning using large language models for partially observable robotic tasks,.

3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning Interactive planning using large language models for partially observable robotic tasks,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:21:25.405220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:21:24.816705Z digest=sha256:4b7974284ede9bd7c1b93e76b64ddecf6bf240f6c01fff6012c91d47d4d80f70

Observation a9d24f4f-73ac-4492-aabf-9e8db7fb4ba3 · outbound

This paper cites InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output.

3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T23:21:24.821792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:21:24.821792Z digest=sha256:d3acd652cacaf7b07ebc7fb9d5f9ff1e0d69b38f7ec30bb875f4308b56fdb0f5

Observation 5561590d-e2a0-4435-958a-18c289f144d5 · outbound

This paper cites MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies.

3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T23:21:24.826991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:21:24.826991Z digest=sha256:0e92fed3c385c5efdcba4acf531ee7769b233f2f41d01fb3fdc59ae9407d9872

Observation bf7997dd-d115-4180-8a32-710dd69bab63 · outbound

This paper cites Visual language maps for robot navigation,.

3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning Visual language maps for robot navigation,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:21:25.389029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:21:24.832463Z digest=sha256:93af487fa4f92d756075325d124478eff8ac51f19d5840bd79398529d69a2861

Observation 2aed1b5c-b8c2-48d6-948e-8cd12ab6780b · outbound

This paper cites Ving: Learning open-world navigation with visual goals,.

3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning Ving: Learning open-world navigation with visual goals,

Reference 40

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-07T23:21:25.177254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:21:24.837326Z digest=sha256:9b72629dbbabb4a856f19e3d2082d896a2baefd3588d25a5602ee61386c97316

Observation ce7cb4ab-9fb5-4273-b845-f81d216c5e45 · outbound

This paper cites PaLM-E: An Embodied Multimodal Language Model.

3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning PaLM-E: An Embodied Multimodal Language Model

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T23:21:24.842290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:21:24.842290Z digest=sha256:a9db25e5336e178cd826221f6919261b48cfe117845512b6607ac34ba2970d74

Observation 53a05805-5044-4462-ab8f-e913e47f8993 · outbound

This paper cites Lora: Low-rank adaptation of large language models,.

3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning Lora: Low-rank adaptation of large language models,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T23:21:24.847764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:21:24.847764Z digest=sha256:df4262bdf0bb98f505742d5af411b0f10f2520e433b7b768583834e12a87d14b

Observation 87ac9c5b-7f41-4d79-9682-7f5f6d967e21 · outbound

This paper cites Patchwork++: Fast and robust ground segmentation solving partial under-segmentation using 3d point cloud,.

3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning Patchwork++: Fast and robust ground segmentation solving partial under-segmentation using 3d point cloud,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:21:25.362772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:21:24.858979Z digest=sha256:51a663f310dd63240e302f2c9fd956ee8371b59a26820af592d36de50e4b629a

Observation ef31a0f3-933b-443c-8e1d-c271e5f3973a · outbound

This paper cites BridgeData V2: A Dataset for Robot Learning at Scale.

3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning BridgeData V2: A Dataset for Robot Learning at Scale

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T23:21:24.869616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:21:24.869616Z digest=sha256:b2794185cf164e9bcf692dd8e3e225eab4d21463b9f1401d86e2828cd0037b36

Observation 5f592ef7-a54d-4ce8-9b5a-dab006be4acf · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning LoRA: Low-Rank Adaptation of Large Language Models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T23:21:24.853042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:21:24.853042Z digest=sha256:7543ffe94dd0fe4712410b0439d28ccfe345051abff62a4465d51fe87df2f80d

Observation bc592cf2-c519-4d80-89cd-abfd98a46259 · outbound

This paper cites Patchwork++: Fast and Robust Ground Segmentation Solving Partial Under-Segmentation Using 3D Point Cloud.

3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning Patchwork++: Fast and Robust Ground Segmentation Solving Partial Under-Segmentation Using 3D Point Cloud

Reference 2022

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T23:21:24.932684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:21:24.864173Z digest=sha256:365e4d3b42bdbd79f326c4aaffdf57990efa1b11e56e78e1e1d99362b635ba11

Pith citing papers

No inbound Pith citation observations are available.