Pith. sign in

Paper Citation Record · LEDGER

Instruct2Act: Mapping Multi-modality Instructions to Robotic Actions with Large Language Model

As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 39 inbound Pith citation observations for arXiv:2305.11176.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2305.11176 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 39 of 39 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 39 of 39 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:28:36.947983Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T07:26:54.507995Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 14ef7f5e-d878-4474-bbb2-24627362f663 · inbound

ConfusionPrompt: Practical Private Inference for Online Large Language Models cites this paper.

ConfusionPrompt: Practical Private Inference for Online Large Language Models Instruct2Act: Mapping Multi-modality Instructions to Robotic Actions with Large Language Model

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-24T04:58:54.849006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-24T04:57:55.197897Z digest=sha256:31793f3117a11b0f1bc1e0458403949f1ac488dbcc7a2a2935be1b8211509f59

Observation 28ec5370-4b3d-4212-9945-02add761aeb4 · inbound

A Survey on Vision-Language-Action Models for Embodied AI cites this paper.

A Survey on Vision-Language-Action Models for Embodied AI Instruct2Act: Mapping Multi-modality Instructions to Robotic Actions with Large Language Model

Reference 107

Resolution
verified exact
arxiv_id, observed 2026-05-24T01:25:54.588575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-24T01:25:10.150459Z digest=sha256:c25b52ad203539e8e82e82c186af80e107215240a944514d75d555b306d046be

Observation ce7f2a23-d673-4b4f-bb93-ee81d151b671 · inbound

Generative Timelines for Instructed Visual Assembly cites this paper.

Generative Timelines for Instructed Visual Assembly Instruct2Act: Mapping Multi-modality Instructions to Robotic Actions with Large Language Model

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T17:47:39.756321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:47:39.756321Z digest=sha256:a81e2d763a2eb6d4504ca5dded6da5d841c7838fb682651cf4943f6aa3c5a48a

Observation 4c4088ad-c387-4a9e-86b6-dd42d59b8338 · inbound

OV-HHIR: Open Vocabulary Human Interaction Recognition Using Cross-modal Integration of Large Language Models cites this paper.

OV-HHIR: Open Vocabulary Human Interaction Recognition Using Cross-modal Integration of Large Language Models Instruct2Act: Mapping Multi-modality Instructions to Robotic Actions with Large Language Model

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T22:56:23.625900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:56:23.625900Z digest=sha256:989eb760bf028e9a333f9659e65fb80e374bd3aab5c2b7ade33fa82da3d3d0cb

Observation 03618dc2-aa94-422a-b006-a9bdd824d7e4 · inbound

Visual Large Language Models for Generalized and Specialized Applications cites this paper.

Visual Large Language Models for Generalized and Specialized Applications Instruct2Act: Mapping Multi-modality Instructions to Robotic Actions with Large Language Model

Reference 184

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:09.612514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:09.612514Z digest=sha256:441c9b41382f070c645e8d144e0274d943bf5e4bb62b96b62e2e56fd539a0ce4

Observation fd0b7d25-7d31-4ed0-a9f1-2f261b8aa9c3 · inbound

Large language models for artificial general intelligence (AGI): A survey of foundational principles and approaches cites this paper.

Large language models for artificial general intelligence (AGI): A survey of foundational principles and approaches Instruct2Act: Mapping Multi-modality Instructions to Robotic Actions with Large Language Model

Reference 258

Resolution
unresolved
no resolver link, observed 2026-08-10T21:56:13.044906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:56:13.044906Z digest=sha256:985b59f89ef530c780b2d31aedfef10a5fdd4259b025d0e09ef32a264c85d7fc

Observation f68a8b5c-2e1c-4b8d-aeba-a6f229672c15 · inbound

Robotic Programmer: Video Instructed Policy Code Generation for Robotic Manipulation cites this paper.

Robotic Programmer: Video Instructed Policy Code Generation for Robotic Manipulation Instruct2Act: Mapping Multi-modality Instructions to Robotic Actions with Large Language Model

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T21:41:29.048861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:41:29.048861Z digest=sha256:466f07dcb20361e49799066d82f77956d7c3396a431aa639c3d8f6ba557404f0

Observation 89f0551f-a3db-4646-a5b3-761676c6817c · inbound

Integrating LMM Planners and 3D Skill Policies for Generalizable Manipulation cites this paper.

Integrating LMM Planners and 3D Skill Policies for Generalizable Manipulation Instruct2Act: Mapping Multi-modality Instructions to Robotic Actions with Large Language Model

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-09T22:44:50.810323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T22:44:50.810323Z digest=sha256:c2f8c61f1c4f0f47927fd8d77db37c29abe8cd0000265d84bfb650fdf37a07f6

Observation 4373253c-8b85-4c23-9d85-3f86929dab2c · inbound

A Real-to-Sim-to-Real Approach to Robotic Manipulation with VLM-Generated Iterative Keypoint Rewards cites this paper.

A Real-to-Sim-to-Real Approach to Robotic Manipulation with VLM-Generated Iterative Keypoint Rewards Instruct2Act: Mapping Multi-modality Instructions to Robotic Actions with Large Language Model

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T00:03:16.363277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:03:16.363277Z digest=sha256:242a29f239cc9d42693cb8711302b3c442b95e7a326d593d48a479bc4dffb7d6

Observation 89abcfe5-4980-4830-a70f-a54d6375978a · inbound

Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models cites this paper.

Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models Instruct2Act: Mapping Multi-modality Instructions to Robotic Actions with Large Language Model

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T23:18:39.880512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:18:39.880512Z digest=sha256:ee38b29d52b3369112165ebcc8989f7f2420b2a72c0b5fce980a60a5464ee460

Observation b5b6eed3-e314-4976-b558-6988bd8ca556 · inbound

Few-Shot Vision-Language Action-Incremental Policy Learning cites this paper.

Few-Shot Vision-Language Action-Incremental Policy Learning Instruct2Act: Mapping Multi-modality Instructions to Robotic Actions with Large Language Model

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T11:28:36.947983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:28:36.947983Z digest=sha256:c01cf57d85bc89d744b0f979a506e4ba33b7cdd0d849a1f053dd2b7b21d42a6a

Observation b763b025-0969-43c8-a389-4c3f5372e3d4 · inbound

Advancing Embodied Agent Security: From Safety Benchmarks to Input Moderation cites this paper.

Advancing Embodied Agent Security: From Safety Benchmarks to Input Moderation Instruct2Act: Mapping Multi-modality Instructions to Robotic Actions with Large Language Model

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-16T11:24:11.651441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:24:11.651441Z digest=sha256:7b55fbe952f0575a914f5813bbbd57f75d711cb41e1a6088281a62393f8f8580

Observation 8740feef-7995-4511-b66c-d275ba4b8552 · inbound

CoordField: Coordination Field for Agentic UAV Task Allocation In Low-altitude Urban Scenarios cites this paper.

CoordField: Coordination Field for Agentic UAV Task Allocation In Low-altitude Urban Scenarios Instruct2Act: Mapping Multi-modality Instructions to Robotic Actions with Large Language Model

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T04:57:48.343486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:57:48.343486Z digest=sha256:4a8ad717234730e308132c764dcd97d34abb4a871f1a94682c49d2160561c16f

Observation 49fde8af-f6c3-43c4-9d94-63a96b1a5f8e · inbound

Semantic Intelligence: Integrating GPT-4 with A Planning in Low-Cost Robotics cites this paper.

Semantic Intelligence: Integrating GPT-4 with A Planning in Low-Cost Robotics Instruct2Act: Mapping Multi-modality Instructions to Robotic Actions with Large Language Model

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T04:10:41.631892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:10:41.631892Z digest=sha256:f6df96cc02322592b215390cf9b380df1583434ff41f4d05087aa9773ecaa7a0

Observation 8ca5bd2e-b098-457d-9d5d-a490089a89d2 · inbound

Conversational Interfaces for Parametric Conceptual Architectural Design: Integrating Mixed Reality with LLM-driven Interaction cites this paper.

Conversational Interfaces for Parametric Conceptual Architectural Design: Integrating Mixed Reality with LLM-driven Interaction Instruct2Act: Mapping Multi-modality Instructions to Robotic Actions with Large Language Model

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:23.742554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:03:23.742554Z digest=sha256:5c6a452f26b75d5a5bbe5cbbb793f17a6c53a410529c5c6f1e1c6eab421fe992

Observation c0190b68-6726-4d51-925c-eb772a4a5cb2 · inbound

ACTLLM: Action Consistency Tuned Large Language Model cites this paper.

ACTLLM: Action Consistency Tuned Large Language Model Instruct2Act: Mapping Multi-modality Instructions to Robotic Actions with Large Language Model

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T22:34:06.759535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:34:06.759535Z digest=sha256:c343a45584000118faf3b2cbeefd3e3d761319a4e32dc5f0030120f43f7d705b

Observation 41837580-5943-489d-86cf-d0ebd4a8e2da · inbound

A Survey on Vision-Language-Action Models: An Action Tokenization Perspective cites this paper.

A Survey on Vision-Language-Action Models: An Action Tokenization Perspective Instruct2Act: Mapping Multi-modality Instructions to Robotic Actions with Large Language Model

Reference 129

Resolution
verified exact
arxiv_id, observed 2026-05-17T14:08:35.437833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-17T14:08:34.893876Z digest=sha256:8ffc74c04030bc98890f561e5d976a0cd17e16bee2147025669bc343b69f7ed5

Observation f149ce8b-d424-4047-8791-14a0fd8573fe · inbound

Foundation Model Driven Robotics: A Comprehensive Review cites this paper.

Foundation Model Driven Robotics: A Comprehensive Review Instruct2Act: Mapping Multi-modality Instructions to Robotic Actions with Large Language Model

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T17:43:53.122220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:43:53.122220Z digest=sha256:80b2530a011ab905edc7355e03d743fe0dbde850cd5d5d5ab0bb65bd18353e65

Observation 8bba5e72-74ee-4bf1-aed0-22d240228810 · inbound

SDD: Self-Degraded Defense against Malicious Fine-tuning cites this paper.

SDD: Self-Degraded Defense against Malicious Fine-tuning Instruct2Act: Mapping Multi-modality Instructions to Robotic Actions with Large Language Model

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T17:55:13.818599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:55:13.818599Z digest=sha256:48ca4fe3ed83c95f60fbaa0b1bef9f5dfc496c01a4933711db123dea9c36018d

Observation fc477058-e9f4-4f03-aa57-ac6113c7ec82 · inbound

Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation cites this paper.

Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation Instruct2Act: Mapping Multi-modality Instructions to Robotic Actions with Large Language Model

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-15T21:28:42.012107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-15T21:28:41.904725Z digest=sha256:1ca5a8678beb75f7efd5dc4eb4368bd5dde1adda54217bbd9833b94cfd1707a6

Observation ba2340f1-f0a0-4a76-8b2a-df76d4981e15 · inbound

FaceAnonyMixer: Cancelable Faces via Identity Consistent Latent Space Mixing cites this paper.

FaceAnonyMixer: Cancelable Faces via Identity Consistent Latent Space Mixing Instruct2Act: Mapping Multi-modality Instructions to Robotic Actions with Large Language Model

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:30.341166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:30.341166Z digest=sha256:a4625cacc576a44eb2616f91df758b4b98f1f55145d78439163e0f70e3a29562

Observation 50b1e20c-be23-4d8f-95b1-8a6def55c66b · inbound

Large Model Empowered Embodied AI: A Survey on Decision-Making and Embodied Learning cites this paper.

Large Model Empowered Embodied AI: A Survey on Decision-Making and Embodied Learning Instruct2Act: Mapping Multi-modality Instructions to Robotic Actions with Large Language Model

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:42.886164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:42.886164Z digest=sha256:733c937593bf228921a97072e1cb7c814741b3e3aa011b4e56be553a6ec08533

Observation 439c4702-a860-4347-a2b3-4074bbf90a83 · inbound

Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey cites this paper.

Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey Instruct2Act: Mapping Multi-modality Instructions to Robotic Actions with Large Language Model

Reference 159

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:28:16.007179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-17T20:28:15.818016Z digest=sha256:aa86845780e7aa90a37be648879c701327beb22599ca553052dce2a253ae65f6

Observation ed0b6296-7bb2-4a49-8dcb-3f090240dd80 · inbound

MARS: Multi-Agent Robotic System with Multimodal Large Language Models for Assistive Intelligence cites this paper.

MARS: Multi-Agent Robotic System with Multimodal Large Language Models for Assistive Intelligence Instruct2Act: Mapping Multi-modality Instructions to Robotic Actions with Large Language Model

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T01:10:34.249058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:05:47.657235Z digest=sha256:690b92c199a378e28f7449a4435766f91e94f4198641d4aa7c1ecadbca336c7f

Observation ff20d6fc-18c7-4d48-bbdb-df59b914bd95 · inbound

ImagineNav++: Prompting Vision-Language Models as Embodied Navigator through Scene Imagination cites this paper.

ImagineNav++: Prompting Vision-Language Models as Embodied Navigator through Scene Imagination Instruct2Act: Mapping Multi-modality Instructions to Robotic Actions with Large Language Model

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-16T21:08:32.486743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-16T21:07:54.497999Z digest=sha256:347e96a87828e3f72454b27b123413e29f8acf131c9a207844834202483f10f4

Observation 288dafc4-074a-4583-abe1-aaee24236c04 · inbound

A Survey on Evaluating Quality and Trustworthiness in LLM-Generated Data cites this paper.

A Survey on Evaluating Quality and Trustworthiness in LLM-Generated Data Instruct2Act: Mapping Multi-modality Instructions to Robotic Actions with Large Language Model

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-03T08:15:20.295381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T08:15:20.295381Z digest=sha256:6af7d3343b4c33bd641036e44ee949b8cb4a86308783693e39d0e3596d447a50

Observation 995f3c54-2c1d-45fa-a8e2-970b786bc6ab · inbound

Evolvable Embodied Agent for Robotic Manipulation via Long Short-Term Reflection and Optimization cites this paper.

Evolvable Embodied Agent for Robotic Manipulation via Long Short-Term Reflection and Optimization Instruct2Act: Mapping Multi-modality Instructions to Robotic Actions with Large Language Model

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:05:24.469130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T13:04:48.932945Z digest=sha256:ce89a39adaea82f0d392e70b4c4814903ba09079c6146229212c3175cd3989b8

Observation be0a79a6-6112-47a0-aa88-ecbeb602750f · inbound

SharedRequest: Privacy-Preserving Model-Agnostic Inference for Large Language Models cites this paper.

SharedRequest: Privacy-Preserving Model-Agnostic Inference for Large Language Models Instruct2Act: Mapping Multi-modality Instructions to Robotic Actions with Large Language Model

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T09:26:51.544337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-28T05:29:47.894226Z digest=sha256:54eddf4a291ed36912bfb71e06a04d3edd9649793865e793e706c8b85c99b35a

Observation 9809110e-0ecf-49d2-a9b3-76dfeb8145ba · inbound

Efficient Skill Grounding via Code Refactoring with Small Language Models cites this paper.

Efficient Skill Grounding via Code Refactoring with Small Language Models Instruct2Act: Mapping Multi-modality Instructions to Robotic Actions with Large Language Model

Reference 78

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T21:07:24.153749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-27T19:55:20.212198Z digest=sha256:67bf1150a7be37c514661c889447d35310ea62177455affd5f6c634e90dff82a

Observation 0e5fb998-6181-44f1-8c8e-c8d955e8c208 · inbound

Functional Cache Grafting: Robust and Rapid Code-Policy Synthesis for Embodied Agents cites this paper.

Functional Cache Grafting: Robust and Rapid Code-Policy Synthesis for Embodied Agents Instruct2Act: Mapping Multi-modality Instructions to Robotic Actions with Large Language Model

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-06-27T05:30:35.769413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-27T05:26:48.431739Z digest=sha256:61d03000b861d65207969ca6a0f64a73b591ab7f38a944026a47d0892d3a4244

Observation 34c3500e-747b-4446-a4f6-566ad15a71d2 · inbound

Functional Cache Grafting: Robust and Rapid Code-Policy Synthesis for Embodied Agents cites this paper.

Functional Cache Grafting: Robust and Rapid Code-Policy Synthesis for Embodied Agents Instruct2Act: Mapping Multi-modality Instructions to Robotic Actions with Large Language Model

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T11:44:03.545449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:44:03.545449Z digest=sha256:e4b1bcc9da0c408df17dab32af79deebc2f0ce4bc637aa62731895e28ba261c0

Observation f29e2a24-7b6b-4337-81ef-652bb752d701 · inbound

Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation cites this paper.

Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation Instruct2Act: Mapping Multi-modality Instructions to Robotic Actions with Large Language Model

Reference 230

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T17:18:43.761394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-27T04:19:26.332718Z digest=sha256:23807305722fcf81cf9cab36884a4b2eab9551da778a59f8111768b0584da125

Observation 1645445c-a951-4492-a399-749d9d4b7ff7 · inbound

Harness VLA: Steering Frozen VLAs into Reliable Manipulation Primitives via Memory-Guided Agents cites this paper.

Harness VLA: Steering Frozen VLAs into Reliable Manipulation Primitives via Memory-Guided Agents Instruct2Act: Mapping Multi-modality Instructions to Robotic Actions with Large Language Model

Reference 59

Resolution
verified exact
local_arxiv, observed 2026-07-10T07:26:54.509395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-10T07:18:57.823444Z digest=sha256:156f3b5ea86bc55bf133bac4c4efdaa09c5c31c32cec40de94db6bbb8f166000

Observation a68f4efd-fefc-40af-aaf3-647509499690 · inbound

Harness VLA: Steering Frozen VLAs into Reliable Manipulation Primitives via Memory-Guided Agents cites this paper.

Harness VLA: Steering Frozen VLAs into Reliable Manipulation Primitives via Memory-Guided Agents Instruct2Act: Mapping Multi-modality Instructions to Robotic Actions with Large Language Model

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-02T07:56:56.620153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:56:56.620153Z digest=sha256:81ecc6d3c3185ff7ece46090a50bd9653dc22f44c62029851160043d0d15037c

Observation c237d157-fea3-4b06-8ba0-d60383f9904d · inbound

Qwen-Audio-VAE Technical Report cites this paper.

Qwen-Audio-VAE Technical Report Instruct2Act: Mapping Multi-modality Instructions to Robotic Actions with Large Language Model

Reference 237

Resolution
unresolved
no resolver link, observed 2026-07-14T03:31:19.309532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T03:31:19.309532Z digest=sha256:fa3bc5c2db0454e2929527741cbdf14604a5ad8b0b263dd7d22c0ef25e43f7b1

Observation 4c58c8a4-f760-43e9-8d79-5fa90acfec7c · inbound

Self-Evolving Just-In-Time Memory for Proactive Embodied Safety cites this paper.

Self-Evolving Just-In-Time Memory for Proactive Embodied Safety Instruct2Act: Mapping Multi-modality Instructions to Robotic Actions with Large Language Model

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T09:54:52.432986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:54:52.432986Z digest=sha256:7464c6eb8ce5042d17b2ed47290a5d22c87b007314855e0afa672085af193391

Observation e3d210c8-b04d-4049-83ea-fe0b50d20bb5 · inbound

Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills cites this paper.

Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills Instruct2Act: Mapping Multi-modality Instructions to Robotic Actions with Large Language Model

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-04T19:45:33.996455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:45:33.996455Z digest=sha256:6192d74d2c818a74471a7e9b8f3e23a4bebae9c4c4933bc6d5f069e8191f8784

Observation 766e97bd-0a28-4853-bb37-341cae407934 · inbound

Is Inter-Seed Cross-Play Enough? Evaluating the Robustness of Zero-Shot Coordination Algorithms to Implementation Details cites this paper.

Is Inter-Seed Cross-Play Enough? Evaluating the Robustness of Zero-Shot Coordination Algorithms to Implementation Details Instruct2Act: Mapping Multi-modality Instructions to Robotic Actions with Large Language Model

Reference 220

Resolution
unresolved
no resolver link, observed 2026-08-05T15:25:40.330321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:25:40.330321Z digest=sha256:ad1e8fb38fef34f4216ca8f9bf48d4e5d8dad4dceb040409b3828a1a77bbac21

Observation 54530e77-b581-4687-a48e-ad8354d3d278 · inbound

ETA: A New Agentic Paradigm for Embodied Tasks cites this paper.

ETA: A New Agentic Paradigm for Embodied Tasks Instruct2Act: Mapping Multi-modality Instructions to Robotic Actions with Large Language Model

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T05:44:44.230086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:44:44.230086Z digest=sha256:165b491c7e603b4b6aab68bee91528235ead8f1c4f85b1cef367d9269c886429