Pith. sign in

Paper Citation Record · LEDGER

In-Context VLA: Endowing Vision-Language-Action Models with Language via In-Context Post-Training and Agentic Tool Use

As of 8 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 0 inbound Pith citation observations for arXiv:2608.05738.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.05738 v1

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T00:30:03.258938Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

46 of 46 outbound references displayed

  • verified exact0
  • verified fuzzy9
  • unresolved37
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f7eb6eb6-0e15-47ef-8527-cdc713ce9b0c · outbound

This paper cites IEEE Transactions on Neural Networks and Learning Systems , year=.

In-Context VLA: Endowing Vision-Language-Action Models with Language via In-Context Post-Training and Agentic Tool Use IEEE Transactions on Neural Networks and Learning Systems , year=

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:30:04.570439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T00:30:03.006647Z digest=sha256:f8a83162b060f92844652074eedc82abd5f1daaaee52a3aeabe244900b745d68

Observation 987ca5b0-57c6-471e-b0e4-7f1e53de1289 · outbound

This paper cites JoyAI-RA 0.1: A Foundation Model for Robotic Autonomy.

In-Context VLA: Endowing Vision-Language-Action Models with Language via In-Context Post-Training and Agentic Tool Use JoyAI-RA 0.1: A Foundation Model for Robotic Autonomy

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T00:30:03.014525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:30:03.014525Z digest=sha256:adf3f6b330a63e1f3fa1a45b668ce7b290845aa2aa8c3d22f1f5bda8c3e5c22a

Observation fa45f4ed-4fb9-453b-b8b6-afb419116e19 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

In-Context VLA: Endowing Vision-Language-Action Models with Language via In-Context Post-Training and Agentic Tool Use OpenVLA: An Open-Source Vision-Language-Action Model

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T00:30:03.020883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:30:03.020883Z digest=sha256:ed9a55c0d7c4153fd446b7232c9c0e4d76a23cfffbcf2601554c7b4c9ae80489

Observation 7082172c-ce99-4768-a7f8-66ccfcb83440 · outbound

This paper cites Proceedings of the 60th annual meeting of the association for computational linguistics (volume 1: long papers) , pages=.

In-Context VLA: Endowing Vision-Language-Action Models with Language via In-Context Post-Training and Agentic Tool Use Proceedings of the 60th annual meeting of the association for computational linguistics (volume 1: long papers) , pages=

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T00:30:03.028302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:30:03.028302Z digest=sha256:d3907fca004ae062a9cb7247070d56df0d6c587e776a646e723cbad31c0c8020

Observation 4ca46f22-ed91-4dab-a9d0-b77ea41e4422 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

In-Context VLA: Endowing Vision-Language-Action Models with Language via In-Context Post-Training and Agentic Tool Use Advances in Neural Information Processing Systems , volume=

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T00:30:03.034197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:30:03.034197Z digest=sha256:4f184e12d440aa8b9f60f4a5e12de67863b8c32fe46091c764fdc86d9830fb2b

Observation e6bbec38-f956-48dc-b36b-dec14083be3a · outbound

This paper cites Do As I Can, Not As I Say: Grounding Language in Robotic Affordances.

In-Context VLA: Endowing Vision-Language-Action Models with Language via In-Context Post-Training and Agentic Tool Use Do As I Can, Not As I Say: Grounding Language in Robotic Affordances

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T00:30:03.040411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:30:03.040411Z digest=sha256:7bdcad0a33b3bf055fcf425011607be27fdddfb56123f36cdf18afbdb4deff3e

Observation 183a0138-3ca1-4870-b752-101aba51a2c4 · outbound

This paper cites 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pages=.

In-Context VLA: Endowing Vision-Language-Action Models with Language via In-Context Post-Training and Agentic Tool Use 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pages=

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T00:30:03.047791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:30:03.047791Z digest=sha256:2ed06b0ec87e197ebe43c59879331e33d2197fd99005f5bc41794e500c4cbcfa

Observation 2ecdc0a2-7efb-4598-a93b-4def83e45495 · outbound

This paper cites VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection.

In-Context VLA: Endowing Vision-Language-Action Models with Language via In-Context Post-Training and Agentic Tool Use VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T00:30:03.053598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:30:03.053598Z digest=sha256:f82d6536dbc6c500fb2265c5f79c58ecf5046397c424c767b11e37995de16583

Observation 939750d8-1c9e-4f39-83ce-3080f9eb4076 · outbound

This paper cites NeurIPS 2025 Workshop on Space in Vision, Language, and Embodied AI , year=.

In-Context VLA: Endowing Vision-Language-Action Models with Language via In-Context Post-Training and Agentic Tool Use NeurIPS 2025 Workshop on Space in Vision, Language, and Embodied AI , year=

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:30:04.516393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T00:30:03.059312Z digest=sha256:8a0f7e704dc1c266fe4615ae0b7da18863611e2a6b26f5f384f5df40e055700e

Observation 5812d0e6-efe0-49fa-8938-448723d2313a · outbound

This paper cites 9th Annual Conference on Robot Learning , year=.

In-Context VLA: Endowing Vision-Language-Action Models with Language via In-Context Post-Training and Agentic Tool Use 9th Annual Conference on Robot Learning , year=

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:30:04.498202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T00:30:03.064459Z digest=sha256:daaf566ddb2c8198a210be4de479efb2dda81cc1a1b7f98af0ab1cc0296265de

Observation fd67b1d6-3ee6-447b-ae0b-1e4aeee59680 · outbound

This paper cites MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation.

In-Context VLA: Endowing Vision-Language-Action Models with Language via In-Context Post-Training and Agentic Tool Use MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T00:30:03.069585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:30:03.069585Z digest=sha256:2f3f8bedacc7647381d33ee1c0144a325068ea233ad6342ab8d6672910d54feb

Observation d4782cff-fe70-40e2-962c-80cd24031efb · outbound

This paper cites European Conference on Computer Vision , pages=.

In-Context VLA: Endowing Vision-Language-Action Models with Language via In-Context Post-Training and Agentic Tool Use European Conference on Computer Vision , pages=

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:30:04.479702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T00:30:03.075063Z digest=sha256:6cd2544aa545878e789d0352ded38737c9149e53530bac3103b3857456808972

Observation 36f58bce-f1f6-4c18-adfe-381d103a2ba4 · outbound

This paper cites PaLM-E: An Embodied Multimodal Language Model.

In-Context VLA: Endowing Vision-Language-Action Models with Language via In-Context Post-Training and Agentic Tool Use PaLM-E: An Embodied Multimodal Language Model

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T00:30:03.081177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:30:03.081177Z digest=sha256:69ee5984f7903fb4bbe19c51ba520ce1926809ac3d7d48a3566ff88df4ccf30a

Observation d78a3072-537f-4a59-9a36-aa028beb1cca · outbound

This paper cites VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models.

In-Context VLA: Endowing Vision-Language-Action Models with Language via In-Context Post-Training and Agentic Tool Use VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T00:30:03.086780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:30:03.086780Z digest=sha256:2c8fcac25a9d7208a9a339c6c501a90410e543f46bd59f7ede56d894949a6fb1

Observation 4170b60c-f2a6-4a5a-93d4-35a4a86896e0 · outbound

This paper cites Inner Monologue: Embodied Reasoning through Planning with Language Models.

In-Context VLA: Endowing Vision-Language-Action Models with Language via In-Context Post-Training and Agentic Tool Use Inner Monologue: Embodied Reasoning through Planning with Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T00:30:03.092074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:30:03.092074Z digest=sha256:2398b0e5c7dd3657eab4aee3f7cb32773ad0d690c8344c43a389fd330453df87

Observation 2b43020b-bb89-4179-b01b-d18a17b3eb1f · outbound

This paper cites SpaceTools: Tool-Augmented Spatial Reasoning via Double Interactive RL.

In-Context VLA: Endowing Vision-Language-Action Models with Language via In-Context Post-Training and Agentic Tool Use SpaceTools: Tool-Augmented Spatial Reasoning via Double Interactive RL

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T00:30:03.097660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:30:03.097660Z digest=sha256:78742db948bc8ac52144419939c87ddab7b3c7dd001054bf7506113ac8424c0d

Observation 1b5ba0d2-d480-4540-a2ae-53b7041dfae8 · outbound

This paper cites Towards Long-horizon Embodied Agents with Tool-Aligned Vision-Language-Action Models.

In-Context VLA: Endowing Vision-Language-Action Models with Language via In-Context Post-Training and Agentic Tool Use Towards Long-horizon Embodied Agents with Tool-Aligned Vision-Language-Action Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T00:30:03.103238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:30:03.103238Z digest=sha256:4c6e7f32a63f2145f3d27d03fced20140eec11e9c3ee0d9814ea1f8863bbdc03

Observation bca905dc-e752-4eab-93fa-e195838f3fb6 · outbound

This paper cites Latent Reasoning VLA: Latent Thinking and Prediction for Vision-Language-Action Models.

In-Context VLA: Endowing Vision-Language-Action Models with Language via In-Context Post-Training and Agentic Tool Use Latent Reasoning VLA: Latent Thinking and Prediction for Vision-Language-Action Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T00:30:03.108171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:30:03.108171Z digest=sha256:717bf6991afe95fdcd705138ff371263785c7b1d9b564d12f04fd1635cd1393d

Observation c23256b3-56ad-49d2-be2e-0c730d7a4f1a · outbound

This paper cites Proceedings of the AAAI Conference on Artificial Intelligence , volume=.

In-Context VLA: Endowing Vision-Language-Action Models with Language via In-Context Post-Training and Agentic Tool Use Proceedings of the AAAI Conference on Artificial Intelligence , volume=

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:30:04.461262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T00:30:03.113661Z digest=sha256:c005154e582e0471b1bef5b3f908f68c42ca1c99d97af7a9d797e971d98eedee

Observation aac46958-56d6-4430-a0de-e5471dff2a8d · outbound

This paper cites arXiv preprint arXiv:2601.11404 , year=.

In-Context VLA: Endowing Vision-Language-Action Models with Language via In-Context Post-Training and Agentic Tool Use arXiv preprint arXiv:2601.11404 , year=

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T00:30:03.118707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:30:03.118707Z digest=sha256:52cfce7527a6bbc821c2714d8b181227d8efdecf5db5b340f49d92a47bed47b0

Observation b73c236e-f001-4b65-a4e7-d42c3af9ab61 · outbound

This paper cites arXiv preprint arXiv:2603.22280 , year=.

In-Context VLA: Endowing Vision-Language-Action Models with Language via In-Context Post-Training and Agentic Tool Use arXiv preprint arXiv:2603.22280 , year=

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T00:30:03.123624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:30:03.123624Z digest=sha256:d5d0a90691302969f0c68bce1aa58c342eaef8e38941cfcb9ff96158d1de211f

Observation 26b69c93-ab4b-4f41-87db-c949e7e991b6 · outbound

This paper cites arXiv preprint arXiv:2603.14523 , year=.

In-Context VLA: Endowing Vision-Language-Action Models with Language via In-Context Post-Training and Agentic Tool Use arXiv preprint arXiv:2603.14523 , year=

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T00:30:03.127765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:30:03.127765Z digest=sha256:a0c76cf193b764716d5a35ad08284344383d22a2361edcc1035a2907d6985a51

Observation f3b75b19-ac0b-4eb3-b5f6-da0aac8eba28 · outbound

This paper cites an unresolved cited work.

In-Context VLA: Endowing Vision-Language-Action Models with Language via In-Context Post-Training and Agentic Tool Use Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T00:30:03.132899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:30:03.132899Z digest=sha256:77cc4901c0008ab7050061d34c514ba298bff4771d43e6be42d684582a28147d

Observation 7c618703-eda2-49e8-b18e-84832367434d · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

In-Context VLA: Endowing Vision-Language-Action Models with Language via In-Context Post-Training and Agentic Tool Use $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T00:30:03.137263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:30:03.137263Z digest=sha256:cd25a5ed86015acaa02239b56b0183a874e7e3f0840acdeec3f23f0b295df92a

Observation d394e4ff-3465-491a-8843-6d3bfcb0ddf8 · outbound

This paper cites InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy.

In-Context VLA: Endowing Vision-Language-Action Models with Language via In-Context Post-Training and Agentic Tool Use InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T00:30:03.142726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:30:03.142726Z digest=sha256:60921b99955a5d2756dc3c3b44a1aa710035b79a817ae87e80d369292daea6d7

Observation b95ac65d-b420-4422-8ce7-55c7859b323a · outbound

This paper cites Forty-third International Conference on Machine Learning , year=.

In-Context VLA: Endowing Vision-Language-Action Models with Language via In-Context Post-Training and Agentic Tool Use Forty-third International Conference on Machine Learning , year=

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:30:04.433143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T00:30:03.147823Z digest=sha256:d011ee73586b28ca1ac60803760668d6acd14b086bf869b8620f5cdd99454c2a

Observation 37f111f6-d03b-47dc-8e04-8cffaddf19ad · outbound

This paper cites arXiv preprint arXiv:2602.10098 , year=.

In-Context VLA: Endowing Vision-Language-Action Models with Language via In-Context Post-Training and Agentic Tool Use arXiv preprint arXiv:2602.10098 , year=

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T00:30:03.153249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:30:03.153249Z digest=sha256:cef888419c21d1e51d51e79cc517852e570bb93ca2e3710a1d27cf8bd2c7ff08

Observation 0bb6acc2-4fcb-4a27-b008-8d0f96ce7529 · outbound

This paper cites Advances in neural information processing systems , volume=.

In-Context VLA: Endowing Vision-Language-Action Models with Language via In-Context Post-Training and Agentic Tool Use Advances in neural information processing systems , volume=

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:30:04.417001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T00:30:03.158542Z digest=sha256:dd5643381605ae8df0b3f163b0b2271ce8050e93325f510a2df069a5d649d655

Observation 9605a5a1-b880-4e73-9d8b-a298bf1a501a · outbound

This paper cites CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation.

In-Context VLA: Endowing Vision-Language-Action Models with Language via In-Context Post-Training and Agentic Tool Use CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T00:30:03.164482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:30:03.164482Z digest=sha256:e74b6add6aa937455109eb60a0e8d947c73039d810a52f48a21987e08e229a36

Observation 53eece55-de9e-4048-9e66-99fad28bd8ea · outbound

This paper cites ABot-M0: VLA Foundation Model for Robotic Manipulation with Action Manifold Learning.

In-Context VLA: Endowing Vision-Language-Action Models with Language via In-Context Post-Training and Agentic Tool Use ABot-M0: VLA Foundation Model for Robotic Manipulation with Action Manifold Learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T00:30:03.169903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:30:03.169903Z digest=sha256:2d64cb6a6542048f5f96d9d57b41639e3a790e9eca0b4d6a418f33525d44d85d

Observation 9bab54ad-2aea-4e55-8cc6-935a0fff07b2 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

In-Context VLA: Endowing Vision-Language-Action Models with Language via In-Context Post-Training and Agentic Tool Use Advances in Neural Information Processing Systems , volume=

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:30:04.401760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T00:30:03.176272Z digest=sha256:40d9b19463ae599ae33f218b12244ba6755a7e62f418fda18f0823c72d240a2e

Observation d26d14f8-be1c-49df-a7d1-c0eb87eb3b33 · outbound

This paper cites arXiv preprint arXiv:2512.16793 , year=.

In-Context VLA: Endowing Vision-Language-Action Models with Language via In-Context Post-Training and Agentic Tool Use arXiv preprint arXiv:2512.16793 , year=

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-08T00:30:03.181620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:30:03.181620Z digest=sha256:61e8c6baa1b2650a9db859fed4645f2ab9feb0589d9a5acccdd5f071a023b874

Observation d3eac13e-9ff4-4a6e-83a7-049686782f5b · outbound

This paper cites arXiv preprint arXiv:2601.14133 , year=.

In-Context VLA: Endowing Vision-Language-Action Models with Language via In-Context Post-Training and Agentic Tool Use arXiv preprint arXiv:2601.14133 , year=

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T00:30:03.187460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:30:03.187460Z digest=sha256:7a6755a6d2d0d7622c72ef08429ac7711dc88efd4b6814b1f22d2a213df94f97

Observation b31fc78e-0ab8-413d-a379-dc434b8a0e3f · outbound

This paper cites F1: A Vision-Language-Action Model Bridging Understanding and Generation to Actions.

In-Context VLA: Endowing Vision-Language-Action Models with Language via In-Context Post-Training and Agentic Tool Use F1: A Vision-Language-Action Model Bridging Understanding and Generation to Actions

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T00:30:03.192785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:30:03.192785Z digest=sha256:bb5dd60eaf8aefe9f70c09ecbb3c8d94b5a9d687bf7fc0f29f9438d68fe23bd4

Observation 03e84617-c671-4265-959b-e2c0102bf50e · outbound

This paper cites SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model.

In-Context VLA: Endowing Vision-Language-Action Models with Language via In-Context Post-Training and Agentic Tool Use SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-08T00:30:03.200006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:30:03.200006Z digest=sha256:a4cb20813d63ddb4e38704b1b6dbf9411b8e95bccdf920f820008728a6e8cbef

Observation cac90b40-68a7-4dc2-8389-5d728c22452b · outbound

This paper cites The International Journal of Robotics Research , volume=.

In-Context VLA: Endowing Vision-Language-Action Models with Language via In-Context Post-Training and Agentic Tool Use The International Journal of Robotics Research , volume=

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-08T00:30:03.205571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:30:03.205571Z digest=sha256:b2a93354ae82cb9d7ef14a49290032942f394c1b7c52b21437159624dcb1f4a5

Observation 10af046f-e2f2-46c3-9b49-32ac84ddabd1 · outbound

This paper cites AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems.

In-Context VLA: Endowing Vision-Language-Action Models with Language via In-Context Post-Training and Agentic Tool Use AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-08T00:30:03.210982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:30:03.210982Z digest=sha256:64052b026864254031c26fbcc504d3da85fd439d6fa89d9347c901a9a3c88377

Observation fcabc654-5e06-4509-8f56-234e2db5dbec · outbound

This paper cites Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=.

In-Context VLA: Endowing Vision-Language-Action Models with Language via In-Context Post-Training and Agentic Tool Use Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T00:30:03.216216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:30:03.216216Z digest=sha256:6f7b55a127b9ef8912dc25b5386b1f0e05c34e7a0a9956a815201dafe977c8b3

Observation 2eae2f62-71b2-429d-8268-38737a1f2000 · outbound

This paper cites European conference on computer vision , pages=.

In-Context VLA: Endowing Vision-Language-Action Models with Language via In-Context Post-Training and Agentic Tool Use European conference on computer vision , pages=

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-08T00:30:03.221703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:30:03.221703Z digest=sha256:c54e4ba78e7f0fd25c8d290c816873797e1f72e3a8226a0434f1fccf844433d9

Observation 30abe21b-8325-4def-a671-800650bead3f · outbound

This paper cites Evaluating Real-World Robot Manipulation Policies in Simulation.

In-Context VLA: Endowing Vision-Language-Action Models with Language via In-Context Post-Training and Agentic Tool Use Evaluating Real-World Robot Manipulation Policies in Simulation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-08T00:30:03.227603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:30:03.227603Z digest=sha256:f0ec705670550354e591e96ad33256f2828506d0aaef9fb4ec84728077fa66e3

Observation 795120de-04e2-44e3-b023-8d31416b2d23 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

In-Context VLA: Endowing Vision-Language-Action Models with Language via In-Context Post-Training and Agentic Tool Use DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-08T00:30:03.233497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:30:03.233497Z digest=sha256:2b7a138dc2108fd825bcad804a5b6912ac2267fe72239e0d3ef68f29ad5cf5bb

Observation d74d6ebf-ad62-4be2-b055-f1c89be2d433 · outbound

This paper cites GR00T N1: An Open Foundation Model for Generalist Humanoid Robots.

In-Context VLA: Endowing Vision-Language-Action Models with Language via In-Context Post-Training and Agentic Tool Use GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-08T00:30:03.239113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:30:03.239113Z digest=sha256:939066e99ce5ea208d589b6f46fd73dabab353a04a17c002fe58888431df18e2

Observation b0d84f58-b6ef-4187-afb3-b069aafdbe19 · outbound

This paper cites Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=.

In-Context VLA: Endowing Vision-Language-Action Models with Language via In-Context Post-Training and Agentic Tool Use Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:30:04.347498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T00:30:03.244179Z digest=sha256:ee982d0908ba2524b4e8b71170d0f02aa91601e995a1d77e851de830e34058cb

Observation 4ac47c81-6e3d-4599-a3af-04d185bdae5f · outbound

This paper cites Advances in neural information processing systems , volume=.

In-Context VLA: Endowing Vision-Language-Action Models with Language via In-Context Post-Training and Agentic Tool Use Advances in neural information processing systems , volume=

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-08T00:30:03.249206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:30:03.249206Z digest=sha256:a07e0a3e0167d0d464e154abf70e830576891eabf3272bd74d87fc9e66c2b230

Observation 973885bc-d72e-4e38-8ea7-39073732b0bd · outbound

This paper cites Advances in neural information processing systems , volume=.

In-Context VLA: Endowing Vision-Language-Action Models with Language via In-Context Post-Training and Agentic Tool Use Advances in neural information processing systems , volume=

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-08T00:30:03.253931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:30:03.253931Z digest=sha256:8d8a501ef86c1d1b6bfd9b8bfc8c4199043ba6112d3959b57571c5320bfd88d2

Observation e9ec2099-f2a6-48dd-a2fa-c8b9eee4228f · outbound

This paper cites arXiv preprint arXiv:2602.01067 , year=.

In-Context VLA: Endowing Vision-Language-Action Models with Language via In-Context Post-Training and Agentic Tool Use arXiv preprint arXiv:2602.01067 , year=

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-08T00:30:03.258938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:30:03.258938Z digest=sha256:d64e96cb673449d680cf9d75ae445f7910cbd10231ce7b4bcef85809cb837e1a

Pith citing papers

No inbound Pith citation observations are available.