Pith. sign in

Paper Citation Record · LEDGER

GROW: Aligning GRPO with State-Action Modeling for Open-World VLM Agents

As of 24 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 0 inbound Pith citation observations for arXiv:2605.20246.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.20246 v2

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-22T09:06:03.221498Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

31 of 31 outbound references displayed

  • verified exact22
  • verified fuzzy8
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 066c5b2e-2e54-4b0f-a100-fbf187bc44b3 · outbound

This paper cites InquireMobile: Teaching VLM-based Mobile Agent to Request Human Assistance via Reinforcement Fine-Tuning.

GROW: Aligning GRPO with State-Action Modeling for Open-World VLM Agents InquireMobile: Teaching VLM-based Mobile Agent to Request Human Assistance via Reinforcement Fine-Tuning

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:06:18.886667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-22T09:06:03.221498Z digest=sha256:084281e130a0174aa4e09036673974ed30eb8ded5a2bd9fc22082c28a58fa3b3

Observation 0f291730-6b29-4c24-8038-779259a856a5 · outbound

This paper cites Qwen2.5-vl technical report.

GROW: Aligning GRPO with State-Action Modeling for Open-World VLM Agents Qwen2.5-vl technical report

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T09:06:20.314370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-22T09:06:03.221498Z digest=sha256:aee68117ad9f6ffb169bbd1b03115117b378f998910091d18491b632c5919ccc

Observation 67000697-7a1d-43ad-853d-60a644c93873 · outbound

This paper cites Qwen2.5-VL Technical Report.

GROW: Aligning GRPO with State-Action Modeling for Open-World VLM Agents Qwen2.5-VL Technical Report

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:06:20.235878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-22T09:06:03.221498Z digest=sha256:d04ce4f2f05dc2359a023d4134c09c69f483440aca9291d6d028583d69dbcefd

Observation a8bffc09-6934-4936-b0d6-e73e97eb518a · outbound

This paper cites Video pretraining (vpt): Learning to act by watching unlabeled online videos.Advances in Neural Information Processing Systems, 35:24639–24654.

GROW: Aligning GRPO with State-Action Modeling for Open-World VLM Agents Video pretraining (vpt): Learning to act by watching unlabeled online videos.Advances in Neural Information Processing Systems, 35:24639–24654

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T09:06:20.318146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-22T09:06:03.221498Z digest=sha256:e11d2a4104ce00fc4fe45cdf9eecf54635c8834ca75a8b5076105d4a63ef7a58

Observation 23336069-bd0d-4da8-bd94-12e81418c2be · outbound

This paper cites Scalable Multi-Task Reinforcement Learning for Generalizable Spatial Intelligence in Visuomotor Agents.

GROW: Aligning GRPO with State-Action Modeling for Open-World VLM Agents Scalable Multi-Task Reinforcement Learning for Generalizable Spatial Intelligence in Visuomotor Agents

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:06:20.226388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-22T09:06:03.221498Z digest=sha256:ac6de641c5cbba5ad63ca9fd1dd20626edcf4681b19ded651739cf3e399b7801

Observation 27697d87-0687-49be-80d8-6a6f8f6acc9a · outbound

This paper cites Freeman, Frédo Durand, Eli Shechtman, and Xun Huang.

GROW: Aligning GRPO with State-Action Modeling for Open-World VLM Agents Freeman, Frédo Durand, Eli Shechtman, and Xun Huang

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T09:06:18.909678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-22T09:06:03.221498Z digest=sha256:8eb90f65382c0f05529d4858bad5f39d516b158e62938054af7078a63c802854

Observation 23a36722-1d0e-4764-8ccb-ffa826d894d3 · outbound

This paper cites Rocket-1: Mastering open-world interaction with visual-temporal context prompting.

GROW: Aligning GRPO with State-Action Modeling for Open-World VLM Agents Rocket-1: Mastering open-world interaction with visual-temporal context prompting

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T09:06:20.307643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-22T09:06:03.221498Z digest=sha256:6c611a12bc45dc5bccf789cc3ccc4059f3e8df704bd8e139fd14dc5f3bb39df3

Observation fe28cfab-1196-4ff2-906d-d4bad9509c0c · outbound

This paper cites Liu, Ram Vasudevan, and Maani Ghaffari.

GROW: Aligning GRPO with State-Action Modeling for Open-World VLM Agents Liu, Ram Vasudevan, and Maani Ghaffari

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:06:20.217335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-22T09:06:03.221498Z digest=sha256:4f56e172cd1bd109da35dc03c730e3be6ba8f807097eba08c016d5d804335c84

Observation bb003685-2125-4a27-9b3e-36cd4c61ca87 · outbound

This paper cites Compassnav: Steering from path imitation to decision understanding in navigation.

GROW: Aligning GRPO with State-Action Modeling for Open-World VLM Agents Compassnav: Steering from path imitation to decision understanding in navigation

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T09:06:20.310942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-22T09:06:03.221498Z digest=sha256:b7d1f116a60796518b8dfa0071896b958eb97a181e8bde3d9ad8e812d7ec076d

Observation 70b77976-9598-43f0-b8db-aa3ea75dff5b · outbound

This paper cites JARVIS-VLA: Post- training large-scale vision language models to play visual games with keyboards and mouse.

GROW: Aligning GRPO with State-Action Modeling for Open-World VLM Agents JARVIS-VLA: Post- training large-scale vision language models to play visual games with keyboards and mouse

Reference 10

Resolution
verified exact
doi, observed 2026-05-22T09:06:18.900244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-22T09:06:03.221498Z digest=sha256:2fe73a57462d1ff37c333e08adb1c99c60eb10e30e228952246abe667629a899

Observation 8221cfde-8f7e-4bef-ba75-c718e1322d5a · outbound

This paper cites Coloragent: Building a robust, personalized, and interactive os agent.

GROW: Aligning GRPO with State-Action Modeling for Open-World VLM Agents Coloragent: Building a robust, personalized, and interactive os agent

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:06:20.207848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-22T09:06:03.221498Z digest=sha256:55e31d2e7a470020350228314f06031712d34eaeb0f9472c04b927cf11929f96

Observation fa8e4814-d2a9-4f46-87d5-b9de4ac112fc · outbound

This paper cites Steve-1: A generative model for text-to-behavior in minecraft.Advances in Neural Information Processing Systems, 36:69900–69929.

GROW: Aligning GRPO with State-Action Modeling for Open-World VLM Agents Steve-1: A generative model for text-to-behavior in minecraft.Advances in Neural Information Processing Systems, 36:69900–69929

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T09:06:20.300899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-22T09:06:03.221498Z digest=sha256:4d49af5fe3b5210b83ef3424415e76888bff934edd5139fdd89a3a1795a44a49

Observation 76f3ea3e-39cc-41b3-84cc-b3999cb62a63 · outbound

This paper cites MCU: An Evaluation Framework for Open-Ended Game Agents.

GROW: Aligning GRPO with State-Action Modeling for Open-World VLM Agents MCU: An Evaluation Framework for Open-Ended Game Agents

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:06:20.148998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-22T09:06:03.221498Z digest=sha256:87a3394175f6660c3d8aab5c5eb92c54015db64ef93485c3b035b0324b4554dc

Observation cc0cdb18-3e2a-4c25-bbde-e8b331d4a0a6 · outbound

This paper cites NaviMaster: Learning a Unified Policy for GUI and Embodied Navigation Tasks.

GROW: Aligning GRPO with State-Action Modeling for Open-World VLM Agents NaviMaster: Learning a Unified Policy for GUI and Embodied Navigation Tasks

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:06:20.198974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-22T09:06:03.221498Z digest=sha256:5c86a3966519e20721b1b2c2614844ca164a5bf65447be85ed68068ded01df92

Observation ea22657c-6db6-4cc3-b7b7-9da40c8ed92a · outbound

This paper cites Interactive Language: Talking to Robots in Real Time.

GROW: Aligning GRPO with State-Action Modeling for Open-World VLM Agents Interactive Language: Talking to Robots in Real Time

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:06:20.194865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-22T09:06:03.221498Z digest=sha256:783202e4277ab3f46b4fa1f07c10a2fb6742e89bb9c9df13f112642b035aad9f

Observation b5aef6f7-a30b-430b-a5ab-7c7f7ad48ee8 · outbound

This paper cites Nitrogen: An open foundation model for generalist gaming agents.

GROW: Aligning GRPO with State-Action Modeling for Open-World VLM Agents Nitrogen: An open foundation model for generalist gaming agents

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T09:06:20.297001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-22T09:06:03.221498Z digest=sha256:17ce8c109b83e9537729b99609541a9c3466bce0e0dc6b9c571b9993515fb3e9

Observation e5cdfbfa-7241-425a-bd87-39e978b04782 · outbound

This paper cites Nitrogen: An open foundation model for generalist gaming agents.

GROW: Aligning GRPO with State-Action Modeling for Open-World VLM Agents Nitrogen: An open foundation model for generalist gaming agents

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:06:20.190443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-22T09:06:03.221498Z digest=sha256:d8091cda2c94ec262d179f2c9a0c7abbfb55fa9809ac47cc2e02ea80ca5c042f

Observation fe16d6f4-0897-4cbd-9059-6e05603b89e4 · outbound

This paper cites Gameworld: Towards standardized and verifiable evaluation of multimodal game agents.

GROW: Aligning GRPO with State-Action Modeling for Open-World VLM Agents Gameworld: Towards standardized and verifiable evaluation of multimodal game agents

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T09:06:20.304168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-22T09:06:03.221498Z digest=sha256:aa40051e689d95828e054639f34b6fed02c2830cbb8f37f8c6f862471ca30f7a

Observation 48e98a13-db5c-4b6d-92a5-97c9c12688b2 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

GROW: Aligning GRPO with State-Action Modeling for Open-World VLM Agents DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:06:20.212533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-22T09:06:03.221498Z digest=sha256:b2c31ca6f738eda14ba7ae45bdb310a795ca672c37057aaccdf167c1f0cdfe06

Observation 63a64880-a6fa-41d1-a4d1-6c8ea82ca854 · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

GROW: Aligning GRPO with State-Action Modeling for Open-World VLM Agents HybridFlow: A Flexible and Efficient RLHF Framework

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:06:20.231172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-22T09:06:03.221498Z digest=sha256:df3aa23f3f5ee7f57fa07960ac91f048a488a608ed936efad537fa88c5fa60a2

Observation d7c09be8-2028-4ab9-b67d-f3049fe923ed · outbound

This paper cites MobileGUI-RL: Advancing Mobile GUI Agent through Reinforcement Learning in Online Environment.

GROW: Aligning GRPO with State-Action Modeling for Open-World VLM Agents MobileGUI-RL: Advancing Mobile GUI Agent through Reinforcement Learning in Online Environment

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:06:20.173865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-22T09:06:03.221498Z digest=sha256:e059cecd255c917b352990838b58e85bd3b81097806af75e952704fb7fe2916d

Observation f59a9f64-b71f-477a-a4e1-eef8872c6b4a · outbound

This paper cites Lumine: An open recipe for building generalist agents in 3d open worlds.

GROW: Aligning GRPO with State-Action Modeling for Open-World VLM Agents Lumine: An open recipe for building generalist agents in 3d open worlds

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:06:20.178376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-22T09:06:03.221498Z digest=sha256:1ad2467fff3a46a6c127a0f77e6b2abf085fb5ffd391c597a7a8e00b3baf6447

Observation 3a6782f4-b59d-491e-ae6d-c3869e43404a · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

GROW: Aligning GRPO with State-Action Modeling for Open-World VLM Agents Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:06:20.182452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-22T09:06:03.221498Z digest=sha256:9e9cd7d4f95a82f1aea59e8b5fc0ba5c7d538fdd4d5147af9acb91c72b96d37d

Observation 318df860-02fc-46be-b865-2464a94a6c03 · outbound

This paper cites Openha: A series of open-source hierarchical agentic models in minecraft.

GROW: Aligning GRPO with State-Action Modeling for Open-World VLM Agents Openha: A series of open-source hierarchical agentic models in minecraft

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:06:20.159090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-22T09:06:03.221498Z digest=sha256:c33d798c541a707e70db70e8de0d6f36cd16fdec60c56c46b9fc153e2438702b

Observation 3394bc5e-ef40-461d-b7c9-7483ed52d96e · outbound

This paper cites Game-tars: Pretrained foundation models for scalable generalist multimodal game agents.

GROW: Aligning GRPO with State-Action Modeling for Open-World VLM Agents Game-tars: Pretrained foundation models for scalable generalist multimodal game agents

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:06:20.164107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-22T09:06:03.221498Z digest=sha256:fa7df42c9be309a3fc583990fbd14eb8b146cb46c824756cb17093f47a93bf7c

Observation bdf00b0d-c6d3-4b39-9db2-1931e3f76182 · outbound

This paper cites Agentgym: Evaluating and training large language model-based agents across diverse environments.

GROW: Aligning GRPO with State-Action Modeling for Open-World VLM Agents Agentgym: Evaluating and training large language model-based agents across diverse environments

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T09:06:20.293099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-22T09:06:03.221498Z digest=sha256:2396a7c9f70f1e8f2b9002c32266322c27233f106b5572948ac711d65583b72b

Observation d3699671-cc8d-426f-b115-87275d78e961 · outbound

This paper cites Etp-r1: Evolving topological planning with reinforcement fine-tuning for vision-language navigation in continuous environments.

GROW: Aligning GRPO with State-Action Modeling for Open-World VLM Agents Etp-r1: Evolving topological planning with reinforcement fine-tuning for vision-language navigation in continuous environments

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:06:20.154173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-22T09:06:03.221498Z digest=sha256:fe1b97adf26043c7ad7f161a05304200ff0e03675dbbd20c967418fe9c078915

Observation 61751152-934b-419a-a5b7-aea962ca45f6 · outbound

This paper cites Agentrl: Scaling agentic reinforcement learning with a multi-turn, multi-task framework.

GROW: Aligning GRPO with State-Action Modeling for Open-World VLM Agents Agentrl: Scaling agentic reinforcement learning with a multi-turn, multi-task framework

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:06:20.168920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-22T09:06:03.221498Z digest=sha256:f23a1b54fcaba5ce02b44791322c980698ba65cc7e6761022ce5509ef90e5a69

Observation f75eb785-be6c-4683-9ee6-de09f904d8bb · outbound

This paper cites Activevln: Towards active exploration via multi-turn rl in vision-and-language navigation.

GROW: Aligning GRPO with State-Action Modeling for Open-World VLM Agents Activevln: Towards active exploration via multi-turn rl in vision-and-language navigation

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:06:20.186483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-22T09:06:03.221498Z digest=sha256:ef167c29ff79924ec3763ab44553aa66732d5667a78a5c0854db12167df30df7

Observation 9c48a5f6-d393-45f8-b7cf-1057e3bbb9da · outbound

This paper cites LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models.

GROW: Aligning GRPO with State-Action Modeling for Open-World VLM Agents LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:06:20.203515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-22T09:06:03.221498Z digest=sha256:8716d92454e0ec1f73dfd00ddcf3f7fde87c6b332e36d5ddd3b3efe6043dfa94

Observation 0f510b5a-f0e7-4489-be3b-178361324bbf · outbound

This paper cites mine iron ore.

GROW: Aligning GRPO with State-Action Modeling for Open-World VLM Agents mine iron ore

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:06:20.221332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-22T09:06:03.221498Z digest=sha256:15dc1f1b8be585e2827343eb71c850ac85d93694ad0642cc71eeb51d79c139d7

Pith citing papers

No inbound Pith citation observations are available.