Pith. sign in

Paper Citation Record · LEDGER

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning

As of 8 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 0 inbound Pith citation observations for arXiv:2606.29915.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.29915 v1

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-30T06:16:51.452335Z

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

56 of 56 outbound references displayed

  • verified exact12
  • verified fuzzy0
  • unresolved44
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 02b8d00c-4f93-4420-83ac-e36dd654eddb · outbound

This paper cites Don’t just assume; look and answer: Overcoming priors for visual question answering.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Don’t just assume; look and answer: Overcoming priors for visual question answering

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-30T06:16:51.452335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:f4dfcd79f00f0c22ad6e4a2accc452b2c28d59936ace76f357128a81462cc87f

Observation f2005225-0a92-4941-b183-d69a499f2df9 · outbound

This paper cites Neural module networks.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Neural module networks

Reference 2

Resolution
unresolved
no resolver link, observed 2026-06-30T06:16:51.452335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:5c9606ef845d7359c81576625ea46d43d5edaccc1c86d451fb434817e80161ab

Observation ee40dcdc-5370-4c95-a3d8-967bd50e467a · outbound

This paper cites Vqa: Visual question answering.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Vqa: Visual question answering

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-30T06:16:51.452335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:59764a661098312d2473954b465c950e4a74f44f271dd0232092e507a0778501

Observation 304897c1-bffe-4823-93be-ce55427eff1c · outbound

This paper cites Qwen2.5-VL Technical Report.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Qwen2.5-VL Technical Report

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-06-30T07:14:22.073999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T16:08:16.864468+00:00.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:512c8d9307d0e5c101ac2c44bfd05081e6454b56ee444f99aa9a4717aa077e4d

Observation d7f7db7e-3c3f-49a4-9b6e-ff5569249d74 · outbound

This paper cites SAM 3: Segment Anything with Concepts.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning SAM 3: Segment Anything with Concepts

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-06-30T06:24:19.583033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:7d303890c2ba8b92145ebc63434f0db81c6e3665f7c7ef9b46777cb05b6177a5

Observation 27fd4ac7-7e01-4009-a5d2-eb51a41c2ce3 · outbound

This paper cites Are we on the right way for evaluating large vision-language models?Advances in Neural Information Processing Systems, 37:27056–27087, 2024.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Are we on the right way for evaluating large vision-language models?Advances in Neural Information Processing Systems, 37:27056–27087, 2024

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-30T06:16:51.452335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:b4569ab0d1187f7ea27d9b6e6476d42feae6c37634619c419c971ab3ac8b0742

Observation 381137bd-9993-4a30-837d-a17fd62a0a69 · outbound

This paper cites Beyond question-based biases: Assessing multimodal shortcut learning in visual question answering.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Beyond question-based biases: Assessing multimodal shortcut learning in visual question answering

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-30T06:16:51.452335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:6bfaa10c2d32e373e19c0c69409baf3be51bd322361bc2b2f2dae69b91d2e3a2

Observation 567c85f0-f740-458e-9cdb-ab0859677745 · outbound

This paper cites Gemini 3 flash: High-efficiency agentic multimodal understanding.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Gemini 3 flash: High-efficiency agentic multimodal understanding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-30T06:16:51.452335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:e85507b86a8ee49c26deadee42f17c60215f11ffd5ea49536a5b806dd0e04c96

Observation 681706ef-c37d-430e-a621-b8633a67aebd · outbound

This paper cites Deepseek-r1 incentivizes reasoning in llms through reinforcement learning.Nature, 645(8081):633–638, 2025.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Deepseek-r1 incentivizes reasoning in llms through reinforcement learning.Nature, 645(8081):633–638, 2025

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-30T06:16:51.452335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:98400815d7422870a9afdb53c333ccf95616a3d6ef6634dbade3a0dfbd4a9616

Observation a41934e7-2ede-4f66-a00b-f76a93870f80 · outbound

This paper cites Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-06-30T07:14:22.071060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:b90e5f4d3d29d671f5a6004741de3f865ffb3b484380265777284b1b1a77595a

Observation 9cb8e320-72d3-4e2f-9a86-0ee4dd2e9462 · outbound

This paper cites Hudson and Christopher D.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Hudson and Christopher D

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-30T06:16:51.452335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:451be9130375ed650fd7ead75a9948566d183ea0e4a3bbe14a92473b8cb66b2a

Observation a8a852af-ba1f-4046-b480-da8f5a67508c · outbound

This paper cites GPT-4o System Card.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning GPT-4o System Card

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-06-30T06:24:19.584388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:3888de2d3e5dad344459bb1a401884ba6768ca5f1a5f3b0d28c21cf5c6345890

Observation cee35312-3f35-48a6-b8d7-35a4c954a5f3 · outbound

This paper cites Raven progressive matrices.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Raven progressive matrices

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-30T06:16:51.452335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:9353916632d5bb4db4049fc8fba919e9e4ac9c38c31beba1373ac844984ccdfe

Observation 77debea3-3348-4396-bc2c-d0d78d85c398 · outbound

This paper cites Clevr: A diagnostic dataset for compositional language and elementary visual reasoning.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Clevr: A diagnostic dataset for compositional language and elementary visual reasoning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-30T06:16:51.452335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:0cacdebc41042f203bb7102e8a6624b91b1053de9cbda66d5cd90fd8b581ad04

Observation bc695b8e-3edf-47f5-b018-e49f0a68ec03 · outbound

This paper cites Visual genome: Connecting language and vision using crowdsourced dense image annotations.International journal of computer vision, 123(1):32–73, 2017.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Visual genome: Connecting language and vision using crowdsourced dense image annotations.International journal of computer vision, 123(1):32–73, 2017

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-30T06:16:51.452335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:f8044590d35af89d45d6a974801b91f4d4fd18353b199df2e90eca638bec47d8

Observation c59b158b-a7c3-4a51-b698-9e6950b63d02 · outbound

This paper cites Imore: Implicit program-guided reasoning for human motion q&a.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Imore: Implicit program-guided reasoning for human motion q&a

Reference 16

Resolution
unresolved
no resolver link, observed 2026-06-30T06:16:51.452335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:1fbe013927f1b01043a3e2c6ff57e83555defa737312814ef06514e57941eb8a

Observation 68e336e7-f3a2-4420-ac31-b599f80d5a74 · outbound

This paper cites Vision-sr1: Self-rewarding vision-language model via reasoning decomposition and multi-reward policy optimization.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Vision-sr1: Self-rewarding vision-language model via reasoning decomposition and multi-reward policy optimization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-30T06:16:51.452335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:9af5151d3fe54ff40002be748489cb51cf629a7c0f9c30ddbaa852d90db0c6f4

Observation fccf2cf1-8ace-4b76-a479-ee35882b6513 · outbound

This paper cites Visual-rft: Visual reinforcement fine-tuning.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Visual-rft: Visual reinforcement fine-tuning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-30T06:16:51.452335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:afba1123ea005cab3b2e96f0b5511715745a70ce05fda52f578f6a431d6df132

Observation ca8cded6-2d71-4a6c-8461-950e1b81a47f · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.Advances in neural information processing systems, 35:2507–2521, 2022.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Learn to explain: Multimodal reasoning via thought chains for science question answering.Advances in neural information processing systems, 35:2507–2521, 2022

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-30T06:16:51.452335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:df3cc3d354fe9c05cedbf117960bcc9bde2190020591c1716f7593ac715025ec

Observation 9bfd8663-6812-4d8c-a8e8-ff796bca7233 · outbound

This paper cites Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-30T06:16:51.452335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:5618994bf91d5963720cf769b23ea029009edc546424dc241f44da92dd126832

Observation 08a5f508-d2ce-4c02-bce3-ef8ce740177d · outbound

This paper cites A computational investigation into the human representation and processing of visual information.WH San Francisco: Freeman and Company, San Francisco, 1(1):4, 1982.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning A computational investigation into the human representation and processing of visual information.WH San Francisco: Freeman and Company, San Francisco, 1(1):4, 1982

Reference 21

Resolution
unresolved
no resolver link, observed 2026-06-30T06:16:51.452335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:77f188c65213fcfb8d28e6fe7823ba59949654af2b23df58ee10ebfafe7d4261

Observation 33942dd2-30ed-426a-acdc-5777d7bab885 · outbound

This paper cites SmolVLM: Redefining small and efficient multimodal models.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning SmolVLM: Redefining small and efficient multimodal models

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-06-30T06:24:19.588487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:d4d4ca31c4c731ebd0f4742e2eb9e929009fc7a571841c1ef90cd3e5113f2889

Observation 1f35ee56-26f9-4737-8384-360dc2ed56b0 · outbound

This paper cites Ok-vqa: A visual question answering benchmark requiring external knowledge.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Ok-vqa: A visual question answering benchmark requiring external knowledge

Reference 23

Resolution
unresolved
no resolver link, observed 2026-06-30T06:16:51.452335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:b481e49323192c51d310f28a5371b976e34f27f654fcd49db4bfa5755b811de1

Observation 59f4b493-5f3e-45a1-9c08-676339e7dbae · outbound

This paper cites Pkr-qa: A benchmark for procedural knowledge reasoning with knowledge module learning.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Pkr-qa: A benchmark for procedural knowledge reasoning with knowledge module learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-06-30T06:16:51.452335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:9c7f426aeb39071b10068830b48376bfda093dee2a5ad18f4a15feff97da1d1e

Observation bf751ef3-1174-4d11-b2c2-3d524abe4764 · outbound

This paper cites an unresolved cited work.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-06-30T06:16:51.452335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:07537c0092221f8a718716ef47c99190503448843e67c8be2458628090eca315

Observation ccbc4a4e-de5e-42a3-a945-99785e6dda8b · outbound

This paper cites Grounding multimodal large language models to the world.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Grounding multimodal large language models to the world

Reference 26

Resolution
unresolved
no resolver link, observed 2026-06-30T06:16:51.452335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:273d88873650e20f00b0e362e853c70386e4192c8960791ba8228c1248c93955

Observation 5e974d39-0d41-4161-a197-8def928049a4 · outbound

This paper cites Dissecting Multimodality in VideoQA Transformer Models by Impairing Modality Fusion.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Dissecting Multimodality in VideoQA Transformer Models by Impairing Modality Fusion

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-06-30T06:24:19.585935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:6c5a55da2a359df6d4c2e76cd09f168779082727f46eecbceca9e7a360a4ae2c

Observation 96d199c0-0ddf-41ce-be50-f39ca01f802b · outbound

This paper cites Sentence-bert: Sentence embeddings using siamese bert-networks.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Sentence-bert: Sentence embeddings using siamese bert-networks

Reference 28

Resolution
unresolved
no resolver link, observed 2026-06-30T06:16:51.452335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:b6e1e95092132abc72df7dd6f008a238b274cbfa1b7a0df38810f409784b3a40

Observation b6c63f4b-5c70-4a23-962e-6d5d7467de5c · outbound

This paper cites an unresolved cited work.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-06-30T06:16:51.452335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:4f3123e796d89752c950fefd679b810cf626fe139f6f7261b25c887235c81900

Observation c8ad17e8-4c85-44a1-bf44-cc49afe30d08 · outbound

This paper cites Grounded reinforcement learning for visual reasoning.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Grounded reinforcement learning for visual reasoning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-06-30T06:16:51.452335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:da05fb1cb99c54817b1ec2228cb4eb02fc7d38b1726ace852f18aab59345b53e

Observation 324d5f17-1452-403d-810d-b9ed2a1b16ee · outbound

This paper cites A-okvqa: A benchmark for visual question answering using world knowledge.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning A-okvqa: A benchmark for visual question answering using world knowledge

Reference 31

Resolution
unresolved
no resolver link, observed 2026-06-30T06:16:51.452335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:40c10b2d91e32dda2d0fcc4a9c539c15c4149a1b0f214048b2a151ce6cc6becd

Observation 1a146134-7fb7-4d46-a9fd-869d1b678983 · outbound

This paper cites an unresolved cited work.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-06-30T06:16:51.452335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:ce54d0debd3f6ae3d7f62e839c544c8b4006bf3dff5b60777f590f96152f88e4

Observation ee428d28-21d9-4781-98d6-00f5aed48bb6 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-06-30T06:24:19.577448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:0333fb6017271cee644775e5afec763eaa09614a4b95f5ffc766e0fc512f79f9

Observation beb3ca3b-32e5-4dfb-9342-9ab25c3dc33e · outbound

This paper cites Vlm-r1: A stable and generalizable r1-style large vision-language model.arXiv e-prints, pages arXiv–2504, 2025.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Vlm-r1: A stable and generalizable r1-style large vision-language model.arXiv e-prints, pages arXiv–2504, 2025

Reference 34

Resolution
unresolved
no resolver link, observed 2026-06-30T06:16:51.452335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:7a010ebeef6703e4d089c1942b3281c16a356b4a504c3d1df38918bd47abd4e4

Observation 3817647c-3a84-414f-ad02-75b509cb8a2b · outbound

This paper cites Language prior is not the only shortcut: A benchmark for shortcut learning in VQA.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Language prior is not the only shortcut: A benchmark for shortcut learning in VQA

Reference 35

Resolution
unresolved
no resolver link, observed 2026-06-30T06:16:51.452335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:3f40e72eae6669cdeb79ce8aead489c44399ab4f518ea61e2478f68e5c1a11dd

Observation f8492fe8-ccf3-4c5d-babe-8f2ea424a8b6 · outbound

This paper cites Robospatial: Teaching spatial understanding to 2d and 3d vision-language models for robotics.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Robospatial: Teaching spatial understanding to 2d and 3d vision-language models for robotics

Reference 36

Resolution
unresolved
no resolver link, observed 2026-06-30T06:16:51.452335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:fe800aff59fe7ea104fce5a5030cab888e1dc35bd204d613d2f875c81724a104

Observation 8edf08a1-565c-4f43-8536-fbd4300c7b4f · outbound

This paper cites Aligning large multimodal models with factually augmented rlhf.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Aligning large multimodal models with factually augmented rlhf

Reference 37

Resolution
unresolved
no resolver link, observed 2026-06-30T06:16:51.452335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:1ff8f8140aea88a2f2539ccb1000101ce00ffea6bf4d16f44f05a8573059bee7

Observation c6c2f491-059a-47b0-bc86-75c48693c019 · outbound

This paper cites Reason-rft: Reinforcement fine-tuning for visual reasoning of vision language models.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Reason-rft: Reinforcement fine-tuning for visual reasoning of vision language models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-06-30T06:16:51.452335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:950db15040e8e3937917e4e389f3fc550e1a87ba53d54c289ba0edfc0c9d35f0

Observation df0e8031-8d0b-4862-becf-f94c559ca7b0 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Gemini: A Family of Highly Capable Multimodal Models

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-06-30T06:24:19.592194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:e4b29fea3f293969994c0c52693bd2d8d2da8147ac3ccac8383b62d93e602f1e

Observation d76f3881-784c-40ee-96e9-31edee27a189 · outbound

This paper cites Gemini Robotics: Bringing AI into the Physical World.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Gemini Robotics: Bringing AI into the Physical World

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-06-30T06:24:19.581646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:3d863466e09ed3a6d06c1bf904725a611e4b32df81a0999b75c1a9228edf3337

Observation 4ef73406-d8cf-452a-bf48-4a59a9975570 · outbound

This paper cites Qwen3.5-Omni Technical Report.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Qwen3.5-Omni Technical Report

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-06-30T06:24:19.590819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:c0eebc77facc4fa11c6e874b6f9036822001ccf09fe6f04a66cacc6886a2908c

Observation 9f84fce8-0f5a-44ee-a14e-82980fd649b9 · outbound

This paper cites Vl-rethinker: Incentivizing self-reflection of vision-language models with reinforcement learning.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Vl-rethinker: Incentivizing self-reflection of vision-language models with reinforcement learning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-06-30T06:16:51.452335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:4f21d03b0f425d1893755e82a990495c44c2db543adec072e324784d6912b0ba

Observation d1bc2cc3-8b94-4cda-897f-1267bc9e62b6 · outbound

This paper cites Procedures as a representation for data in a computer program for understanding natural language.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Procedures as a representation for data in a computer program for understanding natural language

Reference 43

Resolution
unresolved
no resolver link, observed 2026-06-30T06:16:51.452335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:d1aeaefbdbc4255088c17ba8942c4a014a13cfa0f8a9ab116d3cfb2c80933cb0

Observation a06c8446-a93e-497e-abd8-c75bdbe0a6fa · outbound

This paper cites Learning structural descriptions from examples.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Learning structural descriptions from examples

Reference 44

Resolution
unresolved
no resolver link, observed 2026-06-30T06:16:51.452335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:0559d7b1f60f76551ba8547677704a2df62a5478275c2b78b078120475781b65

Observation a4642844-0455-4f73-8e94-04cae7b4956a · outbound

This paper cites STAR: A Benchmark for Situated Reasoning in Real-World Videos.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning STAR: A Benchmark for Situated Reasoning in Real-World Videos

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-06-30T06:24:19.569669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:08fe895e87325e9a19a0a308720ac96d4d8115544cc4c34f2e8c6ce6727e20e4

Observation 5ea1c24a-e558-4ac0-877b-e6af5a798835 · outbound

This paper cites Realworldqa: A benchmark for real-world spatial understanding.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Realworldqa: A benchmark for real-world spatial understanding

Reference 46

Resolution
unresolved
no resolver link, observed 2026-06-30T06:16:51.452335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:6d2ab1493815bd4da20480755c67859f98b8dbd55e6dce247b46870596c02673

Observation 40a64413-d951-4af7-bb72-51832351807d · outbound

This paper cites Next-qa: Next phase of question- answering to explaining temporal actions.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Next-qa: Next phase of question- answering to explaining temporal actions

Reference 47

Resolution
unresolved
no resolver link, observed 2026-06-30T06:16:51.452335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:b685a6b951152542f94fcc528d0b6f2c8390ce106ae2a1d829562afbceef8023

Observation ad88ec18-9174-4c7f-abe2-bb20bd58f2d1 · outbound

This paper cites Neural- symbolic vqa: Disentangling reasoning from vision and language understanding.Advances in neural information processing systems, 31, 2018.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Neural- symbolic vqa: Disentangling reasoning from vision and language understanding.Advances in neural information processing systems, 31, 2018

Reference 48

Resolution
unresolved
no resolver link, observed 2026-06-30T06:16:51.452335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:e66ca7cf67f0f4c6adf52ddca6fe46a2aa68d3508068ed8c41ecd5fc04e02a61

Observation deaf51d4-4060-4d61-b0d3-0f4199430f68 · outbound

This paper cites Rlhf-v: Towards trustworthy mllms via behavior alignment from fine-grained correctional human feedback.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Rlhf-v: Towards trustworthy mllms via behavior alignment from fine-grained correctional human feedback

Reference 49

Resolution
unresolved
no resolver link, observed 2026-06-30T06:16:51.452335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:eb2d0c4882bda90de0000115dfd1b8fbe6b16b78997440b994ad33a2d1432c78

Observation 2fa4283e-d302-4393-83fb-ed25a5e22a12 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi

Reference 50

Resolution
unresolved
no resolver link, observed 2026-06-30T06:16:51.452335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:c1e6c031dfb940eeb603fda28ad2ce75753b5e5bef2780c45504775683003ed9

Observation 25ac8f50-9b0e-4b04-aeef-587f5588266e · outbound

This paper cites Raven: A dataset for relational and analogical visual reasoning.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Raven: A dataset for relational and analogical visual reasoning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-06-30T06:16:51.452335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:89a5491685abedd5a1aa8f27deba15797b65972177d352af89078e38fc22560f

Observation 5d838940-73f2-49c8-87ed-0113bd0f5b32 · outbound

This paper cites Mitigating Easy Option Bias in Multiple-Choice Question Answering.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Mitigating Easy Option Bias in Multiple-Choice Question Answering

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-06-30T06:24:19.587769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:f5a7990bc6a130b6538702a39a3eaf8fac4f7116a49d8380c65f1556dff06da6

Observation 00e4f341-87a6-4d92-97a7-6bf27303ca4a · outbound

This paper cites R1-vl: Learning to reason with multimodal large language models via step-wise group relative policy optimization.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning R1-vl: Learning to reason with multimodal large language models via step-wise group relative policy optimization

Reference 53

Resolution
unresolved
no resolver link, observed 2026-06-30T06:16:51.452335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:ad3d284ba5d1c08c2c41de76c000977d0c6d61d555483d06a78ad80771cbd76a

Observation 1df64fe7-114f-4e9d-97c2-1dd1c9a1d1e2 · outbound

This paper cites Physreason: A comprehensive benchmark towards physics-based reasoning.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Physreason: A comprehensive benchmark towards physics-based reasoning

Reference 54

Resolution
unresolved
no resolver link, observed 2026-06-30T06:16:51.452335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:f1cff3064f64c9874a471ad24ee0b3aa77844cbcdd87a6757346a58907f4fbdb

Observation 87f9e699-cdf2-4b17-80b9-ca1a51bfb1f1 · outbound

This paper cites Multimodal chain-of-thought reasoning in language models.Transactions on Machine Learning Research, 2024, 2024.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Multimodal chain-of-thought reasoning in language models.Transactions on Machine Learning Research, 2024, 2024

Reference 55

Resolution
unresolved
no resolver link, observed 2026-06-30T06:16:51.452335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:9d7e0d3688b5ab567010473b0567b1f69506dc78e94940861dfd9b819be88196

Observation cb99c87c-ba04-4766-b386-b18c03a4e5dc · outbound

This paper cites Visual7w: Grounded question answering in images.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Visual7w: Grounded question answering in images

Reference 56

Resolution
unresolved
no resolver link, observed 2026-06-30T06:16:51.452335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:7dfd4c475b55a1cd51763eb80d5ef5fbb29914dc51440de182465bb611233320

Pith citing papers

No inbound Pith citation observations are available.