Pith. sign in

Paper Citation Record · LEDGER

Does RLVR Extend Reasoning Boundaries? Investigating Capability Expansion in Vision-Language Models

As of 5 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 0 inbound Pith citation observations for arXiv:2511.00710.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2511.00710 v4

Coverage vector

measured 18 of 18 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-18T00:59:29.584299Z

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

18 of 18 outbound references displayed

  • verified exact12
  • verified fuzzy6
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2316059d-5590-43dc-a36c-32534d6de703 · outbound

This paper cites Training language models to follow instructions with human feedback.

Does RLVR Extend Reasoning Boundaries? Investigating Capability Expansion in Vision-Language Models Training language models to follow instructions with human feedback

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T01:00:35.272095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T00:59:29.584299Z digest=sha256:2b9673b418d3c872ad9aa24dcd05684f28e330bdda764d2bdf4f23f548c28ea9

Observation 1a67c20b-5143-44a7-a37e-93bdb3f72e9c · outbound

This paper cites Proximal Policy Optimization Algorithms.

Does RLVR Extend Reasoning Boundaries? Investigating Capability Expansion in Vision-Language Models Proximal Policy Optimization Algorithms

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-18T01:00:34.273634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T00:59:29.584299Z digest=sha256:badd475a204e0c268814a9a463f967d50bdb2fd67ab601a8a62cb58216eb197c

Observation 79ab8197-3bf9-4a69-9c4f-efb6e241d782 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Does RLVR Extend Reasoning Boundaries? Investigating Capability Expansion in Vision-Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-18T01:00:34.278151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T00:59:29.584299Z digest=sha256:c3fb329c3978c07b68c3ad494e78fb3d9ccc101222d5f4bf62c85e3cc60ec419

Observation b1788fd3-b179-4b8f-aa42-b86e6ce6c91f · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Does RLVR Extend Reasoning Boundaries? Investigating Capability Expansion in Vision-Language Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-18T01:00:34.269351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T00:59:29.584299Z digest=sha256:26b8d30ee761e7a068eaf3fd33ff82646f8d43cfcd74349350e32808e7c97b7a

Observation be2d22c2-e302-435b-81e4-3d7679467cb9 · outbound

This paper cites Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models.

Does RLVR Extend Reasoning Boundaries? Investigating Capability Expansion in Vision-Language Models Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-18T01:00:34.249459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T00:59:29.584299Z digest=sha256:94a7ec3f892763559f3efde57c663a993af6aba9ceccaa620b792d7726cf02ca

Observation c16e97bd-ee15-4ce7-8ad3-dd9496696245 · outbound

This paper cites VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning.

Does RLVR Extend Reasoning Boundaries? Investigating Capability Expansion in Vision-Language Models VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-18T01:00:34.287514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T00:59:29.584299Z digest=sha256:e0a309e9c3d55170ad5386e230ff98255033ffaf261de20c599d45fdd86c0b33

Observation d17eea14-dd59-43cf-83b6-309b72468a2c · outbound

This paper cites Srpo: Enhancing multimodal llm reasoning via reflection-aware rein- forcement learning.arXiv preprint arXiv:2506.01713.

Does RLVR Extend Reasoning Boundaries? Investigating Capability Expansion in Vision-Language Models Srpo: Enhancing multimodal llm reasoning via reflection-aware rein- forcement learning.arXiv preprint arXiv:2506.01713

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:00:34.264169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T00:59:29.584299Z digest=sha256:6e5001cd7c7081e82e7c2f36aba5a1f5a7a68e061c82c42b86a4c517f0d02e41

Observation 1c220fca-7c23-4a59-9a6d-c52bcdc1890e · outbound

This paper cites AlphaMaze: Enhancing Large Language Models' Spatial Intelligence via GRPO.

Does RLVR Extend Reasoning Boundaries? Investigating Capability Expansion in Vision-Language Models AlphaMaze: Enhancing Large Language Models' Spatial Intelligence via GRPO

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:00:34.259202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T00:59:29.584299Z digest=sha256:3295f693f565a7d030fd7142bbc100087327aa548acfee693add35be262ca7f1

Observation 423bc05a-b91c-47e2-ae9c-44728c60694d · outbound

This paper cites Learning to navigate in complex environments.

Does RLVR Extend Reasoning Boundaries? Investigating Capability Expansion in Vision-Language Models Learning to navigate in complex environments

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T01:00:35.269094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T00:59:29.584299Z digest=sha256:012e4e483567eafd74d6ddb7dee12974fbcb3c4656b45e9a55809be97e969228

Observation 07ee29a3-e891-4c5d-8ed0-10e437e5c914 · outbound

This paper cites Can Large Vision Language Models Read Maps Like a Human?.

Does RLVR Extend Reasoning Boundaries? Investigating Capability Expansion in Vision-Language Models Can Large Vision Language Models Read Maps Like a Human?

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:00:34.282715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T00:59:29.584299Z digest=sha256:5155010ddca03744c79bd418d21ade4fffb36d418cb7fd03a8f999dc149f10fe

Observation 99a10cf7-9e3a-41c6-854b-16341229ecaa · outbound

This paper cites Can mllms guide me home? a benchmark study on fine-grained visual reasoning from transit maps.

Does RLVR Extend Reasoning Boundaries? Investigating Capability Expansion in Vision-Language Models Can mllms guide me home? a benchmark study on fine-grained visual reasoning from transit maps

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:00:34.254687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T00:59:29.584299Z digest=sha256:bdb35ae0da5998124cfed3863fdaf2dbfa87c62a5126570f1cc670b04b2453c1

Observation 1a3b39e6-9a4b-4400-83b2-1d0f1ce0a28e · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

Does RLVR Extend Reasoning Boundaries? Investigating Capability Expansion in Vision-Language Models Flamingo: a visual language model for few-shot learning

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T01:00:35.260400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T00:59:29.584299Z digest=sha256:49cf143d2f96c32645214b4795bdbfeb162e1c98ffdd93eabda643690b81beb5

Observation e320745e-59cb-43f6-a9a0-966350d7fd0e · outbound

This paper cites Visual instruction tuning.

Does RLVR Extend Reasoning Boundaries? Investigating Capability Expansion in Vision-Language Models Visual instruction tuning

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T01:00:35.275080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T00:59:29.584299Z digest=sha256:dd33413c94fe51cc94b1f2b78478d127c981f5e4f06fc1a88f903ddbf9ec8b65

Observation 0d333c45-4989-4b9c-b800-6f8c453dff9c · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

Does RLVR Extend Reasoning Boundaries? Investigating Capability Expansion in Vision-Language Models InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-18T01:00:34.292214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T00:59:29.584299Z digest=sha256:dd06518c66f57aa8e5db2bb9e8bac4a638f566e3b45d5c9a127e341bb351c8b8

Observation cad7de51-1115-4151-b77d-4de34a206a25 · outbound

This paper cites Chain of thought prompting elicits reasoning in large language models.

Does RLVR Extend Reasoning Boundaries? Investigating Capability Expansion in Vision-Language Models Chain of thought prompting elicits reasoning in large language models

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T01:00:35.265990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T00:59:29.584299Z digest=sha256:e9ea285a811c1e6322deb346b3d42d5b35b3da4cf59aef1de6bb4dc155d9a979

Observation 205bf49e-2912-4f14-97fd-2c0ad26eddea · outbound

This paper cites Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning.

Does RLVR Extend Reasoning Boundaries? Investigating Capability Expansion in Vision-Language Models Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:00:34.244407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T00:59:29.584299Z digest=sha256:8538f2f57da7423bdaf7624d040622417e75ad4bbf05bd964f35d9bcabfc78d9

Observation 2a183e25-41f9-4635-a495-b98ced13be2f · outbound

This paper cites Cot-vla: Visual chain-of-thought reasoning for vision-language-action models.

Does RLVR Extend Reasoning Boundaries? Investigating Capability Expansion in Vision-Language Models Cot-vla: Visual chain-of-thought reasoning for vision-language-action models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T01:00:35.263175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T00:59:29.584299Z digest=sha256:9846e53fb75dad461caf9c016a4acbfa1861101b6b770bcf3afe7ce4e08b3e81

Observation d46a5da2-d0ab-4352-b507-0691fca6735e · outbound

This paper cites Qwen2.5-VL Technical Report.

Does RLVR Extend Reasoning Boundaries? Investigating Capability Expansion in Vision-Language Models Qwen2.5-VL Technical Report

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-18T01:00:34.295974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T00:59:29.584299Z digest=sha256:aff6a84aeff2ad9c547645fa519b136879960dc0f45046c8bf78f44f7fbfa1eb

Pith citing papers

No inbound Pith citation observations are available.