Pith. sign in

Paper Citation Record · LEDGER

SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models

As of 6 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 50 inbound Pith citation observations for arXiv:2504.11468.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.11468 v1

Coverage vector

measured 25 of 25 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-17T15:43:34.652373Z

measured 75 of 75 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 50 of 50 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T23:03:06.609642Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

25 of 25 outbound references displayed

  • verified exact0
  • verified fuzzy12
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 377d1c5d-6337-4d2c-965b-2bcd861d1e8c · outbound

This paper cites description.

SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models description

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T15:43:34.680977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T15:43:34.652373Z digest=sha256:f9a4b8b399f35e9b03ca489657d06feec2cc2a7cfd061b6ead3a707ec00441d6

Observation 2e6fdd68-68aa-4523-b52e-b9daa09501fe · outbound

This paper cites an unresolved cited work.

SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-05-17T15:43:34.688049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T15:43:34.652373Z digest=sha256:4ad67bbc67be026909fd4688b82589b96b6b6a2dd1c0ddb705e261fa6677bec6

Observation 6c138ff6-9497-47a2-bd75-c1f2e83be658 · outbound

This paper cites —— Here is the input: {input} Figure 10: Prompt for answer rewriting with GPT-4-Turbo.

SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models —— Here is the input: {input} Figure 10: Prompt for answer rewriting with GPT-4-Turbo

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T15:43:34.693175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T15:43:34.652373Z digest=sha256:3f75dc264222f21dd0537e94291e6376f53edc1756432ce2e24bc4c3d4eb603d

Observation d8b0a96e-1bc8-42c0-afdd-a53cf6be75dc · outbound

This paper cites an unresolved cited work.

SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-05-17T15:43:34.696741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T15:43:34.652373Z digest=sha256:760e4b1cf8a0f9b5d359314c17fff058b7ef66c861a3b28901f98f5dc347cfab

Observation 1f915c4d-46b9-46a5-a615-e9b11f715519 · outbound

This paper cites an unresolved cited work.

SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-05-17T15:43:34.700302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T15:43:34.652373Z digest=sha256:c939d41ca65c9bc637a715863f7cbd10c45b45bda019f6282d7264c1cba79689

Observation c214b2cc-f479-42a5-ba73-a1a99e336103 · outbound

This paper cites an unresolved cited work.

SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-05-17T15:43:34.703720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T15:43:34.652373Z digest=sha256:f112e2b5df264235f3a0ac7d66385c693c772195cf10dbdb6b5de2621efcfb5e

Observation bf48c659-72e1-44bd-9221-f4de2d795301 · outbound

This paper cites an unresolved cited work.

SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-05-17T15:43:34.707323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T15:43:34.652373Z digest=sha256:4cf5a878ab45fbfe6745edec8aa37d8593be5a76b39ede26193faca0b6121f64

Observation f9d54b84-939a-432b-9eef-6004030516f8 · outbound

This paper cites an unresolved cited work.

SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-05-17T15:43:34.710824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T15:43:34.652373Z digest=sha256:ee7554b88ce3199439d8a3e128fa22d063eac93a01ed9a0beb2d85ef7a89c9a7

Observation 817710e1-b9a3-4930-9f21-4c10b1b614da · outbound

This paper cites A”, “A)”, “(a).

SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models A”, “A)”, “(a)

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T15:43:34.714333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T15:43:34.652373Z digest=sha256:1be5209c3a33a8f6ba7df90ba601bd0db658e64b6ef76f41aca4721e23a3ddb9

Observation fd176a77-ae3b-455e-aec0-13f3cc800221 · outbound

This paper cites - The angle ∠ AOB = 36∘.

SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models - The angle ∠ AOB = 36∘

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T15:43:34.717727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T15:43:34.652373Z digest=sha256:15fa805b557b76f2546b11bae7d3ef4b62b5e7dd7789d68fb8e48a1ef16fd247

Observation 0bb6b597-dbce-475a-ac09-8ec07744f140 · outbound

This paper cites Therefore, ∠ OBA = 90.

SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models Therefore, ∠ OBA = 90

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T15:43:34.721190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T15:43:34.652373Z digest=sha256:882b5acbca51dd38e4dc815e2c9216e38f8ac898086caaab69be8d819d553ed5

Observation f9130c08-2e24-4054-8414-3bfa1debef23 · outbound

This paper cites an unresolved cited work.

SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-05-17T15:43:34.724942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T15:43:34.652373Z digest=sha256:416aaedc5cb0ff57750d4381866417bf384d344c6039deb298bc5125fcb7a3fa

Observation a00ecbf7-cba9-4c20-b997-cfecd68a0b6f · outbound

This paper cites - Points C and D are on the semicircle, with D being the midpoint of arc BC.

SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models - Points C and D are on the semicircle, with D being the midpoint of arc BC

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T15:43:34.728598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T15:43:34.652373Z digest=sha256:577d00127e3ed4079f313a0310d96d3f581cda6fbecffd595a4ccb4e3055f4d4

Observation 11eb9db7-a71f-474e-a285-ef97a3632064 · outbound

This paper cites - Midpoint of Arc: Since D is the midpoint of arc BC, arcs BD and DC are equal.

SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models - Midpoint of Arc: Since D is the midpoint of arc BC, arcs BD and DC are equal

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T15:43:34.733236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T15:43:34.652373Z digest=sha256:8b99e14bdc167a87085b57156c0e5bd0598cd1a862f3c80ea9b94a7d6061ab9f

Observation 4ad2767c-783d-479e-b817-83c8d5a89262 · outbound

This paper cites an unresolved cited work.

SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-05-17T15:43:34.736435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T15:43:34.652373Z digest=sha256:9d911bd5fb8adbec06bb2d3ede3e170fd655f8fdec18d447e24a14c6d024ff5b

Observation a4812136-1b4d-4a3f-b69a-d0d469bfa25a · outbound

This paper cites Let each be.

SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models Let each be

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T15:43:34.739290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T15:43:34.652373Z digest=sha256:b71e7847dcde8bba8e7206074d5f099da55f6a52852fc7507f517af303465c34

Observation 3fd4b36e-47d7-4f48-bc20-062b7fafde3f · outbound

This paper cites - By the Inscribed Angle Theorem.

SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models - By the Inscribed Angle Theorem

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T15:43:34.742450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T15:43:34.652373Z digest=sha256:6c4409db34ed05a687be825d3c40a57715d11a09974fc26cee19ed5adba9c0f9

Observation 9381ce95-e650-48ae-a8af-c2c11af64fc2 · outbound

This paper cites an unresolved cited work.

SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-05-17T15:43:34.745099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T15:43:34.652373Z digest=sha256:67948b53ea6393afe5ba784a8afdacce3b62eadda05173e173e835fbc665b84d

Observation f3ab0f9d-be62-43d5-895c-ff53adfe68b4 · outbound

This paper cites How many objects are left? •Original Answer: 3 Input Image <think> Okay, let's see.

SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models How many objects are left? •Original Answer: 3 Input Image <think> Okay, let's see

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T15:43:34.748216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T15:43:34.652373Z digest=sha256:1f49eae3fefe916fb4da13e7da93c0165e72843d521cec7c620cf36aaf8bec26

Observation 7dc0860c-e552-4caf-88ab-d4217fe9ea96 · outbound

This paper cites To find the value of , I'll substitute the coordinates of the point into the equation.

SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models To find the value of , I'll substitute the coordinates of the point into the equation

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T15:43:34.751072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T15:43:34.652373Z digest=sha256:fe54353e5f252d6bc2d3e05ee5802bc0ccb642915181d9654a247ca8c3a3a4c1

Observation 6d55bae2-aee4-497b-a2dd-16e25d5c1b82 · outbound

This paper cites an unresolved cited work.

SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-05-17T15:43:34.753665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T15:43:34.652373Z digest=sha256:283cf9533e45e173dfef19e7c54cfc0aafbd01f4e3ca69d38c17e08b109930b0

Observation 8f991278-9cf0-436c-8ba8-c5ec09e9d91b · outbound

This paper cites an unresolved cited work.

SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-05-17T15:43:34.756381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T15:43:34.652373Z digest=sha256:95a4bb140bf8c4561a2079da80a303b9d93a4081ea4e1df1100ead70154b7c83

Observation e8ab3baf-dc7d-47a8-85f8-4e8f74c3061a · outbound

This paper cites an unresolved cited work.

SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-05-17T15:43:34.758933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T15:43:34.652373Z digest=sha256:e1fcbd488bd38392adeb155d1e2c19cea5d8ee0f8148f799e5a0e3ad22c2637e

Observation 585d2d22-91e7-483f-91d3-8f3dc50cc1d5 · outbound

This paper cites an unresolved cited work.

SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-05-17T15:43:34.761573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T15:43:34.652373Z digest=sha256:26394729dd0c3c18122bb542e5a673ca070d1f28bd299fc0bb1b533861a66a74

Observation 649384af-e3c0-4f3c-9363-7e67250ae337 · outbound

This paper cites How many objects are left? •Original Answer: 3 Input Image <think> Okay, let's see.

SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models How many objects are left? •Original Answer: 3 Input Image <think> Okay, let's see

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T15:43:34.764485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T15:43:34.652373Z digest=sha256:56af9182af3b9e0f8e8bb7312acfb67487c8c2282fbc8866f575fa21f47bd519

Pith citing papers

Observation 663300b5-aaeb-4403-ae12-0634e5dbdc59 · inbound

OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles cites this paper.

OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-19T06:59:03.351961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T06:59:03.112252Z digest=sha256:d9f6160d7f315cf12a74247cd389efd2e9b22b60bc4cd5180142229843a224b9

Observation 6388e0e0-b733-4138-9ed3-b930cdf92b31 · inbound

GRIT: Teaching MLLMs to Think with Images cites this paper.

GRIT: Teaching MLLMs to Think with Images SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models

Reference 34

Resolution
metadata mismatch
local_arxiv, observed 2026-05-22T13:31:36.123124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T13:29:47.529564Z digest=sha256:535a86aff7a909f30e080de39fd46607c09b61e6eef7debfe63de62e2ff2adf7

Observation edeabbe8-7b09-4e80-bc84-dfeaafec1ec7 · inbound

Blending Supervised and Reinforcement Fine-Tuning with Prefix Sampling cites this paper.

Blending Supervised and Reinforcement Fine-Tuning with Prefix Sampling SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-21T23:40:46.153525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-21T23:39:39.018498Z digest=sha256:d309c4df597a6f9cff782c277fed7753f86ebbb9d89b2035d35b6b9f00f0fa39

Observation ced92bbf-66a5-4550-aaa8-51ebd37cd72b · inbound

WebSailor: Navigating Super-human Reasoning for Web Agent cites this paper.

WebSailor: Navigating Super-human Reasoning for Web Agent SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-17T15:43:34.765363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T15:37:09.572241Z digest=sha256:14a53a1319abd58b3b7331ee4059037743b23e03379047205f8b04cb64b2dead

Observation 21777939-f9bb-4490-ade6-fcc4ea013ce3 · inbound

A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence cites this paper.

A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models

Reference 240

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T15:43:34.765363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-14T22:23:14.621091Z digest=sha256:295c6e00e4beeb7ea37d750589989440e9ff1741dd1913c9089f39215fd66fc0

Observation 52ab568b-7125-4c5f-8753-0f7549bbe07d · inbound

MathReal: We Keep It Real! A Real Scene Benchmark for Evaluating Math Reasoning in Multimodal Large Language Models cites this paper.

MathReal: We Keep It Real! A Real Scene Benchmark for Evaluating Math Reasoning in Multimodal Large Language Models SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T23:03:06.609642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:03:06.609642Z digest=sha256:033a0c3176e37d1f8e36baf323cae09b37422cab50d24842e763831bf82ed9f6

Observation 584c915c-2f4b-4f5e-abde-3e2cf2752b2d · inbound

UAV-VL-R1: Generalizing Vision-Language Models via Supervised Fine-Tuning and Multi-Stage GRPO for UAV Visual Reasoning cites this paper.

UAV-VL-R1: Generalizing Vision-Language Models via Supervised Fine-Tuning and Multi-Stage GRPO for UAV Visual Reasoning SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-18T22:21:53.234074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T22:17:41.758059Z digest=sha256:73b87118802295627637856f9f37f6afb45013f91763090f00bd26a4aab8ebb9

Observation 4f8c97d2-a1cf-4e78-8df9-dba7bdfd113c · inbound

Reinforced Visual Perception with Tools cites this paper.

Reinforced Visual Perception with Tools SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:04.885212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:27:04.885212Z digest=sha256:c91723d2a5551c48fd998fcab4d0f7b6b21ddfbedf2e320ef08948a813139bbe

Observation cd903724-f451-48e7-ab85-592fedc27b69 · inbound

Less is More Tokens: Efficient Math Reasoning via Difficulty-Aware Chain-of-Thought Distillation cites this paper.

Less is More Tokens: Efficient Math Reasoning via Difficulty-Aware Chain-of-Thought Distillation SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T05:34:40.243499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:34:40.243499Z digest=sha256:c7bb3a94377e61b9868c212e3474151ac86d51f073435dbc5a9fb983b0475660

Observation 93596ae4-2186-456b-8029-e254c01e1a43 · inbound

Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding cites this paper.

Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T22:57:44.986207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T22:57:44.986207Z digest=sha256:10bc51ccaac4713f7bc0f8d125d3f0ebf1dc523a5424f1fb06c58a22a245b273

Observation fe009cb9-2b50-4598-83d3-acbabef4f6b3 · inbound

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle cites this paper.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:28.361696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:28.361696Z digest=sha256:6864521d4a79aebba04ea6b1c194c2b458c143cba6c7d0bae9da00316c81bf48

Observation d1d1bde5-9bfa-410f-af95-ce1fe03e63c7 · inbound

Failure Makes the Agent Stronger: Enhancing Accuracy through Structured Reflection for Reliable Tool Interactions cites this paper.

Failure Makes the Agent Stronger: Enhancing Accuracy through Structured Reflection for Reliable Tool Interactions SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-18T15:01:31.296494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T15:00:51.162221Z digest=sha256:63cc47a38e8d077a7646088ac7830ba7b2ac559d60bc53aad98102b7386c04b3

Observation d4b59fc0-9878-4ddd-876b-91f6ccbf08e0 · inbound

Mixture-of-Visual-Thoughts: Exploring Context-Adaptive Reasoning Mode Selection for General Visual Reasoning cites this paper.

Mixture-of-Visual-Thoughts: Exploring Context-Adaptive Reasoning Mode Selection for General Visual Reasoning SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-18T13:36:25.155140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-18T13:33:12.508639Z digest=sha256:6a71d161159ee56a38fb469afd533a3713bfd11ed931a0c19141ec3adf6ac7be

Observation 9b8dd9e8-3217-48f5-9044-2713a86776b4 · inbound

Dynamic-TreeRPO: Breaking the Independent Trajectory Bottleneck with Structured Sampling cites this paper.

Dynamic-TreeRPO: Breaking the Independent Trajectory Bottleneck with Structured Sampling SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-21T21:30:39.192536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T21:30:24.268903Z digest=sha256:a180c59d589249e9bc5b4ce9d5cedac692c126610735d26b895b1d8f7ce73673

Observation 2835058f-2941-4d99-9c7d-e98982c5b5c9 · inbound

Latent Visual Reasoning cites this paper.

Latent Visual Reasoning SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-17T15:43:34.765363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T18:41:30.307521Z digest=sha256:929008dcbda9bb6d800a1ff9cdca526e9c46781258f3a653b0c80a5ef6663741

Observation 8c171be8-466d-4c50-a6b1-6e4562419737 · inbound

VOLD: Reasoning Transfer from LLMs to Vision-Language Models via On-Policy Distillation cites this paper.

VOLD: Reasoning Transfer from LLMs to Vision-Language Models via On-Policy Distillation SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T07:57:22.810986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:57:22.810986Z digest=sha256:093ff712da3c83e70e145f969c1dafc0185929f9f5bc123a81d49b6d6d1a6f02

Observation 94162f6f-42a4-47c7-84a6-2e99cc85a321 · inbound

DeepEyesV2: Toward Agentic Multimodal Model cites this paper.

DeepEyesV2: Toward Agentic Multimodal Model SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-17T15:43:34.765363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T05:32:29.266583Z digest=sha256:632fb66b96fa2014e6cbc7eb41896911a4eaed68bfe9573e04a7864756c37e2c

Observation 227654e6-ca34-4b57-9ab6-70a07116a4b7 · inbound

Reinforcement Learning Improves Traversal of Parametric Knowledge in LLMs cites this paper.

Reinforcement Learning Improves Traversal of Parametric Knowledge in LLMs SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T23:27:34.982009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:27:34.982009Z digest=sha256:5f09d2ea1b6f59c0bc942c0cd887ec79d1224fe521f37560980dc76897bd990d

Observation 4c20dbad-0a8a-46b7-bf7b-cb81c16b9fea · inbound

MultiPriv: Benchmarking Individual-Level Privacy Reasoning in Vision-Language Models cites this paper.

MultiPriv: Benchmarking Individual-Level Privacy Reasoning in Vision-Language Models SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-03T21:09:13.251555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:09:13.251555Z digest=sha256:ced132a4eeba6345f6c4758760856f4ad9e4e208bd3d1cd83a2e996c3bf7bbb0

Observation 6f0c7d33-ee0b-4908-8407-12343980bcf5 · inbound

TeamPath: Building MultiModal Pathology Experts with Reasoning AI Copilots cites this paper.

TeamPath: Building MultiModal Pathology Experts with Reasoning AI Copilots SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models

Reference 62

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:30:11.116491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:29:39.744103Z digest=sha256:adba8b941f6992ce586a546047e1e57d62349dd86ef3c4050c6e50f43d9a76e3

Observation ddf35aeb-ab0b-4711-8dff-c7247bb83273 · inbound

Distilling Counterfactual Reasoning from Language to Vision: Causal Graph Guided Post-Training for Video Understanding cites this paper.

Distilling Counterfactual Reasoning from Language to Vision: Causal Graph Guided Post-Training for Video Understanding SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T20:27:11.411364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:27:11.411364Z digest=sha256:0ad85f0c08a55dadc72dab2760e8b71c74f5d6e4e112e30b45dea232063a2952

Observation feca519a-fadb-41ef-bbfc-b48c1c8ccc14 · inbound

Asking like Socrates: Socrates helps VLMs understand remote sensing images cites this paper.

Asking like Socrates: Socrates helps VLMs understand remote sensing images SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-17T15:43:34.765363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T04:54:57.121481Z digest=sha256:d44956a88bbc6fb099c890f41e7061f1395b63d0b82e4620f939bd0c5bd3f331

Observation 88674871-a74a-4b1b-8776-512f7945ef01 · inbound

Reasoning Within the Mind: Dynamic Multimodal Interleaving in Latent Space cites this paper.

Reasoning Within the Mind: Dynamic Multimodal Interleaving in Latent Space SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-17T15:43:34.765363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T23:02:28.588225Z digest=sha256:3b08b54f40f43cbdae8e1a501b6cb47fc5eea08797d8f35db3087b1c235240c6

Observation 9cf111ae-be03-4f32-b561-3108f6f4307d · inbound

Mimic Human Cognition, Master Multi-Image Reasoning: A Meta-Action Framework for Enhanced Visual Understanding cites this paper.

Mimic Human Cognition, Master Multi-Image Reasoning: A Meta-Action Framework for Enhanced Visual Understanding SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T11:13:16.762815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:13:16.762815Z digest=sha256:6809a7a870b23a2cab84e189160945e6d00abfc9d4652a100099b5a0147d9536

Observation 823cb637-64b5-4a87-90d6-d521c9eadfcb · inbound

ReMoT: Reinforcement Learning with Motion Contrast Triplets cites this paper.

ReMoT: Reinforcement Learning with Motion Contrast Triplets SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T19:58:14.742025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:58:14.742025Z digest=sha256:6587d95a4ec614afb1a5ba799ade881a9f18cd7783e0c4b1c9cf2135755b9591

Observation c70ed5cd-9388-4a7e-92ce-e6a59593d17f · inbound

Teaching an Agent to Sketch One Part at a Time cites this paper.

Teaching an Agent to Sketch One Part at a Time SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-17T15:43:34.765363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T07:53:15.953601Z digest=sha256:9ed785d1da5d71fdbe827991cea4f0e24b5e469f2de314ae7cf9f32ceea44349

Observation 6e169cf3-e0d2-41a3-b71a-0e8a488fd26a · inbound

Understanding the Role of Hallucination in Reinforcement Post-Training of Multimodal Reasoning Models cites this paper.

Understanding the Role of Hallucination in Reinforcement Post-Training of Multimodal Reasoning Models SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-17T15:43:34.765363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T20:48:52.130130Z digest=sha256:c27c1b32095cae168d4f5604204865eef4b7d12d88569d48cc7d66a0c6907e71

Observation 9b8e4736-2200-4c0d-8ff0-d93c1f1727df · inbound

Watch Before You Answer: Learning from Visually Grounded Post-Training cites this paper.

Watch Before You Answer: Learning from Visually Grounded Post-Training SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T15:43:34.765363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T20:01:13.305374Z digest=sha256:96db316af3696c82eb4693a93e4cafe7f799163ecf49f3b3ca12339c2526179b

Observation 53641dc4-bc0b-471e-96dc-847ae3396003 · inbound

Act Wisely: Cultivating Meta-Cognitive Tool Use in Agentic Multimodal Models cites this paper.

Act Wisely: Cultivating Meta-Cognitive Tool Use in Agentic Multimodal Models SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-17T15:43:34.765363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:35:21.514502Z digest=sha256:4d367d94ffb3415087f7769e97db4e13e1b88818082b41207c9571c8a376ef6f

Observation 63c54e0e-f8e2-44a0-8c55-0f24920ef4a1 · inbound

Generalization in LLM Problem Solving: The Case of the Shortest Path cites this paper.

Generalization in LLM Problem Solving: The Case of the Shortest Path SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-17T15:43:34.765363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T10:37:45.355872Z digest=sha256:af4ca18c4d5a38ef6b8a0dc6899086d20b31f52b13f8eda42073931ca31354b8

Observation 7c852431-d69f-4047-8c7c-257aa0598ff3 · inbound

Characterizing Model-Native Skills cites this paper.

Characterizing Model-Native Skills SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models

Reference 76

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T15:43:34.765363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-10T05:42:49.694715Z digest=sha256:4f367ea2a10cda797558fd26cbff1c53532d9f65b7ae405426a01f36bc578d04

Observation 69ca1e44-4ed7-4c75-96e5-13fe8d1825ff · inbound

One Step Forward and K Steps Back: Better Reasoning with Denoising Recursion Models cites this paper.

One Step Forward and K Steps Back: Better Reasoning with Denoising Recursion Models SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models

Reference 144

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T15:43:34.765363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-10T04:56:35.796962Z digest=sha256:615decbbd1e35994e8da6ba2836d91134a971c37a1aba827cfb25ecf2e2a2c93

Observation b1f6c7e3-8b8c-4ebb-a27d-2f6cb7d31507 · inbound

CGC: Compositional Grounded Contrast for Fine-Grained Multi-Image Understanding cites this paper.

CGC: Compositional Grounded Contrast for Fine-Grained Multi-Image Understanding SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-17T15:43:34.765363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T12:26:01.568507Z digest=sha256:066a0d6a0794f5d0597ca9fac30785051010dba8e51132e48af6343394e44a76

Observation e8051922-815f-4367-8c4b-300a29563b98 · inbound

How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors cites this paper.

How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-17T15:43:34.765363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T03:25:04.955816Z digest=sha256:a06d4f4f2ac0b914d5825d42307534bffe296f32b3c5249b99759484cc18e180

Observation f2176436-10eb-4ea0-bbf8-da2b115a9287 · inbound

Learning Agentic Policy from Action Guidance cites this paper.

Learning Agentic Policy from Action Guidance SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-17T15:43:34.765363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:f4d98e490e76e5829561df821bf89f77e883a25c4b0538aa275f5c6a67c8f5d4

Observation 11233d49-b0da-4cfd-9cd2-a9385544f928 · inbound

Teacher-Guided Policy Optimization for On-Policy Reasoning Distillation under Large Policy Divergence cites this paper.

Teacher-Guided Policy Optimization for On-Policy Reasoning Distillation under Large Policy Divergence SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-17T15:43:34.765363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-14T20:12:32.638453Z digest=sha256:b4b1273a217f7012613b13211eb9f066abbfd99e28c9039cede5d293d7280c65

Observation b98e07f9-c4ac-4ba1-9170-79703cd4b01f · inbound

Are Tools Always Beneficial? Learning to Invoke Tools Adaptively for Dual-Mode Multimodal LLM Reasoning cites this paper.

Are Tools Always Beneficial? Learning to Invoke Tools Adaptively for Dual-Mode Multimodal LLM Reasoning SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models

Reference 40

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T06:18:05.178174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:16:47.650748Z digest=sha256:ddcd9b825691c6ce3093261cb0fe54f1ca8634279c21a474f954118d9a95e8a5

Observation 9266a140-c721-4919-ad30-99c840ebd9e8 · inbound

RISE: Reliable Improvement in Self-Evolving Vision-Language Models cites this paper.

RISE: Reliable Improvement in Self-Evolving Vision-Language Models SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-21T05:39:40.523567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T05:38:26.590720Z digest=sha256:2f56ebf3bd464f93505639b53627ad4ec24d7f6b0bfb5786167565360e74e7ec

Observation b0cb786d-7ac9-4b0e-8b91-a44e05412304 · inbound

RISE: Reliable Improvement in Self-Evolving Vision-Language Models cites this paper.

RISE: Reliable Improvement in Self-Evolving Vision-Language Models SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-06-30T17:24:57.491852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T17:19:09.596950Z digest=sha256:a3da3d577d1c24b0f5f59782df465f242299a571e65a5fb51d64a0b46bd39511

Observation 3785af52-345f-441d-a180-2d76d5c2776b · inbound

AnE: Pushing the Reasoning Frontier of Multimodal LLMs via Anchor Evolution cites this paper.

AnE: Pushing the Reasoning Frontier of Multimodal LLMs via Anchor Evolution SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-06-30T00:14:04.569192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T22:56:39.504430Z digest=sha256:3a4993c455e108a70025ed9498e017da763f63d6e1faee08ec34771533001707

Observation 30f6306f-c6fc-4380-a466-508f6fcaa267 · inbound

ToolFG: Towards Well-Grounded Fine-Grained Image Classification cites this paper.

ToolFG: Towards Well-Grounded Fine-Grained Image Classification SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:56:20.349143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T14:52:06.186594Z digest=sha256:1f44c67b59a8591e524cc99844c0b2f9d1a0633c0287490a384003b16bbf56cc

Observation a4d67d8c-9984-4a16-86f5-a83c116e52e9 · inbound

Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients cites this paper.

Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models

Reference 84

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:48:56.121032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T01:08:52.981296Z digest=sha256:8b943082747e79a94955b7e26240469ac513f44a00b06dbb5d9096bac0f56b0d

Observation 1826d3fd-d7dc-49db-a281-4d2ef82897ab · inbound

Beyond Trajectory Imitation: Strategy-Guided Policy Optimization for LLM Reasoning cites this paper.

Beyond Trajectory Imitation: Strategy-Guided Policy Optimization for LLM Reasoning SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-04T16:29:57.013212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-26T00:39:29.703711Z digest=sha256:87ec9cc5d7500f1662d4d5afd0e09606a81cd94f090e27d75570fa2e973d440c

Observation 6c48bdaf-3f3a-4b4a-b366-ee810459815d · inbound

Bridging Physical Reasoning and Task Generalization via Visual Action Outcome Reasoning Alignment cites this paper.

Bridging Physical Reasoning and Task Generalization via Visual Action Outcome Reasoning Alignment SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-08T03:14:31.698320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-08T03:13:07.963343Z digest=sha256:a70efda7836e26eed60520777b834ec729c1c9f5593baf93781144c61fe4bc37

Observation 757dbb39-9cbf-49c1-b8e5-010dd6e76ef0 · inbound

BUS: Brain-Inspired Unsupervised Self-Reflection via Backward Prediction for Multimodal Reasoning cites this paper.

BUS: Brain-Inspired Unsupervised Self-Reflection via Backward Prediction for Multimodal Reasoning SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models

Reference 110

Resolution
metadata mismatch
local_arxiv, observed 2026-07-09T13:56:19.219331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-07-09T13:51:49.149342Z digest=sha256:6fa631d0dd3ff7f70a131fd1ef4ac2ecf62ce270af740469ceaa4c8bc6d3be18

Observation ac7f4469-3187-48a0-ae8c-a5b8904426d0 · inbound

BUS: Brain-Inspired Unsupervised Self-Reflection via Backward Prediction for Multimodal Reasoning cites this paper.

BUS: Brain-Inspired Unsupervised Self-Reflection via Backward Prediction for Multimodal Reasoning SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models

Reference 110

Resolution
unresolved
no resolver link, observed 2026-07-13T06:45:27.857034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T06:45:27.857034Z digest=sha256:cbd2a6a3f6bb3c788ce96f9223cb3a991370b37479f45b9b0a76fa07a59a8120

Observation 3b2f3b5e-95e9-43e5-89b4-e24edf9eb807 · inbound

SIVA-RL: Sensitivity-Invariance Visual Alignment for Multimodal Reinforcement Learning cites this paper.

SIVA-RL: Sensitivity-Invariance Visual Alignment for Multimodal Reinforcement Learning SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-02T03:20:41.224809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:20:41.224809Z digest=sha256:b1915d70e9985ecfe45f9fac827971c0e3fae90caabc08d8f5f881555783978f

Observation 69641fed-a3c3-4135-b179-c2f374505a6a · inbound

Reasoning to Regulate: Chain-of-Thought for Traffic Rule Understanding cites this paper.

Reasoning to Regulate: Chain-of-Thought for Traffic Rule Understanding SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-07-31T21:13:05.201592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:13:05.201592Z digest=sha256:fbc708829599e09f689e0bc53fa76fe9f95278f556773fa150edce4c3c5af5d8

Observation cf4436d1-3642-4344-af48-54e7b424dc8c · inbound

Knowing When to Quit: Diagnosing and Training LLMs to Abort Futile Reasoning cites this paper.

Knowing When to Quit: Diagnosing and Training LLMs to Abort Futile Reasoning SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-03T11:36:49.162119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T11:36:49.162119Z digest=sha256:a485ed20a746e4fc99536802db30bd45be473e26c3338c533871bab6fb1cdbe9

Observation 701a8f4f-69ed-42a1-9375-ccd59810fc97 · inbound

Balancing Efficiency and Efficacy: Training-Free Attention-Guided Switching Between Explicit and Latent Thoughts for MLLMs cites this paper.

Balancing Efficiency and Efficacy: Training-Free Attention-Guided Switching Between Explicit and Latent Thoughts for MLLMs SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T19:03:07.091870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:03:07.091870Z digest=sha256:8cdb58ab840307a1dc3dae1adbde77bf8f79901b6f83b6a7b056149faef81252