Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-17T15:43:34.652373Z
Paper Citation Record · LEDGER
As of 6 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 50 inbound Pith citation observations for arXiv:2504.11468.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-17T15:43:34.652373Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-05T23:03:06.609642Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
25 of 25 outbound references displayed
External citation measurements
0
pith, observed 2026-08-05T02:28:24.338817Z
Observation 377d1c5d-6337-4d2c-965b-2bcd861d1e8c · outbound
SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models description
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2e6fdd68-68aa-4523-b52e-b9daa09501fe · outbound
SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models Unresolved cited work
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6c138ff6-9497-47a2-bd75-c1f2e83be658 · outbound
SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models —— Here is the input: {input} Figure 10: Prompt for answer rewriting with GPT-4-Turbo
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d8b0a96e-1bc8-42c0-afdd-a53cf6be75dc · outbound
SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models Unresolved cited work
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1f915c4d-46b9-46a5-a615-e9b11f715519 · outbound
SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models Unresolved cited work
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c214b2cc-f479-42a5-ba73-a1a99e336103 · outbound
SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models Unresolved cited work
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation bf48c659-72e1-44bd-9221-f4de2d795301 · outbound
SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models Unresolved cited work
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f9d54b84-939a-432b-9eef-6004030516f8 · outbound
SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models Unresolved cited work
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 817710e1-b9a3-4930-9f21-4c10b1b614da · outbound
SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models A”, “A)”, “(a)
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation fd176a77-ae3b-455e-aec0-13f3cc800221 · outbound
SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models - The angle ∠ AOB = 36∘
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0bb6b597-dbce-475a-ac09-8ec07744f140 · outbound
SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models Therefore, ∠ OBA = 90
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f9130c08-2e24-4054-8414-3bfa1debef23 · outbound
SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models Unresolved cited work
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a00ecbf7-cba9-4c20-b997-cfecd68a0b6f · outbound
SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models - Points C and D are on the semicircle, with D being the midpoint of arc BC
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 11eb9db7-a71f-474e-a285-ef97a3632064 · outbound
SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models - Midpoint of Arc: Since D is the midpoint of arc BC, arcs BD and DC are equal
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4ad2767c-783d-479e-b817-83c8d5a89262 · outbound
SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models Unresolved cited work
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a4812136-1b4d-4a3f-b69a-d0d469bfa25a · outbound
SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models Let each be
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3fd4b36e-47d7-4f48-bc20-062b7fafde3f · outbound
SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models - By the Inscribed Angle Theorem
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9381ce95-e650-48ae-a8af-c2c11af64fc2 · outbound
SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models Unresolved cited work
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f3ab0f9d-be62-43d5-895c-ff53adfe68b4 · outbound
SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models How many objects are left? •Original Answer: 3 Input Image <think> Okay, let's see
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 7dc0860c-e552-4caf-88ab-d4217fe9ea96 · outbound
SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models To find the value of , I'll substitute the coordinates of the point into the equation
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6d55bae2-aee4-497b-a2dd-16e25d5c1b82 · outbound
SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models Unresolved cited work
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8f991278-9cf0-436c-8ba8-c5ec09e9d91b · outbound
SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models Unresolved cited work
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e8ab3baf-dc7d-47a8-85f8-4e8f74c3061a · outbound
SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models Unresolved cited work
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 585d2d22-91e7-483f-91d3-8f3dc50cc1d5 · outbound
SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models Unresolved cited work
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 649384af-e3c0-4f3c-9363-7e67250ae337 · outbound
SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models How many objects are left? •Original Answer: 3 Input Image <think> Okay, let's see
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 663300b5-aaeb-4403-ae12-0634e5dbdc59 · inbound
OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6388e0e0-b733-4138-9ed3-b930cdf92b31 · inbound
GRIT: Teaching MLLMs to Think with Images SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation edeabbe8-7b09-4e80-bc84-dfeaafec1ec7 · inbound
Blending Supervised and Reinforcement Fine-Tuning with Prefix Sampling SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ced92bbf-66a5-4550-aaa8-51ebd37cd72b · inbound
WebSailor: Navigating Super-human Reasoning for Web Agent SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 21777939-f9bb-4490-ade6-fcc4ea013ce3 · inbound
A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models
Reference 240
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 52ab568b-7125-4c5f-8753-0f7549bbe07d · inbound
MathReal: We Keep It Real! A Real Scene Benchmark for Evaluating Math Reasoning in Multimodal Large Language Models SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 584c915c-2f4b-4f5e-abde-3e2cf2752b2d · inbound
UAV-VL-R1: Generalizing Vision-Language Models via Supervised Fine-Tuning and Multi-Stage GRPO for UAV Visual Reasoning SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4f8c97d2-a1cf-4e78-8df9-dba7bdfd113c · inbound
Reinforced Visual Perception with Tools SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd903724-f451-48e7-ab85-592fedc27b69 · inbound
Less is More Tokens: Efficient Math Reasoning via Difficulty-Aware Chain-of-Thought Distillation SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93596ae4-2186-456b-8029-e254c01e1a43 · inbound
Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe009cb9-2b50-4598-83d3-acbabef4f6b3 · inbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1d1bde5-9bfa-410f-af95-ce1fe03e63c7 · inbound
Failure Makes the Agent Stronger: Enhancing Accuracy through Structured Reflection for Reliable Tool Interactions SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d4b59fc0-9878-4ddd-876b-91f6ccbf08e0 · inbound
Mixture-of-Visual-Thoughts: Exploring Context-Adaptive Reasoning Mode Selection for General Visual Reasoning SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9b8dd9e8-3217-48f5-9044-2713a86776b4 · inbound
Dynamic-TreeRPO: Breaking the Independent Trajectory Bottleneck with Structured Sampling SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2835058f-2941-4d99-9c7d-e98982c5b5c9 · inbound
Latent Visual Reasoning SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8c171be8-466d-4c50-a6b1-6e4562419737 · inbound
VOLD: Reasoning Transfer from LLMs to Vision-Language Models via On-Policy Distillation SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94162f6f-42a4-47c7-84a6-2e99cc85a321 · inbound
DeepEyesV2: Toward Agentic Multimodal Model SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 227654e6-ca34-4b57-9ab6-70a07116a4b7 · inbound
Reinforcement Learning Improves Traversal of Parametric Knowledge in LLMs SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c20dbad-0a8a-46b7-bf7b-cb81c16b9fea · inbound
MultiPriv: Benchmarking Individual-Level Privacy Reasoning in Vision-Language Models SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f0c7d33-ee0b-4908-8407-12343980bcf5 · inbound
TeamPath: Building MultiModal Pathology Experts with Reasoning AI Copilots SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ddf35aeb-ab0b-4711-8dff-c7247bb83273 · inbound
Distilling Counterfactual Reasoning from Language to Vision: Causal Graph Guided Post-Training for Video Understanding SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation feca519a-fadb-41ef-bbfc-b48c1c8ccc14 · inbound
Asking like Socrates: Socrates helps VLMs understand remote sensing images SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 88674871-a74a-4b1b-8776-512f7945ef01 · inbound
Reasoning Within the Mind: Dynamic Multimodal Interleaving in Latent Space SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9cf111ae-be03-4f32-b561-3108f6f4307d · inbound
Mimic Human Cognition, Master Multi-Image Reasoning: A Meta-Action Framework for Enhanced Visual Understanding SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 823cb637-64b5-4a87-90d6-d521c9eadfcb · inbound
ReMoT: Reinforcement Learning with Motion Contrast Triplets SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c70ed5cd-9388-4a7e-92ce-e6a59593d17f · inbound
Teaching an Agent to Sketch One Part at a Time SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6e169cf3-e0d2-41a3-b71a-0e8a488fd26a · inbound
Understanding the Role of Hallucination in Reinforcement Post-Training of Multimodal Reasoning Models SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9b8e4736-2200-4c0d-8ff0-d93c1f1727df · inbound
Watch Before You Answer: Learning from Visually Grounded Post-Training SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 53641dc4-bc0b-471e-96dc-847ae3396003 · inbound
Act Wisely: Cultivating Meta-Cognitive Tool Use in Agentic Multimodal Models SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 63c54e0e-f8e2-44a0-8c55-0f24920ef4a1 · inbound
Generalization in LLM Problem Solving: The Case of the Shortest Path SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 7c852431-d69f-4047-8c7c-257aa0598ff3 · inbound
Characterizing Model-Native Skills SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 69ca1e44-4ed7-4c75-96e5-13fe8d1825ff · inbound
One Step Forward and K Steps Back: Better Reasoning with Denoising Recursion Models SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models
Reference 144
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b1f6c7e3-8b8c-4ebb-a27d-2f6cb7d31507 · inbound
CGC: Compositional Grounded Contrast for Fine-Grained Multi-Image Understanding SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e8051922-815f-4367-8c4b-300a29563b98 · inbound
How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f2176436-10eb-4ea0-bbf8-da2b115a9287 · inbound
Learning Agentic Policy from Action Guidance SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 11233d49-b0da-4cfd-9cd2-a9385544f928 · inbound
Teacher-Guided Policy Optimization for On-Policy Reasoning Distillation under Large Policy Divergence SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b98e07f9-c4ac-4ba1-9170-79703cd4b01f · inbound
Are Tools Always Beneficial? Learning to Invoke Tools Adaptively for Dual-Mode Multimodal LLM Reasoning SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9266a140-c721-4919-ad30-99c840ebd9e8 · inbound
RISE: Reliable Improvement in Self-Evolving Vision-Language Models SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b0cb786d-7ac9-4b0e-8b91-a44e05412304 · inbound
RISE: Reliable Improvement in Self-Evolving Vision-Language Models SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3785af52-345f-441d-a180-2d76d5c2776b · inbound
AnE: Pushing the Reasoning Frontier of Multimodal LLMs via Anchor Evolution SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 30f6306f-c6fc-4380-a466-508f6fcaa267 · inbound
ToolFG: Towards Well-Grounded Fine-Grained Image Classification SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a4d67d8c-9984-4a16-86f5-a83c116e52e9 · inbound
Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1826d3fd-d7dc-49db-a281-4d2ef82897ab · inbound
Beyond Trajectory Imitation: Strategy-Guided Policy Optimization for LLM Reasoning SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6c48bdaf-3f3a-4b4a-b366-ee810459815d · inbound
Bridging Physical Reasoning and Task Generalization via Visual Action Outcome Reasoning Alignment SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 757dbb39-9cbf-49c1-b8e5-010dd6e76ef0 · inbound
BUS: Brain-Inspired Unsupervised Self-Reflection via Backward Prediction for Multimodal Reasoning SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models
Reference 110
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ac7f4469-3187-48a0-ae8c-a5b8904426d0 · inbound
BUS: Brain-Inspired Unsupervised Self-Reflection via Backward Prediction for Multimodal Reasoning SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models
Reference 110
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b2f3b5e-95e9-43e5-89b4-e24edf9eb807 · inbound
SIVA-RL: Sensitivity-Invariance Visual Alignment for Multimodal Reinforcement Learning SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69641fed-a3c3-4135-b179-c2f374505a6a · inbound
Reasoning to Regulate: Chain-of-Thought for Traffic Rule Understanding SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf4436d1-3642-4344-af48-54e7b424dc8c · inbound
Knowing When to Quit: Diagnosing and Training LLMs to Abort Futile Reasoning SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 701a8f4f-69ed-42a1-9375-ccd59810fc97 · inbound
Balancing Efficiency and Efficacy: Training-Free Attention-Guided Switching Between Explicit and Latent Thoughts for MLLMs SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.