Pith. sign in

Paper Citation Record · LEDGER

Switch-Reasoner: Learn When to Think in Multitask Mixtures via Reinforcement Learning

As of 10 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 0 inbound Pith citation observations for arXiv:2607.08572.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.08572 v1

Coverage vector

measured 25 of 25 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-10T05:07:40.106099Z

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

25 of 25 outbound references displayed

  • verified exact20
  • verified fuzzy1
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch4

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 75206606-2d14-4ee8-9766-409d162a943f · outbound

This paper cites Qwen3-VL Technical Report.

Switch-Reasoner: Learn When to Think in Multitask Mixtures via Reinforcement Learning Qwen3-VL Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-10T05:16:48.114202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T05:07:40.106099Z digest=sha256:99d7fef9c297f99b9685a557ce595d9d5755864597eac8809ce9c7830c1cd416

Observation f47b38ad-c95f-4e43-9eb5-56ae2e3c5701 · outbound

This paper cites Ares: Multimodal adaptive reasoning via difficulty-aware token-level entropy shaping.

Switch-Reasoner: Learn When to Think in Multitask Mixtures via Reinforcement Learning Ares: Multimodal adaptive reasoning via difficulty-aware token-level entropy shaping

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-10T05:16:48.117346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T05:07:40.106099Z digest=sha256:894f716a52906c4890a2201185e077861884ba48c0e4430954b19ab6793c935a

Observation cb889607-63a3-4f79-8594-f425090b6a7c · outbound

This paper cites Video-R1: Reinforcing Video Reasoning in MLLMs.

Switch-Reasoner: Learn When to Think in Multitask Mixtures via Reinforcement Learning Video-R1: Reinforcing Video Reasoning in MLLMs

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T05:16:48.127937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T05:07:40.106099Z digest=sha256:961be473add660e12335b164e8168411377f90328c3fc83816acb02e380d59d9

Observation 73c4187c-fa4e-477a-9202-0e1a3dd727f4 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Switch-Reasoner: Learn When to Think in Multitask Mixtures via Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-10T05:16:48.103262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T05:07:40.106099Z digest=sha256:ae43a9c3ace452fd175f8503955587eb93bbf461d3bac4a00070187adf46c208

Observation 0596a9b0-fd0d-41ec-9eaa-bc89c8382b7e · outbound

This paper cites ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning.

Switch-Reasoner: Learn When to Think in Multitask Mixtures via Reinforcement Learning ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-07-10T05:16:48.069730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T05:07:40.106099Z digest=sha256:0bfff9d3858c010f82c461154a4595c10afbe93d27b8ca5b6c9c45f7a803e534

Observation acad75f2-78db-4485-91aa-9a6d2cca548a · outbound

This paper cites Boosting MLLM Reasoning with Text-Debiased Hint-GRPO.

Switch-Reasoner: Learn When to Think in Multitask Mixtures via Reinforcement Learning Boosting MLLM Reasoning with Text-Debiased Hint-GRPO

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T05:16:48.100099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T05:07:40.106099Z digest=sha256:9fc6083bd31df856d94c49b4162211b802b6771116d8a130898367647ad47b7c

Observation 6182bec2-599a-4f4a-aaa2-ab55463b8b88 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Switch-Reasoner: Learn When to Think in Multitask Mixtures via Reinforcement Learning LLaVA-OneVision: Easy Visual Task Transfer

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-07-10T05:16:48.133413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T05:07:40.106099Z digest=sha256:22196e947338a0061e8418e4b05c187feddb4edfff9a9bf1d6853fa980063f30

Observation 3a2061e6-be90-4961-b661-0e466d79821e · outbound

This paper cites Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle.

Switch-Reasoner: Learn When to Think in Multitask Mixtures via Reinforcement Learning Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-10T05:16:48.097060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T05:07:40.106099Z digest=sha256:41acd38fe787773eee70e44ecb9eed59b47242fd731aba51c1f40d5923a77840

Observation 445bfd8a-f823-46ed-a448-fdd75fb93c32 · outbound

This paper cites Visual-RFT: Visual Reinforcement Fine-Tuning.

Switch-Reasoner: Learn When to Think in Multitask Mixtures via Reinforcement Learning Visual-RFT: Visual Reinforcement Fine-Tuning

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-07-10T05:16:48.105897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T05:07:40.106099Z digest=sha256:a057d9d4510def3f494adef64dba440228ec8828b5eafd7c9859775a40bcde26

Observation 60df4eee-bb17-4688-a3ee-162763452667 · outbound

This paper cites IconQA: A New Benchmark for Abstract Diagram Understanding and Visual Language Reasoning.

Switch-Reasoner: Learn When to Think in Multitask Mixtures via Reinforcement Learning IconQA: A New Benchmark for Abstract Diagram Understanding and Visual Language Reasoning

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-10T05:16:48.108859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T05:07:40.106099Z digest=sha256:5dbf8e99210d91274885b59ae72fbd750a299cfc778b8ce905d5c75b22eda1ae

Observation 23937e72-3d26-4014-8c13-480738d92280 · outbound

This paper cites TERMINATOR: Learning Optimal Exit Points for Early Stopping in Chain-of-Thought Reasoning.

Switch-Reasoner: Learn When to Think in Multitask Mixtures via Reinforcement Learning TERMINATOR: Learning Optimal Exit Points for Early Stopping in Chain-of-Thought Reasoning

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-10T05:16:48.111617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T05:07:40.106099Z digest=sha256:f6de4f0becc1360a1a0f1811c8964ed36fd121e160fc3ecdadb8d500c84303cb

Observation 32d2dd63-4520-40eb-93b5-0f6064373c65 · outbound

This paper cites Safegrpo: Self-rewarded multimodal safety alignment via rule-governed policy optimization.arXiv preprint arXiv:2511.12982,.

Switch-Reasoner: Learn When to Think in Multitask Mixtures via Reinforcement Learning Safegrpo: Self-rewarded multimodal safety alignment via rule-governed policy optimization.arXiv preprint arXiv:2511.12982,

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-10T05:16:48.120378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T05:07:40.106099Z digest=sha256:0ffd8bfc4bc454b4b23b1711f2fac2a8bf71ee9a1ebc50f37b281691143ee798

Observation 6f8f9413-c154-4aec-8677-78b768dce8ad · outbound

This paper cites Proximal Policy Optimization Algorithms.

Switch-Reasoner: Learn When to Think in Multitask Mixtures via Reinforcement Learning Proximal Policy Optimization Algorithms

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-10T05:16:48.125440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T05:07:40.106099Z digest=sha256:9076003349dcb74c020c0333f09d43adc98a368e1997be8e4de2ad7f11ea4b8b

Observation 85f4df1b-6102-4138-b5ce-8e10c571ab83 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Switch-Reasoner: Learn When to Think in Multitask Mixtures via Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-10T05:16:48.136244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T05:07:40.106099Z digest=sha256:e70286fb3ae65d9bdf4810c61780bdfbbdd1bf09f9b1ab98cd9fae78cf00f639

Observation caec8aa2-2cc6-4a14-8bfc-28456e20a3fa · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Switch-Reasoner: Learn When to Think in Multitask Mixtures via Reinforcement Learning Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T05:16:48.130529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T05:07:40.106099Z digest=sha256:2e2183421cf0204ecceea16fca89f0bdf2fd724515378f50e678e63a126ae5b9

Observation cbdb5770-6e00-4341-98d8-226e05c37194 · outbound

This paper cites AdaptR1: Reinforcement Learning Based Adaptive Interleaved Thinking in Multi-hop Question Answering.

Switch-Reasoner: Learn When to Think in Multitask Mixtures via Reinforcement Learning AdaptR1: Reinforcement Learning Based Adaptive Interleaved Thinking in Multi-hop Question Answering

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-10T05:16:48.091457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T05:07:40.106099Z digest=sha256:20d18e1fbd72af3bdb2edc7fdddcd04e15666614d915c5a72c316d8b432cde24

Observation 5e486c2b-ec80-4531-a2e7-5dde2d114f28 · outbound

This paper cites When More is Less: Understanding Chain-of-Thought Length in LLMs.

Switch-Reasoner: Learn When to Think in Multitask Mixtures via Reinforcement Learning When More is Less: Understanding Chain-of-Thought Length in LLMs

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-07-10T05:16:48.094170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T05:07:40.106099Z digest=sha256:6c3bb1d281f1dcb5f550db22ec11a50b578d2f794ee8d0fa962eacc218b1e2a9

Observation ede9c4e3-60de-4973-8663-4d65770c8b26 · outbound

This paper cites Dynamic early exit in reasoning models.arXiv preprint arXiv:2504.15895, 2025a.

Switch-Reasoner: Learn When to Think in Multitask Mixtures via Reinforcement Learning Dynamic early exit in reasoning models.arXiv preprint arXiv:2504.15895, 2025a

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-07-10T05:16:48.085721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T05:07:40.106099Z digest=sha256:1463478eadbbcd12a83132d6179dae811b872ab92ac085c8b6323c9c90c90e29

Observation 172c5265-982d-45bd-88fc-ec64add2b199 · outbound

This paper cites R-4B: Incentivizing General-Purpose Auto-Thinking Capability in MLLMs via Bi-Mode Annealing and Reinforce Learning.

Switch-Reasoner: Learn When to Think in Multitask Mixtures via Reinforcement Learning R-4B: Incentivizing General-Purpose Auto-Thinking Capability in MLLMs via Bi-Mode Annealing and Reinforce Learning

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-07-10T05:16:48.088969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T05:07:40.106099Z digest=sha256:61818b11571be99b2eb6964898e3f0f92b16564eb6f3d8aab45217eb1bf44a9e

Observation 86302e5a-d931-4bed-a8b0-c36aaa99e3f9 · outbound

This paper cites R1-ShareVL: Incentivizing Reasoning Capability of Multimodal Large Language Models via Share-GRPO.

Switch-Reasoner: Learn When to Think in Multitask Mixtures via Reinforcement Learning R1-ShareVL: Incentivizing Reasoning Capability of Multimodal Large Language Models via Share-GRPO

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-07-10T05:16:48.067239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T05:07:40.106099Z digest=sha256:fa589f3e2a15e5eeeed7277c5d1e1cf3903c9aeb663b681564286a3f686bce7e

Observation 03dc2a73-379d-46d3-b1bd-3180698589b6 · outbound

This paper cites A Survey of Safety on Large Vision-Language Models: Attacks, Defenses and Evaluations.

Switch-Reasoner: Learn When to Think in Multitask Mixtures via Reinforcement Learning A Survey of Safety on Large Vision-Language Models: Attacks, Defenses and Evaluations

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-07-10T05:16:48.123078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T05:07:40.106099Z digest=sha256:6aa36b139e46b3a4a5f1025694dab13e2c75b5f2108d0277654e1931bf2c9976

Observation 1a8593ca-a459-4bfa-bb0a-a2a8f5dcbca8 · outbound

This paper cites R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization.

Switch-Reasoner: Learn When to Think in Multitask Mixtures via Reinforcement Learning R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization

Reference 22

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T05:16:48.081912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T05:07:40.106099Z digest=sha256:afdae4dd779513756501bc4b7dfdfcaffb019853da95f84beb1d7b369577909c

Observation e61c7e6f-ad04-4b8a-a5a2-404dba7ec733 · outbound

This paper cites Reinforced MLLM: A Survey on RL-Based Reasoning in Multimodal Large Language Models.

Switch-Reasoner: Learn When to Think in Multitask Mixtures via Reinforcement Learning Reinforced MLLM: A Survey on RL-Based Reasoning in Multimodal Large Language Models

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-07-10T05:16:48.079134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T05:07:40.106099Z digest=sha256:98285be180bbc67d747d94c5f6f58512a34a12257dfcba93845728e925bc3592

Observation dc21e25c-2802-4e4a-977c-f18ea9d45bc3 · outbound

This paper cites Think in Blocks: Adaptive Reasoning from Direct Response to Deep Reasoning.

Switch-Reasoner: Learn When to Think in Multitask Mixtures via Reinforcement Learning Think in Blocks: Adaptive Reasoning from Direct Response to Deep Reasoning

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-07-10T05:16:48.076132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T05:07:40.106099Z digest=sha256:e93133c57384d02f5586c4b7d240f35d76813252f0fde7e2221bbef267b11196

Observation e9edec57-32d4-4589-99aa-0d0537938e3c · outbound

This paper cites Mathe- matics includes Geometry3k and MathVista; Chart/Doc comprises ChartQA and DocVQA; Ground- ing consists of RefAdv; and the remaining benchmarks are grouped under General.

Switch-Reasoner: Learn When to Think in Multitask Mixtures via Reinforcement Learning Mathe- matics includes Geometry3k and MathVista; Chart/Doc comprises ChartQA and DocVQA; Ground- ing consists of RefAdv; and the remaining benchmarks are grouped under General

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T05:16:48.395291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T05:07:40.106099Z digest=sha256:d3a197ffabf28f81517632a27cdb1aa62de1919e84c32f6b9913bd54b70b7534

Pith citing papers

No inbound Pith citation observations are available.