Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:33:14.622826Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 12 inbound Pith citation observations for arXiv:2505.14810.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:33:14.622826Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T19:46:47.158980Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T01:27:30.780707Z
54 of 54 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 44568d9e-23c9-485f-8e98-0824b2d06ecf · outbound
Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models A survey of efficient reasoning for large reasoning models: Language, multimodality, and beyond.arXiv preprint arXiv:2503.21614, 2025
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65eaa0d6-7f87-479e-9758-3cc235820994 · outbound
Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models Introducing openai o3 and o4-mini
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 47a35041-50e0-4fdc-8ffb-97c75c48fe23 · outbound
Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c7a2c4f-99c9-4c9a-ae76-32bfbf88e23c · outbound
Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models Unresolved cited work
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b26cc8b-3281-4877-8961-b99584ef6053 · outbound
Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models Olympiadbench: A challenging benchmark for promoting agi with olympiad-level bilingual multimodal scientific problems, 2024
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9b5b964-b0ef-4d25-b0a2-003cf15c219e · outbound
Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models Measuring Mathematical Problem Solving With the MATH Dataset
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b153dfa-26f1-4809-be95-e2b9b1e63776 · outbound
Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models Aime problem set 1983-2024, 2023
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2dd24725-91cb-47cd-b3a2-efc412a34f00 · outbound
Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models Unresolved cited work
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18bd3729-b3b4-4c53-b40a-9fd402adf3d2 · outbound
Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models Chain-of-thought prompting elicits reasoning in large language models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 242ded00-ce83-493d-b0a4-cd4920915107 · outbound
Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models Crossing the reward bridge: Expanding rl with verifiable rewards across diverse domains, 2025
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f6dd2e34-7add-447c-b744-8234be23c9b3 · outbound
Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models A survey on llm-as-a-judge, 2025
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a33b4e24-dc98-45f0-a3e2-f428740bf069 · outbound
Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models Instruction-Following Evaluation for Large Language Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e794753-fe77-449e-8d01-208d32f78d11 · outbound
Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5964b2cc-872b-4985-9287-f2872170f483 · outbound
Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models s1: Simple test-time scaling
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb7d4fd3-53a1-4468-a2f1-8d85f9fd58e9 · outbound
Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models Limo: Less is more for reasoning, 2025
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d988746a-3274-4130-af2b-b319f12ac6f4 · outbound
Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models Demystifying long chain-of-thought reasoning in llms, 2025
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27eceab6-b648-490e-8d28-ce69017e3b3d · outbound
Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models SFT memorizes, RL generalizes: A comparative study of foundation model post-training
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f344a563-8832-4594-b617-c98677f4577a · outbound
Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models There may not be aha moment in r1-zero-like training — a pilot study
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a25740f0-4537-4d6f-8f54-d60862085901 · outbound
Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d772c21-66fd-496d-8fcd-fbe6ccde9a3d · outbound
Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models Process Reinforcement through Implicit Rewards
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92ce5b9b-bf31-467e-aafe-b599b15e1fdb · outbound
Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models Learning to reason under off-policy guidance, 2025
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de45736f-69f1-464f-8f18-338b7dba298d · outbound
Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models Thinking Preference Optimization
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0162cf0f-ef5e-40e7-bfb3-9c06587da67c · outbound
Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models Hashimoto
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a56b2ab7-afc0-4d9c-8d4e-716add35b17c · outbound
Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models Gonzalez, Ion Stoica, and Eric P
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd55063f-3abd-499a-95a8-4b6a268df8c6 · outbound
Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models FOFO: A benchmark to evaluate LLMs’ format-following capability
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e94cee61-4939-4de0-b284-97e4471d10c5 · outbound
Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models Unresolved cited work
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8bb70e02-d4e0-48a9-bcff-a8cc018a5b53 · outbound
Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models Multi-IF: Benchmarking LLMs on Multi-Turn and Multilingual Instructions Following
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17fbfcef-6d59-4686-aad8-7d66ebad32a5 · outbound
Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models StructFlowBench: A Structured Flow Benchmark for Multi-turn Instruction Following
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 022c8737-095b-40c0-96c4-b31b2ec9154c · outbound
Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models Can language models follow multiple turns of entangled instructions?arXiv preprint arXiv:2503.13222, 2025
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e46e7b3-0481-483d-8932-7b2ab6cebf05 · outbound
Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28e29008-b9b2-4a36-879c-89128bc6f3d3 · outbound
Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models LIFBench: Evaluating the Instruction Following Performance and Stability of Large Language Models in Long-Context Scenarios
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41151d93-1676-4dfb-97cf-ecc2cea02dcc · outbound
Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models Xifbench: Evaluating large language models on multilingual instruction following
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cbd99d66-e41f-4dc3-ac5d-80aa18f35783 · outbound
Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models IHEval: Evaluating language models on following the instruction hierarchy
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 53abe998-55b1-4760-922b-e120406ac5de · outbound
Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models Chain-of-instructions: Compositional instruction tuning on large language models
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 53656fb6-10c5-43f2-b5c3-cdc3f7ae0f64 · outbound
Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models RefuteBench: Evaluating refuting instruction-following for large language models
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a81c3877-1496-43b9-8e72-46959fe21fe1 · outbound
Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models Refutebench 2.0 – agentic benchmark for dynamic evaluation of llm responses to refutation instruction, 2025
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1dc49d07-536f-49d5-af35-f5ed750fec17 · outbound
Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models Benchmarking complex instruction-following with multiple constraints composition.Advances in Neural Information Processing Systems, 37:137610– 137645, 2024
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c15a7049-461a-4614-9ebf-2a802448278b · outbound
Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models GPT-4o System Card
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10907a57-b443-4eff-8db6-2aa6a0c68a11 · outbound
Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models Training Verifiers to Solve Math Word Problems
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3db8243-d059-40b1-8ece-eb9e9b71d518 · outbound
Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models Minerva: Accelerating data analysis in next-generation ssds
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 81bdc687-70e6-4f71-af8d-29c061b47e1a · outbound
Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models Gonzalez, Hao Zhang, and Ion Stoica
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11a41faa-a906-4367-b22e-7db1df5f14e5 · outbound
Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models Qwen3, April 2025
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55f2cf9e-82ae-4bbc-9a03-aa64511f35d8 · outbound
Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79fa9c6b-3bd0-4099-9b08-824846a7d1f1 · outbound
Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f49f2bd5-2c6d-4f20-8d8d-a2dddf292488 · outbound
Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Li Erran Li, Raluca Ada Popa, and Ion Stoica
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dcec4b4d-7758-4c3d-9930-0c4c215f67e4 · outbound
Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0cac9504-c05f-4349-a0ec-c1b24cc81c41 · outbound
Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6dc24a3-efa5-479b-adbe-fac5ed3701fe · outbound
Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models Open r1: A fully open reproduction of deepseek-r1, January 2025
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f43c058a-ad96-4a97-8379-471d5590d1fc · outbound
Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models Qwq-32b: Embracing the power of reinforcement learning, March 2025
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79f0a2f6-98d9-45a0-ac80-fb90fb618d74 · outbound
Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models Open-reasoner-zero: An open source approach to scaling up reinforcement learning on the base model, 2025
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c73eeb6f-0dd0-42b2-85bc-58f3076f5df9 · outbound
Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063, 2024
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2bd12cc8-133b-4fd6-8560-9ffb777bc94e · outbound
Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ebf27aa1-fe25-4327-a976-0eac331754cd · outbound
Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models The impact of reasoning step length on large language models
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2ff083fc-448b-4229-a8eb-0199ed4cec8c · outbound
Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models Unresolved cited work
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e8854ce-8f38-4ae7-9e39-60f8befb076a · inbound
Learning to Reason under Off-Policy Guidance Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cd451d3a-100f-4632-acb7-9f442874c142 · inbound
Activation Steering for Chain-of-Thought Compression Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87a3f983-2e7c-44fb-b190-219dd66e7440 · inbound
Towards Concise and Adaptive Thinking in Large Reasoning Models: A Survey Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4143950-2e76-456e-a840-66278d750e9e · inbound
AutoTIR: Autonomous Tools Integrated Reasoning via Reinforcement Learning Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df7b24f1-6691-41a3-8471-0b59b258279a · inbound
A Survey of Reinforcement Learning for Large Reasoning Models Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models
Reference 146
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1244e1a4-09af-45f5-bc65-927153f45db9 · inbound
Retrieval-of-Thought: Efficient Reasoning via Reusing Thoughts Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cd80a8e1-9df3-4476-a93e-309dba05edbc · inbound
Position: The Hidden Costs and Measurement Gaps of Reinforcement Learning with Verifiable Rewards Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5b236504-6262-4cd1-b48f-05ee07634c87 · inbound
Reasoning Up the Instruction Ladder for Controllable Language Models Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 035981a2-0272-45d4-abc0-f931a2a19ce8 · inbound
ImpRIF: Stronger Implicit Reasoning Leads to Better Complex Instruction Following Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5d43761b-0d0d-4906-a652-117b55d16885 · inbound
Expert-Aware Refusal Steering Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a5be2a59-61f1-4517-b01d-8da2e092dabe · inbound
When Built-in Thinking Helps and Hurts: Constraint-Level Error Shifts in Instruction Following Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f2244e3d-ae28-4746-96ff-a48ba458ff32 · inbound
Structured Thoughts For Improved Reasoning And Context Pruning Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.