Pith. sign in

Paper Citation Record · LEDGER

Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models

As of 7 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 12 inbound Pith citation observations for arXiv:2505.14810.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.14810 v2

Coverage vector

measured 54 of 54 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:33:14.622826Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:46:47.158980Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T01:27:30.780707Z

Reference resolution

54 of 54 outbound references displayed

  • verified exact0
  • verified fuzzy8
  • unresolved45
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 44568d9e-23c9-485f-8e98-0824b2d06ecf · outbound

This paper cites A survey of efficient reasoning for large reasoning models: Language, multimodality, and beyond.arXiv preprint arXiv:2503.21614, 2025.

Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models A survey of efficient reasoning for large reasoning models: Language, multimodality, and beyond.arXiv preprint arXiv:2503.21614, 2025

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:10.157381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:33:10.157381Z digest=sha256:ccb8d27e79297ef21756c6b18501e14713f1f38ae3962e527ca6807b03a2fd5d

Observation 65eaa0d6-7f87-479e-9758-3cc235820994 · outbound

This paper cites Introducing openai o3 and o4-mini.

Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models Introducing openai o3 and o4-mini

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:33:17.316784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:33:10.208482Z digest=sha256:ee8e9f9fd2cfea303356ff9a54351127738945acafc517c15d9bf03577ae830e

Observation 47a35041-50e0-4fdc-8ffb-97c75c48fe23 · outbound

This paper cites Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025.

Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:10.246442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:33:10.246442Z digest=sha256:31db064196220c1e74f9b916c16e6de7b28e7549707353c590bdd58ad3cd65d4

Observation 2c7a2c4f-99c9-4c9a-ae76-32bfbf88e23c · outbound

This paper cites an unresolved cited work.

Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:10.307082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:33:10.307082Z digest=sha256:4eb463cbef10cf470f5d4eb172a287a13244403a2c5ffc9c231f2003cf6f0725

Observation 5b26cc8b-3281-4877-8961-b99584ef6053 · outbound

This paper cites Olympiadbench: A challenging benchmark for promoting agi with olympiad-level bilingual multimodal scientific problems, 2024.

Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models Olympiadbench: A challenging benchmark for promoting agi with olympiad-level bilingual multimodal scientific problems, 2024

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:10.370272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:33:10.370272Z digest=sha256:4a6ab74a18a2cf63621df3ee6064e8356b328f9ff44b54e87ad0c9ca69d9ee1f

Observation d9b5b964-b0ef-4d25-b0a2-003cf15c219e · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models Measuring Mathematical Problem Solving With the MATH Dataset

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:10.452157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:33:10.452157Z digest=sha256:e8e74e55be70f5619549f1f4aad85d086b3ee5dcd6387ccdb5626b0426018597

Observation 7b153dfa-26f1-4809-be95-e2b9b1e63776 · outbound

This paper cites Aime problem set 1983-2024, 2023.

Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models Aime problem set 1983-2024, 2023

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:10.518270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:33:10.518270Z digest=sha256:335f7938f0c1e82dae1ace7e2f768591489bf46dc980db04e7ae5e96cfa7bba7

Observation 2dd24725-91cb-47cd-b3a2-efc412a34f00 · outbound

This paper cites an unresolved cited work.

Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:10.568995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:33:10.568995Z digest=sha256:ad3df0b79baddfa0f3d1a8b65817a6d2fad9e237dc823e7e7e338c32f51d5f1c

Observation 18bd3729-b3b4-4c53-b40a-9fd402adf3d2 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models Chain-of-thought prompting elicits reasoning in large language models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:10.639989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:33:10.639989Z digest=sha256:4bcff9129663f6d6e41685a0e5a8fd9987e39bc64c82527aa0275e8241c543c7

Observation 242ded00-ce83-493d-b0a4-cd4920915107 · outbound

This paper cites Crossing the reward bridge: Expanding rl with verifiable rewards across diverse domains, 2025.

Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models Crossing the reward bridge: Expanding rl with verifiable rewards across diverse domains, 2025

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:33:17.080686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:33:10.722778Z digest=sha256:ca6e958dc1f7bead36637e5a2cb6b862ed458accf2982bb6be42a94589f3e213

Observation f6dd2e34-7add-447c-b744-8234be23c9b3 · outbound

This paper cites A survey on llm-as-a-judge, 2025.

Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models A survey on llm-as-a-judge, 2025

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:10.783973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:33:10.783973Z digest=sha256:d365d2fae8b291c1c01937d4a34fc8612157f2e53fedc84b13a7faf73aefc651

Observation a33b4e24-dc98-45f0-a3e2-f428740bf069 · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models Instruction-Following Evaluation for Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:10.861362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:33:10.861362Z digest=sha256:9ea79be180af4be6c8a5c824615d596917c6c5e7d8a41cf8fee506d3b6adfec7

Observation 6e794753-fe77-449e-8d01-208d32f78d11 · outbound

This paper cites FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models.

Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:10.940404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:33:10.940404Z digest=sha256:eb809748ae73a5c6f881703d822e540b39e23c50d43a3ecd9685af4534e6d46c

Observation 5964b2cc-872b-4985-9287-f2872170f483 · outbound

This paper cites s1: Simple test-time scaling.

Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models s1: Simple test-time scaling

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:10.981211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:33:10.981211Z digest=sha256:447dd27a467a97547a955d285be6843650667f8ed9902c6c4b66a09ed2ef7ac7

Observation bb7d4fd3-53a1-4468-a2f1-8d85f9fd58e9 · outbound

This paper cites Limo: Less is more for reasoning, 2025.

Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models Limo: Less is more for reasoning, 2025

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:11.057906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:33:11.057906Z digest=sha256:66ca7ce9c107b479ce514a097a2cd626d1d87dabdfd63475ac4ff080c9baa8c6

Observation d988746a-3274-4130-af2b-b319f12ac6f4 · outbound

This paper cites Demystifying long chain-of-thought reasoning in llms, 2025.

Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models Demystifying long chain-of-thought reasoning in llms, 2025

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:11.148146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:33:11.148146Z digest=sha256:7b7d74f257d999f94343f8567b6ecbc922f778a911682590dc1ca31aabbdf631

Observation 27eceab6-b648-490e-8d28-ce69017e3b3d · outbound

This paper cites SFT memorizes, RL generalizes: A comparative study of foundation model post-training.

Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models SFT memorizes, RL generalizes: A comparative study of foundation model post-training

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:11.225305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:33:11.225305Z digest=sha256:a6455169b7a11cbd8850125ed0c2f2a8fa4ce1eec0786e7b8ea6d69224b7c18a

Observation f344a563-8832-4594-b617-c98677f4577a · outbound

This paper cites There may not be aha moment in r1-zero-like training — a pilot study.

Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models There may not be aha moment in r1-zero-like training — a pilot study

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:11.311589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:33:11.311589Z digest=sha256:1efcc5ca18bca620ae396fc8fb8fb5c6f0388767d1b38c293f1402446dc9ba16

Observation a25740f0-4537-4d6f-8f54-d60862085901 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:11.407955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:33:11.407955Z digest=sha256:7706df3a94c04a7e915d272e9165fd05f4e74896c6782fd73560ab7c4692e8f3

Observation 4d772c21-66fd-496d-8fcd-fbe6ccde9a3d · outbound

This paper cites Process Reinforcement through Implicit Rewards.

Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models Process Reinforcement through Implicit Rewards

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:11.486382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:33:11.486382Z digest=sha256:f11169c45206caf0cb223effc5f0ec2fa3ffb61a29dde3c998f2283c19f98ab0

Observation 92ce5b9b-bf31-467e-aafe-b599b15e1fdb · outbound

This paper cites Learning to reason under off-policy guidance, 2025.

Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models Learning to reason under off-policy guidance, 2025

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:11.560154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:33:11.560154Z digest=sha256:9f5f79d91cb951d088b9e034a2c6c604ca7f18f41b9f7a602b530cc0fdb69bbe

Observation de45736f-69f1-464f-8f18-338b7dba298d · outbound

This paper cites Thinking Preference Optimization.

Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models Thinking Preference Optimization

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:11.638080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:33:11.638080Z digest=sha256:5682106200245516183d3eb1902445b0fb5e56856befc572e6a91d51ab74852a

Observation 0162cf0f-ef5e-40e7-bfb3-9c06587da67c · outbound

This paper cites Hashimoto.

Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models Hashimoto

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:11.738840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:33:11.738840Z digest=sha256:ca330b040ab6905ac073802f4fa7893aebd29f9579f194c9af71a996f3a1e090

Observation a56b2ab7-afc0-4d9c-8d4e-716add35b17c · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models Gonzalez, Ion Stoica, and Eric P

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:11.809584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:33:11.809584Z digest=sha256:4e217000c7edf2e1d0f0863224d71e6d06924da3eda07948694feb5aecee85a8

Observation dd55063f-3abd-499a-95a8-4b6a268df8c6 · outbound

This paper cites FOFO: A benchmark to evaluate LLMs’ format-following capability.

Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models FOFO: A benchmark to evaluate LLMs’ format-following capability

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:33:16.800930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:33:11.922967Z digest=sha256:8c25aac7a21dd5812ad00755269fa622246927c2693c851064a0a4f46170f9be

Observation e94cee61-4939-4de0-b284-97e4471d10c5 · outbound

This paper cites an unresolved cited work.

Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:33:16.608320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:33:12.043697Z digest=sha256:6e95ab51186ae22518b49354315442534ffa34ac999ca13f6b84a781b3ae0a0e

Observation 8bb70e02-d4e0-48a9-bcff-a8cc018a5b53 · outbound

This paper cites Multi-IF: Benchmarking LLMs on Multi-Turn and Multilingual Instructions Following.

Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models Multi-IF: Benchmarking LLMs on Multi-Turn and Multilingual Instructions Following

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:12.154089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:33:12.154089Z digest=sha256:9bd390bd0284515531929973e07bb1b5b6340dffb2cd22290732fe423daeb9a7

Observation 17fbfcef-6d59-4686-aad8-7d66ebad32a5 · outbound

This paper cites StructFlowBench: A Structured Flow Benchmark for Multi-turn Instruction Following.

Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models StructFlowBench: A Structured Flow Benchmark for Multi-turn Instruction Following

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:12.228882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:33:12.228882Z digest=sha256:feb7fbf08a3768b7a8dcf33ebe1bcc85f9c71e2dee58b7a72796f07bd14b1998

Observation 022c8737-095b-40c0-96c4-b31b2ec9154c · outbound

This paper cites Can language models follow multiple turns of entangled instructions?arXiv preprint arXiv:2503.13222, 2025.

Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models Can language models follow multiple turns of entangled instructions?arXiv preprint arXiv:2503.13222, 2025

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:12.305782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:33:12.305782Z digest=sha256:4b914d2b7b53a9de4cf19b7dd8d875d62f5e904bb12f9719b9b23f1af1672579

Observation 6e46e7b3-0481-483d-8932-7b2ab6cebf05 · outbound

This paper cites MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs.

Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:12.420401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:33:12.420401Z digest=sha256:11546a24f4117d39335bbcacd8ade4a2a18d8b94bd50750b29aa8695767f7c9e

Observation 28e29008-b9b2-4a36-879c-89128bc6f3d3 · outbound

This paper cites LIFBench: Evaluating the Instruction Following Performance and Stability of Large Language Models in Long-Context Scenarios.

Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models LIFBench: Evaluating the Instruction Following Performance and Stability of Large Language Models in Long-Context Scenarios

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:12.524178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:33:12.524178Z digest=sha256:e55049cc3330162a66ccbade5ff71b951cbca5197eec8851edc1781f4535092b

Observation 41151d93-1676-4dfb-97cf-ecc2cea02dcc · outbound

This paper cites Xifbench: Evaluating large language models on multilingual instruction following.

Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models Xifbench: Evaluating large language models on multilingual instruction following

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:12.608031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:33:12.608031Z digest=sha256:953679cdf0797017fc75da054c686256d0cc5b8ca3e647f5aca949e72e7938cc

Observation cbd99d66-e41f-4dc3-ac5d-80aa18f35783 · outbound

This paper cites IHEval: Evaluating language models on following the instruction hierarchy.

Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models IHEval: Evaluating language models on following the instruction hierarchy

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:33:16.432287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:33:12.671400Z digest=sha256:4c6b745ffa80a04c65583485a023e81dc45c839d20f879eedece5fa36667c0f4

Observation 53abe998-55b1-4760-922b-e120406ac5de · outbound

This paper cites Chain-of-instructions: Compositional instruction tuning on large language models.

Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models Chain-of-instructions: Compositional instruction tuning on large language models

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:33:16.265814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:33:12.869124Z digest=sha256:11628dfaa373bdffdd5364635bf14652c481e8aa29bc97d9dd68b29ef351f071

Observation 53656fb6-10c5-43f2-b5c3-cdc3f7ae0f64 · outbound

This paper cites RefuteBench: Evaluating refuting instruction-following for large language models.

Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models RefuteBench: Evaluating refuting instruction-following for large language models

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:33:16.093504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:33:12.970356Z digest=sha256:b25fd6cfdc0bd14db5805d164f72e1850dd75f6ab23dd475a5f7e27d12f046ce

Observation a81c3877-1496-43b9-8e72-46959fe21fe1 · outbound

This paper cites Refutebench 2.0 – agentic benchmark for dynamic evaluation of llm responses to refutation instruction, 2025.

Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models Refutebench 2.0 – agentic benchmark for dynamic evaluation of llm responses to refutation instruction, 2025

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:33:15.938425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:33:13.063240Z digest=sha256:7c4db74b376d6cc0b8517fc11fe3f76ac7541034505713144162d3b01a55aa43

Observation 1dc49d07-536f-49d5-af35-f5ed750fec17 · outbound

This paper cites Benchmarking complex instruction-following with multiple constraints composition.Advances in Neural Information Processing Systems, 37:137610– 137645, 2024.

Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models Benchmarking complex instruction-following with multiple constraints composition.Advances in Neural Information Processing Systems, 37:137610– 137645, 2024

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:13.132176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:33:13.132176Z digest=sha256:5ebf7e5fbe4965f8c979daef384976ea527fd18dd0ce117558a455385d8c576e

Observation c15a7049-461a-4614-9ebf-2a802448278b · outbound

This paper cites GPT-4o System Card.

Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models GPT-4o System Card

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:13.201382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:33:13.201382Z digest=sha256:a2827bfca14e3f9412f53d78484ca2b42ad67d17030e484aa7e8e21a4ec4713a

Observation 10907a57-b443-4eff-8db6-2aa6a0c68a11 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models Training Verifiers to Solve Math Word Problems

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:13.276979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:33:13.276979Z digest=sha256:db2899a0ec67e11265428ecba363363d83fa694492f33ebf39eb9831ea9d6404

Observation e3db8243-d059-40b1-8ece-eb9e9b71d518 · outbound

This paper cites Minerva: Accelerating data analysis in next-generation ssds.

Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models Minerva: Accelerating data analysis in next-generation ssds

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:33:15.745229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:33:13.341352Z digest=sha256:e296d5a0255d0643e4871ef5c4da1dbdc95521dc1dc7b27a7c171d3dbf13899a

Observation 81bdc687-70e6-4f71-af8d-29c061b47e1a · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models Gonzalez, Hao Zhang, and Ion Stoica

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:13.429582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:33:13.429582Z digest=sha256:208a456dfb76738899440942c59436e708883f7d2cf0711cbbf6b7acdd0b4958

Observation 11a41faa-a906-4367-b22e-7db1df5f14e5 · outbound

This paper cites Qwen3, April 2025.

Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models Qwen3, April 2025

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:13.532662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:33:13.532662Z digest=sha256:23f27cf655f781ae128904cdd932a8ffa1b4fa43bd8d1eadab42c472ff6a8af6

Observation 55f2cf9e-82ae-4bbc-9a03-aa64511f35d8 · outbound

This paper cites SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild.

Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:13.621188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:33:13.621188Z digest=sha256:ce2851ee0f6caf4561cae9a61ac66d4ce2cb4da07885789fea68b8b86d21da0f

Observation 79fa9c6b-3bd0-4099-9b08-824846a7d1f1 · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:13.735556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:33:13.735556Z digest=sha256:5d34be2a617e19baa73df2b4e69224476fee8b43e765a8c2c0c8101ef7ece86d

Observation f49f2bd5-2c6d-4f20-8d8d-a2dddf292488 · outbound

This paper cites Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Li Erran Li, Raluca Ada Popa, and Ion Stoica.

Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Li Erran Li, Raluca Ada Popa, and Ion Stoica

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:13.835666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:33:13.835666Z digest=sha256:ce205c37cce29bc53379d28f87c16e3394447f816547f106131a0afa8eea5242

Observation dcec4b4d-7758-4c3d-9930-0c4c215f67e4 · outbound

This paper cites L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning.

Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:13.910728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:33:13.910728Z digest=sha256:ba2bcf2f3254ca582a6694dbecc5e844406d76843870a44b0a99243f140a5ca2

Observation 0cac9504-c05f-4349-a0ec-c1b24cc81c41 · outbound

This paper cites Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model.

Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:14.005401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:33:14.005401Z digest=sha256:4f1e9b4cd43a9e8659c16955259bd1adc4ad49f0806bfcc992c87e5f8dd4271c

Observation d6dc24a3-efa5-479b-adbe-fac5ed3701fe · outbound

This paper cites Open r1: A fully open reproduction of deepseek-r1, January 2025.

Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models Open r1: A fully open reproduction of deepseek-r1, January 2025

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:14.083554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:33:14.083554Z digest=sha256:ea144b8396b1a898ddd5039b7d398cb8777e3b65e407c0ab0a3f0e95d322f138

Observation f43c058a-ad96-4a97-8379-471d5590d1fc · outbound

This paper cites Qwq-32b: Embracing the power of reinforcement learning, March 2025.

Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models Qwq-32b: Embracing the power of reinforcement learning, March 2025

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:14.221930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:33:14.221930Z digest=sha256:730f255d1a1b01c9a2342b8407b725e232b6ff106bb19b1fc897a90e9625cda4

Observation 79f0a2f6-98d9-45a0-ac80-fb90fb618d74 · outbound

This paper cites Open-reasoner-zero: An open source approach to scaling up reinforcement learning on the base model, 2025.

Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models Open-reasoner-zero: An open source approach to scaling up reinforcement learning on the base model, 2025

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:14.341149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:33:14.341149Z digest=sha256:a536081655e221b04af07b28bf90154e04289d43fb74371b06a34462a6079d46

Observation c73eeb6f-0dd0-42b2-85bc-58f3076f5df9 · outbound

This paper cites Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063, 2024.

Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063, 2024

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:14.438481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:33:14.438481Z digest=sha256:9228e45ddad86ebde96d521546b344f8a23c65f66406166c901ea46f30ea35d4

Observation 2bd12cc8-133b-4fd6-8560-9ffb777bc94e · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:14.543415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:33:14.543415Z digest=sha256:86091b5f8b46ace172dc33f6b92c769d037ea88a9e3f39251fdded9c0dd0e194

Observation ebf27aa1-fe25-4327-a976-0eac331754cd · outbound

This paper cites The impact of reasoning step length on large language models.

Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models The impact of reasoning step length on large language models

Reference 53

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T15:33:15.469178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:33:14.622826Z digest=sha256:dbbd2c1ab16e224b33aa877a9fe0b89d712e58c5b178556c82f82d7a1557c530

Observation 2ff083fc-448b-4229-a8eb-0199ed4cec8c · outbound

This paper cites an unresolved cited work.

Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models Unresolved cited work

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:12.796349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:33:12.796349Z digest=sha256:459e13e469bceddd95893e3f2bdbcf051220fedf2603f8d49b5a65e35bf5c94d

Pith citing papers

Observation 1e8854ce-8f38-4ae7-9e39-60f8befb076a · inbound

Learning to Reason under Off-Policy Guidance cites this paper.

Learning to Reason under Off-Policy Guidance Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-15T23:17:02.898250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T23:17:02.701393Z digest=sha256:326339683068ff7952f70ecaa831797db7d13536cd75198e54fce299b948daa3

Observation cd451d3a-100f-4632-acb7-9f442874c142 · inbound

Activation Steering for Chain-of-Thought Compression cites this paper.

Activation Steering for Chain-of-Thought Compression Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:47.158980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:46:47.158980Z digest=sha256:940e19afc6f0c6fb7fcc8432f0e831b1a7deca74e35dd168cbb6059fe9628fd2

Observation 87a3f983-2e7c-44fb-b190-219dd66e7440 · inbound

Towards Concise and Adaptive Thinking in Large Reasoning Models: A Survey cites this paper.

Towards Concise and Adaptive Thinking in Large Reasoning Models: A Survey Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T17:53:47.566505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:53:47.566505Z digest=sha256:524f4f7482c00030ca5529a08e82fa016bef565f7abd935d095dbd32d7a890ef

Observation f4143950-2e76-456e-a840-66278d750e9e · inbound

AutoTIR: Autonomous Tools Integrated Reasoning via Reinforcement Learning cites this paper.

AutoTIR: Autonomous Tools Integrated Reasoning via Reinforcement Learning Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T12:23:55.086609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T12:23:55.086609Z digest=sha256:6a37da5263a27a2ecfa89913dd41252a867fc99747c34d174b46c809bd36e2bc

Observation df7b24f1-6691-41a3-8471-0b59b258279a · inbound

A Survey of Reinforcement Learning for Large Reasoning Models cites this paper.

A Survey of Reinforcement Learning for Large Reasoning Models Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models

Reference 146

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:02:25.397403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-18T00:02:24.352947Z digest=sha256:800ae6d562ac72e5900a401ad2e974be4b049e87c1bc4a84133aa83ace48ae3e

Observation 1244e1a4-09af-45f5-bc65-927153f45db9 · inbound

Retrieval-of-Thought: Efficient Reasoning via Reusing Thoughts cites this paper.

Retrieval-of-Thought: Efficient Reasoning via Reusing Thoughts Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:42:38.628885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T13:42:07.883909Z digest=sha256:8331a6fa541dd63081526ea9e083b4dcf36819b73f124735530643bab44e9a78

Observation cd80a8e1-9df3-4476-a93e-309dba05edbc · inbound

Position: The Hidden Costs and Measurement Gaps of Reinforcement Learning with Verifiable Rewards cites this paper.

Position: The Hidden Costs and Measurement Gaps of Reinforcement Learning with Verifiable Rewards Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-18T14:26:28.356258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T14:24:48.666197Z digest=sha256:8f05c0dcf2729470d1ecefef3f64e52951638fe30b3eb5f62056804620f4c190

Observation 5b236504-6262-4cd1-b48f-05ee07634c87 · inbound

Reasoning Up the Instruction Ladder for Controllable Language Models cites this paper.

Reasoning Up the Instruction Ladder for Controllable Language Models Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T07:07:52.626304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:07:52.626304Z digest=sha256:c7b2bba9ed93517e50c604b62dca71301ad314e0c6702c064498c4df947fe9fc

Observation 035981a2-0272-45d4-abc0-f931a2a19ce8 · inbound

ImpRIF: Stronger Implicit Reasoning Leads to Better Complex Instruction Following cites this paper.

ImpRIF: Stronger Implicit Reasoning Leads to Better Complex Instruction Following Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T08:00:44.619454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T07:59:35.303413Z digest=sha256:92e21bcf236c5555be78cedec258f086b631222324b70f8fe9c6d3020754f459

Observation 5d43761b-0d0d-4906-a652-117b55d16885 · inbound

Expert-Aware Refusal Steering cites this paper.

Expert-Aware Refusal Steering Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-07-02T03:26:29.106120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T10:04:13.562338Z digest=sha256:b44c7f696ea2130a54869006c58df21e0821d642fe602b360e8118bd12fcc0ef

Observation a5be2a59-61f1-4517-b01d-8da2e092dabe · inbound

When Built-in Thinking Helps and Hurts: Constraint-Level Error Shifts in Instruction Following cites this paper.

When Built-in Thinking Helps and Hurts: Constraint-Level Error Shifts in Instruction Following Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T01:27:30.782222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T16:33:59.661761Z digest=sha256:97accc45001d3340fc3dd5ea4ee621c1627130f358e033b10ab125bd68d2f77b

Observation f2244e3d-ae28-4746-96ff-a48ba458ff32 · inbound

Structured Thoughts For Improved Reasoning And Context Pruning cites this paper.

Structured Thoughts For Improved Reasoning And Context Pruning Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-07-14T12:08:05.502310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T12:08:05.502310Z digest=sha256:12d0c402e41661b835338a5f44c32e50e84b78178f3c6e3fbcb99550d5ec43f8