Pith. sign in

Paper Citation Record · LEDGER

General-Reasoner: Advancing LLM Reasoning Across All Domains

As of 8 August 2026, this Paper Citation Record lists 34 of 34 outbound references and 30 inbound Pith citation observations for arXiv:2505.14652.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.14652 v5

Coverage vector

measured 34 of 34 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:34:10.160986Z

measured 64 of 64 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 30 of 30 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:35:52.681204Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T14:58:32.537541Z

Reference resolution

34 of 34 outbound references displayed

  • verified exact0
  • verified fuzzy4
  • unresolved29
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8660dee9-4aa0-4f2b-8b3f-ccbb38e9523f · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

General-Reasoner: Advancing LLM Reasoning Across All Domains DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:05.940907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:05.940907Z digest=sha256:50c3ce5652cd33fe5b92e39e4e65cb19756fa6b638329c12a03313c2416434ca

Observation 5f4d5688-991c-4ec3-b892-f229d8b214b6 · outbound

This paper cites SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild.

General-Reasoner: Advancing LLM Reasoning Across All Domains SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:06.045215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:06.045215Z digest=sha256:54546802b1e55e21debb96aaa5c66debde94d7e3548a035849e418a73e97fed9

Observation fa17cf45-45ea-48df-8151-bfd00c7d1077 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

General-Reasoner: Advancing LLM Reasoning Across All Domains DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:06.180306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:06.180306Z digest=sha256:aca29f535c35a7ceecb56df055a0bafec25357c632aee992619dce922d736f7f

Observation e4f54a7f-a643-4f7e-8794-c1dc59c3c474 · outbound

This paper cites Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Li Erran Li, Raluca Ada Popa, and Ion Stoica.

General-Reasoner: Advancing LLM Reasoning Across All Domains Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Li Erran Li, Raluca Ada Popa, and Ion Stoica

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:06.321248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:06.321248Z digest=sha256:5eb38a23d8a980217a61d7586009e9802b2e650d4186dbc90ee43f990d873bde

Observation 36bcc3cf-19b4-4d15-9b86-948cdbcd9e05 · outbound

This paper cites Deepcoder: A fully open-source 14b coder at o3-mini level, 2025.

General-Reasoner: Advancing LLM Reasoning Across All Domains Deepcoder: A fully open-source 14b coder at o3-mini level, 2025

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:06.486096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:06.486096Z digest=sha256:58c11ac23ffabf1f8c4cddd1cbff2bca1fa7b085ee91c6a612f78e8de7c4db39

Observation c7765d8d-b838-4ad8-aefc-8b5984f9f77a · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

General-Reasoner: Advancing LLM Reasoning Across All Domains DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:06.606581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:06.606581Z digest=sha256:0dd0d28720a87d1806e35674cec7ac983b7c6940c350ac9945186e6e01547d80

Observation 0b41e973-c2c2-4820-94a8-6df9c61bc57d · outbound

This paper cites s1: Simple test-time scaling.

General-Reasoner: Advancing LLM Reasoning Across All Domains s1: Simple test-time scaling

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:06.758536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:06.758536Z digest=sha256:832cfbf2a1202681567f94976c9fd7df4e77f60d9c26e6ce97e230589dc273eb

Observation 5396bf05-66ad-46fe-889f-81fc88969b09 · outbound

This paper cites MMLU-pro: A more robust and challenging multi-task language understanding benchmark.

General-Reasoner: Advancing LLM Reasoning Across All Domains MMLU-pro: A more robust and challenging multi-task language understanding benchmark

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:34:11.796009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:34:06.891461Z digest=sha256:e7c39189ef205590e22f2efaefbe26f3739b496beafcd440e933e83c97e2eb58

Observation f7ec4993-0070-4ba3-a03a-6d02924c8ddd · outbound

This paper cites MAmmoTH2: Scaling instructions from the web.

General-Reasoner: Advancing LLM Reasoning Across All Domains MAmmoTH2: Scaling instructions from the web

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:34:11.552054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:34:06.975131Z digest=sha256:0ee5c2b532dfd90e3d49181fe8b7abc0bac717239ac42b64e3ed81db300bffb9

Observation 75a69ac9-c4f3-4bb9-90b5-f4bdcdbbc664 · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

General-Reasoner: Advancing LLM Reasoning Across All Domains GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:07.116592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:07.116592Z digest=sha256:19a21ff3004c540defa7ab5dae9fdb8827c3203b8e2da1ff6e8e87000b54ec83

Observation 64a4f703-91d3-4723-845e-7053bd05a2bd · outbound

This paper cites SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines.

General-Reasoner: Advancing LLM Reasoning Across All Domains SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:07.251201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:07.251201Z digest=sha256:0973c558baeb3dcc2ab2f3c37cb44e200e0d4335373765f5e630b799f2187e21

Observation 752dd351-798e-4dca-884f-00620bd47ab3 · outbound

This paper cites TheoremQA: A theorem-driven question answering dataset.

General-Reasoner: Advancing LLM Reasoning Across All Domains TheoremQA: A theorem-driven question answering dataset

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:07.416960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:07.416960Z digest=sha256:9bc16050dfd3b9ea1d0f5657a9d33fec7c62d53ac61d6697bcf05967f11cfa32

Observation 7abd6682-5d79-4dbe-8f1b-764533b579b5 · outbound

This paper cites BIG-Bench Extra Hard.

General-Reasoner: Advancing LLM Reasoning Across All Domains BIG-Bench Extra Hard

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:07.541003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:07.541003Z digest=sha256:bb7eb55017d706800beb01db94dbf07fe9fc6d41f4856e0c44d4df8e476ecc87

Observation b58e34ca-405b-4694-a0b3-f38922369352 · outbound

This paper cites Measuring mathematical problem solving with the MATH dataset.

General-Reasoner: Advancing LLM Reasoning Across All Domains Measuring mathematical problem solving with the MATH dataset

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:07.679696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:07.679696Z digest=sha256:a6e0434644ffadb9a7e63f14f941138cdaaed42379e88f3990c3cd67873df7f6

Observation feedb8fe-64ce-4896-8fca-d57edbbb2586 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

General-Reasoner: Advancing LLM Reasoning Across All Domains Training Verifiers to Solve Math Word Problems

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:07.774827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:07.774827Z digest=sha256:eaebb7735ec104736502f2403cf6fae6fa07a3327d9fbefcb0ea514612f8de09

Observation 88ea4c29-168b-4bbf-989f-6909cb26e6cf · outbound

This paper cites OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

General-Reasoner: Advancing LLM Reasoning Across All Domains OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:07.896656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:07.896656Z digest=sha256:1c2dee045772b02f030acd6a47e065691995327e8eed65e9b21f2c190f6ad75c

Observation 61b49554-a17f-4557-a281-7c0e1044a9b2 · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

General-Reasoner: Advancing LLM Reasoning Across All Domains Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:08.034860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:08.034860Z digest=sha256:b8268cdea9b87d07c8f12ce58f38869e09f4d7033a48186c3a40bb3c6026ce95

Observation f65fb8a5-15cb-4e2b-99fa-5e41bafe52cb · outbound

This paper cites A Survey on Test-Time Scaling in Large Language Models: What, How, Where, and How Well?.

General-Reasoner: Advancing LLM Reasoning Across All Domains A Survey on Test-Time Scaling in Large Language Models: What, How, Where, and How Well?

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:08.168952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:08.168952Z digest=sha256:08bb76dfeb88faeae7d8ed0132be147060dbd5433e2aed32707948f8b0c7857b

Observation 20d67b80-af51-4d82-8d60-d8e885ce2f81 · outbound

This paper cites Efficient Test-Time Scaling via Self-Calibration.

General-Reasoner: Advancing LLM Reasoning Across All Domains Efficient Test-Time Scaling via Self-Calibration

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:08.320486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:08.320486Z digest=sha256:c9c9b55338ea0e0ba410371be9383cf10dc68dab25bc3995280d963389106ab9

Observation 3ae24153-daf7-4435-ad31-8f4ebbc6e3d8 · outbound

This paper cites OpenAI o1 System Card.

General-Reasoner: Advancing LLM Reasoning Across All Domains OpenAI o1 System Card

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:08.428347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:08.428347Z digest=sha256:325d860f8a88b66fa6b2f7873adebf86352cf4ea83bf99efad2f5d9dade4b3e6

Observation 70039290-fb81-4070-9396-c52d1f418df3 · outbound

This paper cites Qwen2.5 Technical Report.

General-Reasoner: Advancing LLM Reasoning Across All Domains Qwen2.5 Technical Report

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:08.546263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:08.546263Z digest=sha256:a47042a1edf5a14cda7e94de3119e5cbd3bc5d273302babad8b2e00e47062ab2

Observation 62152091-6a38-487d-b1c9-8c1bf0e489ed · outbound

This paper cites QwQ-32B: Embracing the power of reinforcement learning, March 2025.

General-Reasoner: Advancing LLM Reasoning Across All Domains QwQ-32B: Embracing the power of reinforcement learning, March 2025

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:08.659011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:08.659011Z digest=sha256:e1a94caa48f5fe46e01cf14f3bddf977fc99c42a3adea38754274af311d5c025

Observation 39b7bfed-5928-4752-af39-907b057f3687 · outbound

This paper cites MiMo: Unlocking the Reasoning Potential of Language Model -- From Pretraining to Posttraining.

General-Reasoner: Advancing LLM Reasoning Across All Domains MiMo: Unlocking the Reasoning Potential of Language Model -- From Pretraining to Posttraining

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:08.771561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:08.771561Z digest=sha256:942d6b225698d0cd853faeea152a92fac311dc42e28ebfaee418ddd69a31b3fb

Observation 301f934e-4a3f-4208-96e2-f87d2b259427 · outbound

This paper cites Training language models to follow instructions with human feedback.

General-Reasoner: Advancing LLM Reasoning Across All Domains Training language models to follow instructions with human feedback

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:08.878975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:08.878975Z digest=sha256:dea2139911bb9ad2395d65bd4bd4fdb273043f77528ff5c17a2daf23d49bff63

Observation c18bf4b9-14a5-4292-a002-0cab5289ebac · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

General-Reasoner: Advancing LLM Reasoning Across All Domains Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:09.034625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:09.034625Z digest=sha256:f00e4ca56e09057a9b6d364ac92451ea6e22994ded3e1f9cd9fde1f8eb6aa35f

Observation 7475d5f0-a5f1-4a9c-8f75-cff476ba9780 · outbound

This paper cites Nemotron-CrossThink: Scaling self-learning beyond math reasoning, 2025.

General-Reasoner: Advancing LLM Reasoning Across All Domains Nemotron-CrossThink: Scaling self-learning beyond math reasoning, 2025

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:09.154833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:09.154833Z digest=sha256:1623ba110949c86fe2649886e92af20813a2f75e36227b7e4a20f5c26d0d1810

Observation 5e84209f-bbff-45ee-a5d2-77e813fae459 · outbound

This paper cites Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains.

General-Reasoner: Advancing LLM Reasoning Across All Domains Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:09.276522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:09.276522Z digest=sha256:412680e794337d86390430338294cdde878989ec76e2a7bc52656c3738264a4b

Observation ec55bc7c-da4d-417c-8666-434d516fb66d · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

General-Reasoner: Advancing LLM Reasoning Across All Domains Gemini: A Family of Highly Capable Multimodal Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:09.410088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:09.410088Z digest=sha256:f9c77ac3df55960fc5fc6b349c5c6d0276fff3e3cd73d95b6bf3786243d1493a

Observation de6e2a22-695a-4287-8156-4b2377720a6e · outbound

This paper cites Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them.

General-Reasoner: Advancing LLM Reasoning Across All Domains Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:09.547506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:09.547506Z digest=sha256:2eca39dbe0bfef8edd1579b771b123fd91d693a844392284a11e812749c47314

Observation a4488639-9137-4a1c-919d-2d22188fc6c2 · outbound

This paper cites Solving quantitative reasoning problems with language models.

General-Reasoner: Advancing LLM Reasoning Across All Domains Solving quantitative reasoning problems with language models

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:34:11.354508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:34:09.650113Z digest=sha256:e1e695d9afd4193b8a01f0f1fe4d32279b4db5ed1b189203347cde4f4eb2d3ee

Observation e7393aa3-7260-4634-8426-7e27861835c0 · outbound

This paper cites GPT-4o System Card.

General-Reasoner: Advancing LLM Reasoning Across All Domains GPT-4o System Card

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:09.789116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:09.789116Z digest=sha256:acbd0ac92c54b696e08b3918220d94a0e4029650c6444f022db2dfca547f05b0

Observation 6cab01a3-205b-415b-86cd-173a2afd898f · outbound

This paper cites Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model.

General-Reasoner: Advancing LLM Reasoning Across All Domains Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model

Reference 32

Resolution
malformed identifier
no resolver link, observed 2026-08-07T15:34:09.929322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:09.929322Z digest=sha256:2cece68ffeba4b5d58a26c0da3ee6ebc6f09e5d588c9311b56917c41bec1d596

Observation 45cb9444-f20f-40f2-879a-2cba68c92911 · outbound

This paper cites an unresolved cited work.

General-Reasoner: Advancing LLM Reasoning Across All Domains Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:34:11.129480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:34:10.006075Z digest=sha256:5c48d78bda652c8b764e6204b483681b49d5ebbd0905147c6f5a3a4e261dfcf4

Observation 0476a43f-538a-4ae5-9ce0-ab9aa34d982e · outbound

This paper cites Final Decision: Yes A.5 Detailed Hyper-Parameters We provide the detailed hyperparameters for training our General-Reasoner variants in Table 9.

General-Reasoner: Advancing LLM Reasoning Across All Domains Final Decision: Yes A.5 Detailed Hyper-Parameters We provide the detailed hyperparameters for training our General-Reasoner variants in Table 9

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:34:10.933677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:34:10.160986Z digest=sha256:882e09fce031a0f0661433e26bd9da822306840c8b17f028f7b1dac7485acbcc

Pith citing papers

Observation e8604899-443b-4c3e-8189-b8cef03b172b · inbound

Reinforcing General Reasoning without Verifiers cites this paper.

Reinforcing General Reasoning without Verifiers General-Reasoner: Advancing LLM Reasoning Across All Domains

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:52.681204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:52.681204Z digest=sha256:02c8484c4abd0a1d5cca925a244f9dd3d309836f1931f44e690eaa57a68ab608

Observation c458e86f-bf83-4911-8d45-be01a33a7d17 · inbound

Unleashing the Reasoning Potential of Pre-trained LLMs by Critique Fine-Tuning on One Problem cites this paper.

Unleashing the Reasoning Potential of Pre-trained LLMs by Critique Fine-Tuning on One Problem General-Reasoner: Advancing LLM Reasoning Across All Domains

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:13:53.817051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:13:53.817051Z digest=sha256:eed812fffcbb7545c8cf4c80579fa1d248d7cfacd1f00358a6f44d16f4cf8df7

Observation b84d26e8-1849-4665-96d1-d46135f5a881 · inbound

Reinforcement Pre-Training cites this paper.

Reinforcement Pre-Training General-Reasoner: Advancing LLM Reasoning Across All Domains

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:22.131730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:26:22.131730Z digest=sha256:f6b3ccecd11ed6b141cd02291c3fd04d169f68b8607c67b30dfd03e572edbd9a

Observation f78fb8e2-03c2-47d9-a4ab-11b796b2e74a · inbound

One Token to Fool LLM-as-a-Judge cites this paper.

One Token to Fool LLM-as-a-Judge General-Reasoner: Advancing LLM Reasoning Across All Domains

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T18:15:37.994518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:15:37.994518Z digest=sha256:fa2fd986928717d612b232e88339da7fdef8eb242e63495ba2b486e1369ad973

Observation 1a004ef3-4a04-48a4-845c-295a6671fe8a · inbound

VerifyBench: A Systematic Benchmark for Evaluating Reasoning Verifiers Across Domains cites this paper.

VerifyBench: A Systematic Benchmark for Evaluating Reasoning Verifiers Across Domains General-Reasoner: Advancing LLM Reasoning Across All Domains

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T17:49:48.886824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:49:48.886824Z digest=sha256:20c9f28eb7b8c3c739ee7a4c29b1a7378b421e5d810dcd3616700b612c39baf1

Observation 1c1d8365-832a-456a-b54f-62345e323905 · inbound

MUR: Momentum Uncertainty guided Reasoning for Large Language Models cites this paper.

MUR: Momentum Uncertainty guided Reasoning for Large Language Models General-Reasoner: Advancing LLM Reasoning Across All Domains

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T03:37:01.273840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T03:34:06.058415Z digest=sha256:0fb9a0c0c47febf259ab3b5e495ac77b847926147712a34465d697565bbeb8dd

Observation 8f17ff87-35b6-4bee-b944-dab26b743bfa · inbound

Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains cites this paper.

Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains General-Reasoner: Advancing LLM Reasoning Across All Domains

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:07:56.804620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T06:07:56.678339Z digest=sha256:34f36310530d451305b5bc827bd49b68f14c3239195ac9adcbec6767a86bed8f

Observation 0242803c-7fac-4bad-bf40-2fb9cad42d41 · inbound

Hermes 4 Technical Report cites this paper.

Hermes 4 Technical Report General-Reasoner: Advancing LLM Reasoning Across All Domains

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:54.568612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:32:54.568612Z digest=sha256:76e6f54cd068464c1a881543c2e6d9f5d633087c98652eac45de7ee2f273c16c

Observation f3cbd91c-f7c8-41f4-983f-73cbf847eeea · inbound

Coupled Variational Reinforcement Learning for Language Model General Reasoning cites this paper.

Coupled Variational Reinforcement Learning for Language Model General Reasoning General-Reasoner: Advancing LLM Reasoning Across All Domains

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T16:43:56.063243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:43:56.063243Z digest=sha256:a41137fbaaf03ef976394f04f11874308a6dd58f851affeccf0ee648c6436dce

Observation c5967dec-a513-4d0c-a844-d051323217f3 · inbound

Outcome-Grounded Advantage Reshaping for Fine-Grained Credit Assignment in Mathematical Reasoning cites this paper.

Outcome-Grounded Advantage Reshaping for Fine-Grained Credit Assignment in Mathematical Reasoning General-Reasoner: Advancing LLM Reasoning Across All Domains

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T11:10:26.169628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:10:26.169628Z digest=sha256:77fda3f7babc46f9aeedb6ad280ffa1aafb039263d63e343ed77103a4a853cf9

Observation 62f02fdb-4ad0-4d2f-a77c-6f7a85099092 · inbound

MoCo: A One-Stop Shop for Model Collaboration Research cites this paper.

MoCo: A One-Stop Shop for Model Collaboration Research General-Reasoner: Advancing LLM Reasoning Across All Domains

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:17:43.759451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T10:17:37.129753Z digest=sha256:5b7fedbe396749cd6d22c0ead8491f57b03f05974424c65d4b287a788722541d

Observation 64427d66-ef48-4ef3-bb07-287120e20721 · inbound

Knowing Bias, Doing Better: Mitigating Social Bias in LLMs via Know-Bias Neuron Enhancement cites this paper.

Knowing Bias, Doing Better: Mitigating Social Bias in LLMs via Know-Bias Neuron Enhancement General-Reasoner: Advancing LLM Reasoning Across All Domains

Reference 589

Resolution
unresolved
no resolver link, observed 2026-08-03T06:47:17.081766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:47:17.081766Z digest=sha256:e79c944322b52473061ba14861596e466a780b1795625f146923f0eb5d781b79

Observation ed0e52c4-2ff1-4323-b54f-08c5820904c1 · inbound

GradAlign: Gradient-Aligned Data Selection for LLM Reinforcement Learning cites this paper.

GradAlign: Gradient-Aligned Data Selection for LLM Reinforcement Learning General-Reasoner: Advancing LLM Reasoning Across All Domains

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T21:05:25.922061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:05:25.922061Z digest=sha256:8cd5042a1cc95a7e51b83e26ab7053413b3d15224d401bb0aa8b4aa6234748b9

Observation 797f6862-7acc-4dc6-bcd3-b4badfd13adf · inbound

Vocabulary Dropout for Curriculum Diversity in LLM Co-Evolution cites this paper.

Vocabulary Dropout for Curriculum Diversity in LLM Co-Evolution General-Reasoner: Advancing LLM Reasoning Across All Domains

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-13T19:28:10.107358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T19:23:10.901441Z digest=sha256:f773243f0cc74ba2c9969f5c99ee192ff85c7205fe01dad0c78346f6004391b5

Observation 7794019f-94f7-48c4-803c-b6858482df27 · inbound

Vocabulary Dropout for Curriculum Diversity in LLM Co-Evolution cites this paper.

Vocabulary Dropout for Curriculum Diversity in LLM Co-Evolution General-Reasoner: Advancing LLM Reasoning Across All Domains

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T16:52:28.340493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:52:28.340493Z digest=sha256:f62f491dcb72b1316464e8e42ddcecbed92eaba0e38533d078768ce8348534ef

Observation 2f87208b-cc9f-464b-a250-c0f76a6b5909 · inbound

TDA-RC: Task-Driven Alignment for Knowledge-Based Reasoning Chains in Large Language Models cites this paper.

TDA-RC: Task-Driven Alignment for Knowledge-Based Reasoning Chains in Large Language Models General-Reasoner: Advancing LLM Reasoning Across All Domains

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:25:36.017296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T12:21:39.237267Z digest=sha256:bc5517eca7bcd585e94e2b901958fdc6eeaf4bb060b0da51daa164b818471b92

Observation 96ad7fa4-1ced-4789-8816-d6e831d70999 · inbound

HEALing Entropy Collapse: Enhancing Exploration in Few-Shot RLVR via Hybrid-Domain Entropy Dynamics Alignment cites this paper.

HEALing Entropy Collapse: Enhancing Exploration in Few-Shot RLVR via Hybrid-Domain Entropy Dynamics Alignment General-Reasoner: Advancing LLM Reasoning Across All Domains

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T09:23:37.532442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T05:23:08.478393Z digest=sha256:09bfd76bfb4f1aad90b1e2edbeebb3c301b6978ca6353215453bf2988b2dedb8

Observation 8764e888-71d2-4741-8a44-3cee1f0a77c0 · inbound

Rethinking RL for LLM Reasoning: It's Sparse Policy Selection, Not Capability Learning cites this paper.

Rethinking RL for LLM Reasoning: It's Sparse Policy Selection, Not Capability Learning General-Reasoner: Advancing LLM Reasoning Across All Domains

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:01:09.633344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T10:33:02.741857Z digest=sha256:e44cffe6a9bb8f98bbb3cd8748a50ef60ada31b24565dbcc61bde1069aa51f8a

Observation 02260b02-af6e-4c34-b6e9-be7e7667b468 · inbound

Rethinking RL for LLM Reasoning: It's Sparse Policy Selection, Not Capability Learning cites this paper.

Rethinking RL for LLM Reasoning: It's Sparse Policy Selection, Not Capability Learning General-Reasoner: Advancing LLM Reasoning Across All Domains

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:26:24.090924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T01:10:54.313745Z digest=sha256:01e89e414f458c953305b865e28a9fe27fd2a4ab9fcd5dedfaf70fe082dcd41c

Observation b561991d-0d04-4f2b-9389-a584b724ae35 · inbound

How Well Do LLMs Perform on the Simplest Long-Chain Reasoning Tasks: An Empirical Study on the Equivalence Class Problem cites this paper.

How Well Do LLMs Perform on the Simplest Long-Chain Reasoning Tasks: An Empirical Study on the Equivalence Class Problem General-Reasoner: Advancing LLM Reasoning Across All Domains

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:45:51.968561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-11T01:29:33.453354Z digest=sha256:ccc71ed8f65170bbe72854b532167783378582838c9371bf71924b235bcf85cd

Observation 8c27bc82-1920-4710-8c03-1da9f3780b3d · inbound

CLR-voyance: Reinforcing Open-Ended Reasoning for Inpatient Clinical Decision Support with Outcome-Aware Rubrics cites this paper.

CLR-voyance: Reinforcing Open-Ended Reasoning for Inpatient Clinical Decision Support with Outcome-Aware Rubrics General-Reasoner: Advancing LLM Reasoning Across All Domains

Reference 153

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:11:23.746522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-12T04:32:16.930291Z digest=sha256:2e9717f31761ad5aeac1e5c576579db2a52ea7ffb198629129138ce9849fd2d3

Observation 8c98ad48-f85f-40c2-8df1-1f32771137ad · inbound

M2A: Synergizing Mathematical and Agentic Reasoning in Large Language Models cites this paper.

M2A: Synergizing Mathematical and Agentic Reasoning in Large Language Models General-Reasoner: Advancing LLM Reasoning Across All Domains

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:01:22.716327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T04:43:50.384980Z digest=sha256:be49585fef8092c8df600ed068728728defb90c2d8b6db251372336bd4e37e02

Observation 1da4d8c6-7b3a-4288-9ab6-9c323b7e5c41 · inbound

Uncovering the Representation Geometry of Minimal Cores in Overcomplete Reasoning Traces cites this paper.

Uncovering the Representation Geometry of Minimal Cores in Overcomplete Reasoning Traces General-Reasoner: Advancing LLM Reasoning Across All Domains

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-06-30T21:05:03.840065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T21:04:02.263300Z digest=sha256:923dff613ffc2bfc5b111b848d8abf4e7bbaf28eaaa3d882d0ab687420252d3a

Observation 4c8bd9bd-d072-436a-9cd3-e510aee9a90f · inbound

Harnessing LLM Agents with Skill Programs cites this paper.

Harnessing LLM Agents with Skill Programs General-Reasoner: Advancing LLM Reasoning Across All Domains

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:28:14.365903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T11:26:19.382463Z digest=sha256:17e734a832fb4aa2d5a86956a76086ecb09e228e07abcdb05caf880ba629eeeb

Observation 817f0464-c865-4d97-bf12-1a6adf7a3269 · inbound

CARE-RL: Capability-Aware Reinforcement Learning for Mitigating Cross-Domain Conflicts cites this paper.

CARE-RL: Capability-Aware Reinforcement Learning for Mitigating Cross-Domain Conflicts General-Reasoner: Advancing LLM Reasoning Across All Domains

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T19:02:34.037389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T19:01:14.340754Z digest=sha256:e5f9a278e59b1b1a4916e73248b4f0912158b3b9ca611fd57466fb10812e86b1

Observation d6bd743e-635f-48b5-a0e2-b434492a0924 · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation General-Reasoner: Advancing LLM Reasoning Across All Domains

Reference 180

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T20:46:14.301614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:f0177fc0cbe126ab94ee6209d63303f491e48b501ae023919e059975b905ad89

Observation ad228ab9-3f72-466f-adb9-f8cd1dd759d0 · inbound

ResMerge: Residual-based Spectral Merging of Large Language Models cites this paper.

ResMerge: Residual-based Spectral Merging of Large Language Models General-Reasoner: Advancing LLM Reasoning Across All Domains

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T22:46:19.489013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T15:03:27.081409Z digest=sha256:da07636d3bf173154e98cbe34484f59e99495f92e40faff9e6c488944691592e

Observation 0ea3f5d0-a361-44e1-bfef-4275aac5c3cf · inbound

Invariant Gradient Alignment for Robust Reasoning Distillation cites this paper.

Invariant Gradient Alignment for Robust Reasoning Distillation General-Reasoner: Advancing LLM Reasoning Across All Domains

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T07:26:46.043395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T06:56:05.310135Z digest=sha256:a7495f3cc2abc8711449d976a86bdefeb60a8c4233f9349479e9d01b15f3ff8a

Observation ab7a2902-ecef-4c65-ad83-8059c68c19e7 · inbound

Scaling with Confidence: Calibrating Confidence of LLMs for Adaptive Test Time Scaling cites this paper.

Scaling with Confidence: Calibrating Confidence of LLMs for Adaptive Test Time Scaling General-Reasoner: Advancing LLM Reasoning Across All Domains

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T14:58:32.539111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-07-03T14:49:33.364596Z digest=sha256:91ea82980abcb41adde38095585e7361234562366a697c7e65886458d67f4bd9

Observation 5e19bfea-0437-473c-8d2b-e9a4c1e01820 · inbound

ArchEval: Measuring AI Agents as Computer Architects cites this paper.

ArchEval: Measuring AI Agents as Computer Architects General-Reasoner: Advancing LLM Reasoning Across All Domains

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-12T01:16:50.927415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:16:50.927415Z digest=sha256:003ddd8a8c5ca89787c36d8bc0c03ea8ee2e44383fa67c055b7bb8041829e7f9