Pith. sign in

Paper Citation Record · LEDGER

MOTIF: Modular Thinking via Reinforcement Fine-tuning in LLMs

As of 7 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 0 inbound Pith citation observations for arXiv:2507.02851.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.02851 v1

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:24:00.072659Z

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

21 of 21 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved21
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d1ab79bf-bfa9-4f86-8244-f623cc80202b · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

MOTIF: Modular Thinking via Reinforcement Fine-tuning in LLMs Training Verifiers to Solve Math Word Problems

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:58.425741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:58.425741Z digest=sha256:ddfca1000a141169210f3877bc22b7ab6424377050bc12510e97adb4cdf8e1a1

Observation 4bbb6bad-8780-49ae-8cff-76d35cbd5203 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

MOTIF: Modular Thinking via Reinforcement Fine-tuning in LLMs DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:58.618190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:58.618190Z digest=sha256:3f7576b2fcda7f27dfa8f6e43ca560b3d842e62de38f731ebe8a635dbdd36ef3

Observation 10d88f10-faae-461e-92be-b5b33ae3038e · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

MOTIF: Modular Thinking via Reinforcement Fine-tuning in LLMs Measuring Mathematical Problem Solving With the MATH Dataset

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:58.701438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:58.701438Z digest=sha256:d90a127a4adbcad447337bf9cc323beb9d53ffb7c912ee6495172d117fd7da53

Observation eeac5b9b-b6d6-4df1-8850-e86650132768 · outbound

This paper cites Lost in the Middle: How Language Models Use Long Contexts.

MOTIF: Modular Thinking via Reinforcement Fine-tuning in LLMs Lost in the Middle: How Language Models Use Long Contexts

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:59.003617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:59.003617Z digest=sha256:fd1f66b7a983335f3e7106200876f97f4d2d6a74050fbdbc26a8a968ee4e1732

Observation 25bb37db-d3ba-485c-b145-102e33531c4c · outbound

This paper cites Distributed Mixture-of-Agents for Edge Inference with Large Language Models.

MOTIF: Modular Thinking via Reinforcement Fine-tuning in LLMs Distributed Mixture-of-Agents for Edge Inference with Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:59.097832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:59.097832Z digest=sha256:283f85c0ead85559ee46884d400f2996c1ad3b32a9fd9c2174ff97e928ba46e4

Observation 13df14a6-0ba3-4bfd-9742-b2bafe2e196f · outbound

This paper cites s1: Simple test-time scaling.

MOTIF: Modular Thinking via Reinforcement Fine-tuning in LLMs s1: Simple test-time scaling

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:59.192668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:59.192668Z digest=sha256:20054241b0e6298e70a544c7f55368003adb3a454acf9788469ed49a680d8a29

Observation 5e128e2a-9d15-41db-8ed4-447e90630a8e · outbound

This paper cites Leave No Context Behind: Efficient Infinite Context Transformers with Infini-attention.

MOTIF: Modular Thinking via Reinforcement Fine-tuning in LLMs Leave No Context Behind: Efficient Infinite Context Transformers with Infini-attention

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:59.260212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:59.260212Z digest=sha256:e2d56fd406f834152c31a3861fe20982722a8cb239555873ce4b6a7aa6956633

Observation 54dd50c3-df24-42f5-b8f8-ec4744002cf9 · outbound

This paper cites Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning.

MOTIF: Modular Thinking via Reinforcement Fine-tuning in LLMs Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:59.391394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:59.391394Z digest=sha256:a4df20c952eabeba43ab8eece7be93d13d10a23b59c48d5687c18f9f1d7023c1

Observation 6f21235f-aaf2-4b8b-93d7-aa0cc0404c4c · outbound

This paper cites Iteration of Thought: Leveraging Inner Dialogue for Autonomous Large Language Model Reasoning.

MOTIF: Modular Thinking via Reinforcement Fine-tuning in LLMs Iteration of Thought: Leveraging Inner Dialogue for Autonomous Large Language Model Reasoning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:59.459779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:59.459779Z digest=sha256:7c2149d7f36d426b9e6a165d1c509974af026c9192b36aa7277e4d9949a32fb2

Observation 5624c9c5-cef8-4112-969c-fec3d4d8411b · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

MOTIF: Modular Thinking via Reinforcement Fine-tuning in LLMs DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:59.546311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:59.546311Z digest=sha256:856097c929d58dbac05e225f2e0ce72b0f5773d625849ba672e03d3f511cd348

Observation 4acbca54-faff-4b52-9264-555fd99b6fe8 · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

MOTIF: Modular Thinking via Reinforcement Fine-tuning in LLMs Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:59.612071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:59.612071Z digest=sha256:fbb7c14c7ab120c10d3becd7f4583b087a404cea4e75f5c87a6bed3b7326ad53

Observation ec8f81c2-c71e-48f4-8be7-a5f134b74053 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

MOTIF: Modular Thinking via Reinforcement Fine-tuning in LLMs Gemini: A Family of Highly Capable Multimodal Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:59.704363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:59.704363Z digest=sha256:2f0b7078b71eb14ff00e2ec1c7af0b8752d7aa97e823119f7c36cf0c4e99a2c3

Observation 4c107b09-0a26-4ad1-b454-e6d62fdca359 · outbound

This paper cites Think Twice: Enhancing LLM Reasoning by Scaling Multi-round Test-time Thinking.

MOTIF: Modular Thinking via Reinforcement Fine-tuning in LLMs Think Twice: Enhancing LLM Reasoning by Scaling Multi-round Test-time Thinking

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:59.774601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:59.774601Z digest=sha256:de83bb980e7b20b47e9cbcbfd8c342ba959f0dcc90b4cd1270d134f03fb7c27b

Observation b5b6da0e-5c72-4e63-9fd5-00f508c0bc21 · outbound

This paper cites Mixture-of-Agents Enhances Large Language Model Capabilities.

MOTIF: Modular Thinking via Reinforcement Fine-tuning in LLMs Mixture-of-Agents Enhances Large Language Model Capabilities

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:59.868672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:59.868672Z digest=sha256:05f7cc0c4ff5eb67dc282905457f447efac9f139fed4de5c69c75cf7987f47b4

Observation 3e115bb4-420d-4c02-b88a-a10199752927 · outbound

This paper cites Inftythink: Breaking the length limits of long-context reasoning in large language models.

MOTIF: Modular Thinking via Reinforcement Fine-tuning in LLMs Inftythink: Breaking the length limits of long-context reasoning in large language models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:59.964450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:59.964450Z digest=sha256:7c53f9a2bc3275b8faa9bcbfbefead7f97c1e24fff3a7a369d7a89c0bb296869

Observation 19b1f7ff-1647-42a6-8aef-00e65e33485f · outbound

This paper cites Infinite Retrieval: Attention Enhanced LLMs in Long-Context Processing.

MOTIF: Modular Thinking via Reinforcement Fine-tuning in LLMs Infinite Retrieval: Attention Enhanced LLMs in Long-Context Processing

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:00.072659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:00.072659Z digest=sha256:c11b19d3a775e3802e62d727911e20c70522ae15a5899910c470beaa80a143bc

Observation d697b007-2cf2-4794-be4f-2283440f5ce2 · outbound

This paper cites Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models.

MOTIF: Modular Thinking via Reinforcement Fine-tuning in LLMs Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:58.328473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:58.328473Z digest=sha256:b6fe938c57887d5ebf252e7b50c6aaa421922f614b8980a0f1d28771503ca6a7

Observation 1673fa13-dfda-421b-8440-b4887d3289ba · outbound

This paper cites OpenAI o1 System Card.

MOTIF: Modular Thinking via Reinforcement Fine-tuning in LLMs OpenAI o1 System Card

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:58.821730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:58.821730Z digest=sha256:43443b6b2a7d979cb364b713e31d04cf04656feb0b09626960474159acf7869f

Observation bd220da0-e9ad-43e9-b6fd-e1996fb0e4de · outbound

This paper cites Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use.

MOTIF: Modular Thinking via Reinforcement Fine-tuning in LLMs Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:58.523176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:58.523176Z digest=sha256:92623125b84ee53b9d92ac89c2c0382fedbddd4dc11260aa305276bbccac2cbf

Observation f67615ef-5366-4947-a05b-f5a7327e10e0 · outbound

This paper cites SMoA: Improving Multi-agent Large Language Models with Sparse Mixture-of-Agents.

MOTIF: Modular Thinking via Reinforcement Fine-tuning in LLMs SMoA: Improving Multi-agent Large Language Models with Sparse Mixture-of-Agents

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:58.917129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:58.917129Z digest=sha256:1d4c752914ba6f37e04f728a5e0f0c1b3f2d1b55fb2dbd519f5ee086c3a82774

Observation 127a6973-4793-4e5f-bb46-1190510a15f9 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

MOTIF: Modular Thinking via Reinforcement Fine-tuning in LLMs Evaluating Large Language Models Trained on Code

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:58.197865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:58.197865Z digest=sha256:7130d04b87f946be08d1c9d84337923af5d13d4430e23a77a376d9e55fce4058

Pith citing papers

No inbound Pith citation observations are available.