Pith. sign in

Paper Citation Record · LEDGER

MASPRM: Multi-Agent System Process Reward Model

As of 23 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2510.24803.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2510.24803 v3

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T07:56:48.047690Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b7692c4d-85f9-4910-a5a7-d62e1a7e7240 · outbound

This paper cites online" 'onlinestring :=.

MASPRM: Multi-Agent System Process Reward Model online" 'onlinestring :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T07:56:45.789075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:56:45.789075Z digest=sha256:41d08d4bab5f9fe35538bddc760b2e227f95769b5c772040032418783339f6d6

Observation 684fef4c-a351-48f4-868c-d894a4bc037f · outbound

This paper cites write newline.

MASPRM: Multi-Agent System Process Reward Model write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T07:56:45.833177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:56:45.833177Z digest=sha256:a0c4c77b7d8a23a7672ffbc49966ff47fdcabe9724dd2495d6a6cac33d224a2e

Observation b3f51e71-208f-4874-9e3b-34949f3ee422 · outbound

This paper cites an unresolved cited work.

MASPRM: Multi-Agent System Process Reward Model Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T07:56:45.894456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:56:45.894456Z digest=sha256:19f57704f0794bfdf6002a76ff0cc3c6e90b4bff98d77c80007e0aee82946692

Observation fb73fb84-8b1e-4f77-bbaf-95cbc41138c7 · outbound

This paper cites Scaling Autonomous Agents via Automatic Reward Modeling And Planning.

MASPRM: Multi-Agent System Process Reward Model Scaling Autonomous Agents via Automatic Reward Modeling And Planning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T07:56:45.944867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:56:45.944867Z digest=sha256:8faf81997674878de1b897b35b9ac2c4d2affd152bdcb12652677b0045426e0d

Observation 69ce9a63-a132-439c-b4ef-6b5382d7eb0c · outbound

This paper cites Process Reward Models for LLM Agents: Practical Framework and Directions.

MASPRM: Multi-Agent System Process Reward Model Process Reward Models for LLM Agents: Practical Framework and Directions

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T07:56:46.023170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:56:46.023170Z digest=sha256:9803f98b22fa2fd683190ff185f33056ede12c472555620174b3e8144aec4864

Observation 53283494-a6e1-490e-9314-547247250b80 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

MASPRM: Multi-Agent System Process Reward Model Training Verifiers to Solve Math Word Problems

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T07:56:46.107061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:56:46.107061Z digest=sha256:de6ca749f1f651cfe3a959002599aa6171ff449ab50b8ec9c29ab6a91e210d54

Observation 881e3b73-1110-4772-9f13-8601f252e676 · outbound

This paper cites Process Reinforcement through Implicit Rewards.

MASPRM: Multi-Agent System Process Reward Model Process Reinforcement through Implicit Rewards

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T07:56:46.211072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:56:46.211072Z digest=sha256:559a1622b136a218e8734bde2e904b9e1edcedb78e7cc4ea59038227231ab0a3

Observation 43d7f39d-9f5c-48a5-833c-98136f934a23 · outbound

This paper cites an unresolved cited work.

MASPRM: Multi-Agent System Process Reward Model Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T07:56:46.312568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:56:46.312568Z digest=sha256:bd021e6eac58ea23fb415ba6a6030252d036bb6eb0587b16c1882b898df8e32e

Observation 0c1c197b-2910-421a-b4f6-af84b593fac8 · outbound

This paper cites rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking.

MASPRM: Multi-Agent System Process Reward Model rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T07:56:46.376327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:56:46.376327Z digest=sha256:8a7466f05737e1b370e59091e11970e4913e4845e14b96e1f14b6afd1ce9e07b

Observation b70fef98-0b36-4922-815d-b009d1471623 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

MASPRM: Multi-Agent System Process Reward Model Measuring Mathematical Problem Solving With the MATH Dataset

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T07:56:46.446556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:56:46.446556Z digest=sha256:c3786f7cbaed61cc9da657ac3cca8bb42faf1e2cae307e3b1568ded8add3de3d

Observation 021f8d04-102a-46b9-8d57-c22f8d372532 · outbound

This paper cites an unresolved cited work.

MASPRM: Multi-Agent System Process Reward Model Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T07:56:46.502349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:56:46.502349Z digest=sha256:5f85d520065cc1764373ff98b2c5103b40e5272915a074b1e93de788d99ffb5c

Observation 7a8fa645-9e8a-43c9-b9ee-0c03bf070808 · outbound

This paper cites Six Challenges for Neural Machine Translation.

MASPRM: Multi-Agent System Process Reward Model Six Challenges for Neural Machine Translation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T07:56:46.579280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:56:46.579280Z digest=sha256:5ccb529f5ec86bc2bc6fe1e8e1008cd4b6e7eb7e53e90c5cf213ef00c28a4969

Observation a4c744a8-7fc4-4fc7-a50f-d1ac819ee512 · outbound

This paper cites MARFT: Multi-Agent Reinforcement Fine-Tuning.

MASPRM: Multi-Agent System Process Reward Model MARFT: Multi-Agent Reinforcement Fine-Tuning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T07:56:46.663818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:56:46.663818Z digest=sha256:1d049e51af683823fe09509d69716778d5db5ff405860a78fd41be4a6a1943a1

Observation ccd2f73c-dac3-476f-90d9-542b1c7f04b8 · outbound

This paper cites an unresolved cited work.

MASPRM: Multi-Agent System Process Reward Model Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T07:56:46.737622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:56:46.737622Z digest=sha256:7a9615f0f2d17f9f3b7ac62a4b138e70a28628c23c96798a18672c0a9404b1d5

Observation 61540494-045a-46b4-b6e9-baa27c71ad69 · outbound

This paper cites Speaking the Language of Teamwork: LLM-Guided Credit Assignment in Multi-Agent Reinforcement Learning.

MASPRM: Multi-Agent System Process Reward Model Speaking the Language of Teamwork: LLM-Guided Credit Assignment in Multi-Agent Reinforcement Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T07:56:46.797195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:56:46.797195Z digest=sha256:e06fabdf1a8ab53ba26fd355d2d0392ec33e15c2949fdf4642560ef0955b72ef

Observation 48877fb9-e9cd-4017-825f-afe11f9ab73f · outbound

This paper cites Improve Mathematical Reasoning in Language Models by Automated Process Supervision.

MASPRM: Multi-Agent System Process Reward Model Improve Mathematical Reasoning in Language Models by Automated Process Supervision

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T07:56:46.894742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:56:46.894742Z digest=sha256:cd51b8889e6b4140253a8ed6d9ff768920b9d452de107333ffa1dd7013c87318

Observation 1ef2ecb1-3502-4295-9f9a-d9a99d76b874 · outbound

This paper cites Leveraging Large Language Models for Effective and Explainable Multi-Agent Credit Assignment.

MASPRM: Multi-Agent System Process Reward Model Leveraging Large Language Models for Effective and Explainable Multi-Agent Credit Assignment

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T07:56:46.965356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:56:46.965356Z digest=sha256:2daafd187773bf8a2efde2021714ebff530207505e4fb7291c7fbb0f8801477c

Observation aa185488-4eb0-48e8-8ae1-8d51f677c548 · outbound

This paper cites Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning.

MASPRM: Multi-Agent System Process Reward Model Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T07:56:47.011678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:56:47.011678Z digest=sha256:3ad1a586c4e3da3251c39bd4f0cef870d25df01c0536304d52c8e3c8ec45e124

Observation 479e194e-4b08-45f6-b0b4-2035b59a8ee8 · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

MASPRM: Multi-Agent System Process Reward Model Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T07:56:47.083055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:56:47.083055Z digest=sha256:475d2a3b4acd24e2727afe64a7c69efd06bcaec31395f526d2ca3eb8074d956c

Observation e81255b8-ca8b-4fa4-a25f-17e737f2a235 · outbound

This paper cites SPA-RL: Reinforcing LLM Agents via Stepwise Progress Attribution.

MASPRM: Multi-Agent System Process Reward Model SPA-RL: Reinforcing LLM Agents via Stepwise Progress Attribution

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T07:56:47.157332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:56:47.157332Z digest=sha256:a7922aaa32999eb74c64f1786021a1431712b97f36b24f7476a72ca86ab8764a

Observation cf0f3b42-89d9-4fe8-b040-d0d9f591fd22 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

MASPRM: Multi-Agent System Process Reward Model Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T07:56:47.291212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:56:47.291212Z digest=sha256:c71b5d0f1071b70b331e1422a92c92f177a44013a43aad6aa8486a3790ca78a2

Observation 1525c82c-fdbb-405e-a361-e760530203a1 · outbound

This paper cites Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation.

MASPRM: Multi-Agent System Process Reward Model Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T07:56:47.435790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:56:47.435790Z digest=sha256:f2665a42f0c3bbb30c9af1ba228b2710b16f067da8c07bc5c1fca6db919b3bae

Observation 0f56fec8-7d61-4596-9512-7fee015ef8da · outbound

This paper cites Breaking the Beam Search Curse: A Study of (Re-)Scoring Methods and Stopping Criteria for Neural Machine Translation.

MASPRM: Multi-Agent System Process Reward Model Breaking the Beam Search Curse: A Study of (Re-)Scoring Methods and Stopping Criteria for Neural Machine Translation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T07:56:47.604970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:56:47.604970Z digest=sha256:3bd5f8d7d5a48b1bc74e888afa9cdfb1e004b27c11152d1a8ab2445484144968

Observation 4314658c-e392-409d-8aa4-778ee98c19c7 · outbound

This paper cites an unresolved cited work.

MASPRM: Multi-Agent System Process Reward Model Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T07:56:47.764276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:56:47.764276Z digest=sha256:eb16370659d138617185c11c952718b94d906cb99e5f20cb5f38bd556d18673e

Observation f4f44dfe-ee49-448d-9142-681886b99854 · outbound

This paper cites VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data.

MASPRM: Multi-Agent System Process Reward Model VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T07:56:47.849322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:56:47.849322Z digest=sha256:7e95e5d3f8426ac778de4d26bcb1d3e4e19f0954850752f131a7937eb0ac6dc5

Observation 525660d8-24af-4269-b8ea-866c6bd7e667 · outbound

This paper cites G-Designer: Architecting Multi-agent Communication Topologies via Graph Neural Networks.

MASPRM: Multi-Agent System Process Reward Model G-Designer: Architecting Multi-agent Communication Topologies via Graph Neural Networks

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T07:56:47.964740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:56:47.964740Z digest=sha256:ba9988997c9db8166a5637018304e3a326098cd1469dec38c90df8ab3783d2eb

Observation efbc73d8-16ad-4a69-b59b-a06923dfcd0d · outbound

This paper cites an unresolved cited work.

MASPRM: Multi-Agent System Process Reward Model Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T07:56:48.047690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:56:48.047690Z digest=sha256:a7268e3bfd6f6c0e6b5cec232cdc8e723559f96140a94682abb4e72862d4acb1

Pith citing papers

No inbound Pith citation observations are available.