Pith. sign in

Paper Citation Record · LEDGER

VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data

As of 21 August 2026, this Paper Citation Record lists 74 of 74 outbound references and 6 inbound Pith citation observations for arXiv:2502.06737.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.06737 v2

Coverage vector

measured 74 of 74 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T14:32:10.732509Z

measured 80 of 80 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T00:59:54.314068Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T20:49:00.766137Z

Reference resolution

74 of 74 outbound references displayed

  • verified exact0
  • verified fuzzy17
  • unresolved55
  • parse uncertain2
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0a59c4f0-7c83-4f7c-8573-125e34c76d9f · outbound

This paper cites an unresolved cited work.

VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data Unresolved cited work

Reference 4

Resolution
parse uncertain
raw_fallback, observed 2026-08-08T14:32:12.075011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:32:10.323686Z digest=sha256:32a6c9cef5560411c87b4aa9faed470685158b1737d28d44ce3556fe6f3e7494

Observation a443ee0f-8a18-4135-8098-dcda6d4450ad · outbound

This paper cites an unresolved cited work.

VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:32:12.061130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:32:10.331717Z digest=sha256:8ee089b75661d302ea46ce7d61ddfc568fcaed61bed8cfa97a38e8bed89bcaef

Observation c4d99383-4717-4272-9768-9304a07fae6d · outbound

This paper cites A student’s final answer is considered correct if it matches the ground truth answer or only differs due to differences in how the answer is rounded.

VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data A student’s final answer is considered correct if it matches the ground truth answer or only differs due to differences in how the answer is rounded

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:32:12.049173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:32:10.338652Z digest=sha256:bca1fc176a4762e5d096514b375edf9da2d9aa0eb1a78f320e043df3121390f7

Observation c145d8ad-dd97-4f75-845d-7a474727b55e · outbound

This paper cites Your task is to: 1.

VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data Your task is to: 1

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:32:12.036917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:32:10.344183Z digest=sha256:dc9d4b42ce89d9078a7f9db6d7975e5702ceca96234988005143824100605213

Observation 490ee6f6-02f7-4a08-aa6c-38118b608103 · outbound

This paper cites This should include: - Identifying a step where the reasoning could naturally deviate.

VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data This should include: - Identifying a step where the reasoning could naturally deviate

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:32:12.024869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:32:10.348830Z digest=sha256:e19339a26dd6d75c1d388f7123936247e10019d9274797b4ddcceafe3a08a5a7

Observation add17453-6820-4374-9d6a-9ad110861df1 · outbound

This paper cites This incorrect step should: - Reflect a deviation in reasoning that significantly harms the correctness.

VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data This incorrect step should: - Reflect a deviation in reasoning that significantly harms the correctness

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:32:12.010447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:32:10.353000Z digest=sha256:46fd9477d647a8733760b1826a9b60e1bf3124ff646ca70a63fc595d67ddc806

Observation 7da8f460-c833-4641-924c-bedbf61c2fdf · outbound

This paper cites an unresolved cited work.

VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:32:11.988657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:32:10.357375Z digest=sha256:caf64ff4eb7a33a9abd5ffddc95d4b901ec5447859d6b89517bf382182729d22

Observation 4cab58e6-d00f-4814-b20e-d33f9de27aae · outbound

This paper cites - Verifiable: The step can be verified using common knowledge, simple calculations, or a quick refer- ence (e.g., recalling a basic theorem).

VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data - Verifiable: The step can be verified using common knowledge, simple calculations, or a quick refer- ence (e.g., recalling a basic theorem)

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:32:12.115698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:32:10.362337Z digest=sha256:897092586766984285acec5bcc7aa789f3867b4b7c7e576777832ce428ece0f0

Observation 10266fca-597a-4bd1-b513-1a16d6102ce5 · outbound

This paper cites Good job!.

VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data Good job!

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:32:12.102248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:32:10.366565Z digest=sha256:e30f1dbf6da6e065e7cf8a644b697abc9074b87ed1e25864d5b1607b2ce63b58

Observation 16381305-d567-4437-854b-6c7697b182d7 · outbound

This paper cites - Is Hard to Verify: Requires significant effort to confirm due to poor explanation.

VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data - Is Hard to Verify: Requires significant effort to confirm due to poor explanation

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:32:12.088003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:32:10.370712Z digest=sha256:b359f7ae549b11dbf618c1f6852d9f70808b3377bba58d0eab312440be53192e

Observation 88c8f66c-59a8-4a4e-9e17-87b0644f7bde · outbound

This paper cites Otherwise, return the index of -1 (which denotes all steps are GOOD or OK).

VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data Otherwise, return the index of -1 (which denotes all steps are GOOD or OK)

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:32:11.961006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:32:10.375419Z digest=sha256:b94f94b3daeb1165496f11335f8caaed3a4d989235e13b31844d4496ec73c073

Observation 78c6957c-c362-42c8-9070-36149de33a42 · outbound

This paper cites an unresolved cited work.

VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:32:11.944718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:32:10.379848Z digest=sha256:4d066c882dee34e00a562c35813cc8d80dfb90d4ee0fbf27c0aa9e1320484bee

Observation c5bf14cd-7e07-444e-8760-15a2e9e104f8 · outbound

This paper cites This process continues until a terminal node is reached.

VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data This process continues until a terminal node is reached

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:32:11.914376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:32:10.383770Z digest=sha256:24f004b7a51dfd2f2e099f2f098a7e2263c796b16990ad35a25167801defb725

Observation 322952c3-ba45-4532-aba3-3a4ce8e08d78 · outbound

This paper cites These steps are repeated for a fixed number of iterations or until a computational or time limit is reached.

VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data These steps are repeated for a fixed number of iterations or until a computational or time limit is reached

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:32:11.864138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:32:10.387834Z digest=sha256:c8ef381c820dedd3c259f01210d4e894a94c53a2e7392789c2a2721d1b276f63

Observation 9a9d6bfa-8c38-4834-bbae-c6b662a43ff4 · outbound

This paper cites an unresolved cited work.

VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:32:11.812151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:32:10.392881Z digest=sha256:7fa7e830bc99a9e8f853d3d0a88562f4938f06e0e529463a0df3fba440722651

Observation 4ae3eb8e-c469-46e2-a7c8-a6013535c7a2 · outbound

This paper cites an unresolved cited work.

VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:32:11.771436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:32:10.397444Z digest=sha256:a4bbd3b499b786fb2a466b4c069ce006e5a381b9aadc9d20bb0b1b424e2975b6

Observation ceba9e57-3def-45d1-9fdd-92a50c900ab8 · outbound

This paper cites an unresolved cited work.

VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:32:11.744326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:32:10.402174Z digest=sha256:c8f126900b118df9d4e371ce75acb0164108c4e1e3ff2cdf6f0c76ae28cb702f

Observation b78cbf04-8536-4fab-80fd-eee15b7e2570 · outbound

This paper cites an unresolved cited work.

VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:32:11.732081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:32:10.407012Z digest=sha256:066b5c8ae176b3e7e56767267f2f4426b06f8e5f33d8de93e43f7ec25104bf99

Observation da8c17fb-56b1-47b4-926f-237694543d2e · outbound

This paper cites an unresolved cited work.

VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:32:11.720151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:32:10.412018Z digest=sha256:8e0ae8ba648d0354f612b42c2f82ba5e8ad92513e89a2bcefc85517f3babd9b1

Observation 2a0feb1d-ee04-4b71-80fc-816389f549f6 · outbound

This paper cites 33 VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data.

VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data 33 VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:32:11.705556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:32:10.417084Z digest=sha256:97c369a1411d95343997c560b8c515388a7ed80a9dcd3e8eb8baa81c8bb6d097

Observation 43f30992-d450-4301-8563-c57ef45148f8 · outbound

This paper cites an unresolved cited work.

VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:32:11.690181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:32:10.421826Z digest=sha256:5cdc8b68e1538925c91397694d5759e8c30107486d7ead2e32dccfd34935d026

Observation fd97ac58-2655-43bd-b689-d73ab0d01832 · outbound

This paper cites Math PRM rewards: 0.94, 0.92, 0.96, 0.87 VersaPRM rewards: 1.00, 0.43, 1.00, 0.09 Explanation: The Math PRM fails to detect that the wrong hearsay exception was applied.

VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data Math PRM rewards: 0.94, 0.92, 0.96, 0.87 VersaPRM rewards: 1.00, 0.43, 1.00, 0.09 Explanation: The Math PRM fails to detect that the wrong hearsay exception was applied

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:32:11.677372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:32:10.426243Z digest=sha256:b42cb6946ee7e2b4124d3c9257bfc1fa626171854db6b92e20ad7b48dbc3ca3b

Observation 5a42e14c-7d81-47fc-913f-1cc8c1fad88c · outbound

This paper cites an unresolved cited work.

VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:32:11.659699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:32:10.431463Z digest=sha256:85d837df56964915d8bf609b5a8c8983999a7a9bc7bf9e74cc470c4cad800865

Observation 734db4ff-4e45-44d2-8a11-ee3fbf6d2a58 · outbound

This paper cites an unresolved cited work.

VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:32:11.642577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:32:10.436521Z digest=sha256:918b28fab889dbfb3b263350d9a6761f165397c69acf794922b6ab989b856d37

Observation 3fdefb31-3c15-477d-aa28-9ddb6e117a98 · outbound

This paper cites an unresolved cited work.

VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:32:11.626789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:32:10.440337Z digest=sha256:1f8064a0eedb2aa23c068099c062ba1178efab13ac71267aec224c9b208b8266

Observation 4c9ec132-0c80-4ed4-902d-416fe830a837 · outbound

This paper cites an unresolved cited work.

VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:32:11.612589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:32:10.444055Z digest=sha256:0c508b323e97d6e7239167c2388078048e552c4f87befb2b39bd293605cc8a2d

Observation e8ac57aa-eb6e-4cfd-9bbe-16e299fb431c · outbound

This paper cites an unresolved cited work.

VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:32:11.595421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:32:10.448203Z digest=sha256:8e6b22ff9241509744416952a15999ccd587d2a39e8c879c198e9bbdc920b7ca

Observation 33adc291-6033-46c8-90d4-ff55092700ed · outbound

This paper cites an unresolved cited work.

VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:32:11.573835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:32:10.452473Z digest=sha256:90ca8696f439a1e098f09cde8cdc1173b48a727bab2a3097208cbc110b276f85

Observation f7223351-2a9c-42ef-9988-f5f335b7b850 · outbound

This paper cites an unresolved cited work.

VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:32:11.545051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:32:10.456164Z digest=sha256:c1d787a4bcff4daa2b6ccc3f610607aae3648898433ea80168891280d46a7e3e

Observation faf8697d-eaaa-4430-a1dd-254df8d095e8 · outbound

This paper cites race to the bottom.

VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data race to the bottom

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:32:11.516446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:32:10.459940Z digest=sha256:b5bff312418b764870e8553e15322d2501787e6e275fd84fdf01c99cc5707fe4

Observation f65efcb2-935e-40aa-b160-16926f351c88 · outbound

This paper cites an unresolved cited work.

VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:32:11.489253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:32:10.464080Z digest=sha256:d432de738c6cc1dd7b9f711d27a8d169bfe04a98f3c86433e960a0057284eae6

Observation 885a6794-cbbb-46e5-ab7d-20364bc631cd · outbound

This paper cites This is an example of devolution because the federal government is limiting the power of the state to act on its own, effectively ”devolving” power back to the federal level.

VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data This is an example of devolution because the federal government is limiting the power of the state to act on its own, effectively ”devolving” power back to the federal level

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:32:11.468521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:32:10.468168Z digest=sha256:d84a702de667e1a09e56ce76bd6c9619c0a15cf60b3f5a99391daabbeff30301

Observation 85138796-aed2-4c9e-b69a-9ffabb79f3fa · outbound

This paper cites an unresolved cited work.

VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:32:11.454999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:32:10.478428Z digest=sha256:8de494fe035da143815e416b57452bec42bb064da63cfb75aa6b2a31683eb359

Observation 6b1ffba6-334b-42d4-bfaf-c6b531ed8906 · outbound

This paper cites an unresolved cited work.

VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:32:11.442925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:32:10.483665Z digest=sha256:c78be388528f104f2c92f39830355b3347cc9bfd4389fd555328bd0d1bee35d3

Observation 1e8f8dd8-9544-473c-aec7-e1725496a15a · outbound

This paper cites an unresolved cited work.

VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:32:11.428158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:32:10.488428Z digest=sha256:408c9502ef333b3095f2caac6b8bc6f068ca4d187b9671aadbe641c7bf268399

Observation 87788937-a00b-45f1-8fbd-d5335070a93b · outbound

This paper cites an unresolved cited work.

VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data Unresolved cited work

Reference 39

Resolution
parse uncertain
raw_fallback, observed 2026-08-08T14:32:11.414629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:32:10.493939Z digest=sha256:7db2c3a889d1b3d8cde54eb1dbf81193e67900c666c24aabcb68d57c4591f082

Observation d59d4d11-4eda-48ff-8d0d-022666874e21 · outbound

This paper cites an unresolved cited work.

VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:32:11.399881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:32:10.498981Z digest=sha256:ce7755ee73331efb7f284e82a7189abf3c2d16c6b3e54b15b873d374a3b27b21

Observation 9d64dc5e-8711-4bdb-b742-cb8b127ba305 · outbound

This paper cites an unresolved cited work.

VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:32:11.378232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:32:10.503473Z digest=sha256:0ece05e8af9df023455d7585854dfbb940e48721cfd82e33841af69ee3e67f45

Observation 8885c703-11ba-4b62-9de2-0db2a6ea14df · outbound

This paper cites an unresolved cited work.

VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:32:11.360707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:32:10.510762Z digest=sha256:69d12b81776e7eef64f2f64f6589993a821db7094c1b13972aaf800f8ecf487f

Observation fad6c245-8daa-4b0d-988d-dce0b528d1f9 · outbound

This paper cites an unresolved cited work.

VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:32:11.346678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:32:10.515548Z digest=sha256:da46fac219a3b88dfea906c1ea1cca815a4e48751eddb9cbcc48f3e554ac5d59

Observation df6ea495-c2bd-49ec-bd17-d8cff99400a1 · outbound

This paper cites an unresolved cited work.

VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:32:11.316955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:32:10.526878Z digest=sha256:202273c1f7fe03d698f6bf4c025901378c21b8424dd2c4707b81e6798ff5de84

Observation 384f4f90-24aa-4b5e-b65a-a65a18e545ec · outbound

This paper cites an unresolved cited work.

VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:32:11.301830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:32:10.532001Z digest=sha256:ecc93fa4fc40e383aa8bc5971618f4f84fb43a2e87fb70fb9d85364e73964656

Observation 66e17569-4f34-432d-8004-e9b5d5f79eb3 · outbound

This paper cites The new concentration of OH − ions is 0.05 + 2x.

VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data The new concentration of OH − ions is 0.05 + 2x

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:32:11.285965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:32:10.536458Z digest=sha256:dca2c56d8833c317fb22c99d028f4e7d4f33dda99b343dc3b54f3234de15ce18

Observation fbd7bfa9-f1b1-4043-865b-7ca01076ae5a · outbound

This paper cites an unresolved cited work.

VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:32:11.264134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:32:10.541145Z digest=sha256:a220334b82426457673be602fe11d59780aa7dccacf98e28d4510b3451146704

Observation 2bc622cf-fd00-4673-b3e6-897db0ed8310 · outbound

This paper cites an unresolved cited work.

VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:32:11.332224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:32:10.547791Z digest=sha256:7afc952446c73ec521d29ee65c656ad70efe08a53e1f95f918aa9cd1c286e87e

Observation f06d1f56-715d-4ca2-ad11-bbd45cb1f4a5 · outbound

This paper cites an unresolved cited work.

VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:32:11.249381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:32:10.552237Z digest=sha256:1936499eb5d6eaa38847ee7cfc756ea1c8bffc7d1e1fab4b73f22bb6e6d58541

Observation 5812b2ad-d33b-49c2-96b2-8a5bea1fe320 · outbound

This paper cites The concentration of OH − ions is much larger than the concentration of M g2+ ions, so we can assume that 0.05 + 2x ≈ 0.05.

VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data The concentration of OH − ions is much larger than the concentration of M g2+ ions, so we can assume that 0.05 + 2x ≈ 0.05

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:32:11.235860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:32:10.556930Z digest=sha256:e9703a5c9d155e9e863113314bd636c8e0af3c94f30ec2c2685888411513af34

Observation 322512e6-abc7-402d-b3d9-65f7513be571 · outbound

This paper cites an unresolved cited work.

VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:32:11.218240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:32:10.562477Z digest=sha256:ec290d70fd2d6116d908070ae90a5947f208bff8a089690295f6023b966bab13

Observation b76409ed-0649-4215-970f-6ca5b5ba1d83 · outbound

This paper cites an unresolved cited work.

VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:32:11.201973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:32:10.567936Z digest=sha256:b656c6009f4bd377a1d7b5b06ffe22d8ae295791253e2cfad58c84ea68a43248

Observation 761fb5c2-021d-4e7d-b382-2610d9bb6683 · outbound

This paper cites an unresolved cited work.

VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:32:11.186911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:32:10.578688Z digest=sha256:cd51edc1f2c26426f428606cdd5add51b1da73ae437186cc34ada041be54234a

Observation cb0428c3-bb1c-45b0-a519-848d3951d4c2 · outbound

This paper cites an unresolved cited work.

VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:32:11.160947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:32:10.585422Z digest=sha256:2e623dd800151ad3329a6cde1eb31f7dacd4c331cac8cc7686210832d2801933

Observation bdd2ab91-3491-4f39-98ad-a9ff796a2e24 · outbound

This paper cites an unresolved cited work.

VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:32:11.145533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:32:10.590512Z digest=sha256:a8267414c27ee810dc5cb54fdb29a763d13f039aeea58734022768bcd7843475

Observation 88b84cf3-a438-40c2-bd11-51d45c95d25e · outbound

This paper cites an unresolved cited work.

VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:32:11.125941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:32:10.595142Z digest=sha256:9a27dd850b06446bc7b837d4bc6d6ae9096ed370490188117dfd8679cbd021da

Observation 1eaeb2dd-5e0e-41a6-903c-b9f4a65a21a3 · outbound

This paper cites an unresolved cited work.

VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:32:11.113211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:32:10.601217Z digest=sha256:252b52e682fc4ebc0b07d96d0db116c92ec5ac220d22cc81554e4a8aca218f03

Observation 64e8d002-503e-4021-8290-71fab4218d6a · outbound

This paper cites an unresolved cited work.

VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:32:11.080047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:32:10.611340Z digest=sha256:389850f07bc65a0860305a317ab7833749aad359646bd04aba9bc6daaff2aa76

Observation aaf88d98-2cc0-4b36-b55a-66d0036d56a2 · outbound

This paper cites an unresolved cited work.

VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data Unresolved cited work

Reference 61

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:32:11.063092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:32:10.618264Z digest=sha256:b278270b299ee615bf812065816571ce80ab570d9982b00ecdba560990c03744

Observation 68430d26-ff17-4e09-8f61-099739d01aea · outbound

This paper cites an unresolved cited work.

VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data Unresolved cited work

Reference 62

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:32:11.098527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:32:10.625867Z digest=sha256:ada684456980cc9c9dc68bce2bdd9161a2b7a09153dff984b7c4fbf1bd8ae8c5

Observation 268a5365-f874-4686-8add-4f4e43e86f16 · outbound

This paper cites an unresolved cited work.

VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:32:11.048738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:32:10.631596Z digest=sha256:e5ce5bfa401c62334c2b1587a29aaf56262a11c39cb1c6e4cc27d0b9049504a9

Observation 3c07f981-5666-4b11-a645-3051bc28e441 · outbound

This paper cites an unresolved cited work.

VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data Unresolved cited work

Reference 64

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:32:11.034584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:32:10.641039Z digest=sha256:d14ce2f0e1b7344280e41709403373eb3e637a8f03c6b4196399c687c725276f

Observation b0b2bd8b-a944-4f56-bb35-0664fa94df3d · outbound

This paper cites an unresolved cited work.

VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data Unresolved cited work

Reference 65

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:32:11.016477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:32:10.649547Z digest=sha256:28b4719acee6c7a009dfaf8ea3f4719625cd0658dafa9b939ffddfb79a1eba93

Observation b08bdc63-ca0b-4407-a177-ce613b804e7c · outbound

This paper cites an unresolved cited work.

VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data Unresolved cited work

Reference 66

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:32:10.987923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:32:10.658995Z digest=sha256:acd93cb7737e1ccf8b8e00844e3c5cd72958ea6539380d6c0ee8e4df139d3654

Observation 75b7e747-9f1b-429a-bd2a-c5a817620a33 · outbound

This paper cites an unresolved cited work.

VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data Unresolved cited work

Reference 67

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:32:10.968751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:32:10.663541Z digest=sha256:ea2c971d14545ab29ebd752567f455e17b09ef3c7d20447f479231d0a4fbc401

Observation 852a70a5-147b-4390-b895-4e2bae37481f · outbound

This paper cites an unresolved cited work.

VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data Unresolved cited work

Reference 68

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:32:10.954897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:32:10.669827Z digest=sha256:8f8a7ab1f9beeaa56e726c0ae15e376688a9b80fcf09f4ba27479b166862dd7f

Observation 98e54ac9-3f21-40b7-b8de-fbcb8b3c414a · outbound

This paper cites an unresolved cited work.

VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data Unresolved cited work

Reference 69

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:32:10.941013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:32:10.675092Z digest=sha256:ba0219d7e0315268a2d51970242a6f7303bc56dfa8fc88390c73bbf600b64e67

Observation 7c2fefdf-a252-4d15-8997-c902c770038f · outbound

This paper cites an unresolved cited work.

VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data Unresolved cited work

Reference 70

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:32:10.925989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:32:10.687292Z digest=sha256:565bed13c3075d0564555684afe9782ca4d8813909fd89efcdb58dd8db23550f

Observation ccfa85dd-937a-4ec2-950a-735615ed8695 · outbound

This paper cites an unresolved cited work.

VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data Unresolved cited work

Reference 71

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:32:10.910362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:32:10.692485Z digest=sha256:e67b546f60cd5d3304afc6a8dc7baa7ef00e4d519a4b4ad3810afb0e8c605e78

Observation 93ff9bc8-ee16-4216-a512-33f57822e0ae · outbound

This paper cites an unresolved cited work.

VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data Unresolved cited work

Reference 72

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:32:10.895485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:32:10.697916Z digest=sha256:efb9d55006a4affdc191efb813343dd3f8a767af2ffa59a172b6b847c5ac4efd

Observation f17120c2-d00c-419e-98f6-9cfa6fde514a · outbound

This paper cites an unresolved cited work.

VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data Unresolved cited work

Reference 73

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:32:10.881094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:32:10.702536Z digest=sha256:de08cb87f794128fba54a231ce89d1ddb3bc46c25d5329ca5fac41a9adbfc3bf

Observation 30c20af5-c890-4ec2-8f71-c4f3eee869bc · outbound

This paper cites an unresolved cited work.

VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data Unresolved cited work

Reference 74

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:32:10.860983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:32:10.706851Z digest=sha256:671b95da4caca91aeaec4bd83cb925fafbc915d2986b818ccc372db2097ab50f

Observation 84e55778-c315-451e-b814-179053aa2980 · outbound

This paper cites an unresolved cited work.

VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data Unresolved cited work

Reference 75

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:32:10.841817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:32:10.712735Z digest=sha256:c2ec8525ee733e9731f6b54b9a361fe2398132609e23adbc39a9b408efffe6f4

Observation cbf2af6d-b626-475b-9452-ca6c595cd88a · outbound

This paper cites an unresolved cited work.

VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data Unresolved cited work

Reference 76

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:32:10.817118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:32:10.717394Z digest=sha256:09d6ba168d2833beee6b7cf385864e518d6e23f5aba254a5d203c32341a4f1c4

Observation 1e3d19b2-a24b-4a09-b637-fc649b63066f · outbound

This paper cites an unresolved cited work.

VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data Unresolved cited work

Reference 77

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:32:10.800279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:32:10.722142Z digest=sha256:e92f455ec02cc2c4e8ff61b8bb2f09d7c5fb120754c3c893f481566a6bf06b29

Observation fa9c7ed9-fe6a-4a65-87a6-8e7fcb4dd5b7 · outbound

This paper cites an unresolved cited work.

VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data Unresolved cited work

Reference 78

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:32:10.787366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:32:10.726615Z digest=sha256:d204c844c010ff50890a9e1d2c550e9bd23f1e7221d904135429f559350b3d9c

Observation 95df2d65-3585-46ec-adfa-8e1d438d6a4d · outbound

This paper cites which of the following [X] is correct.

VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data which of the following [X] is correct

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:32:10.773615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T14:32:10.732509Z digest=sha256:e4405e43ca2db7ff63812c5425f78b743adcbded8d5cbd95d13fe37fb927ec48

Pith citing papers

Observation bd0b5798-ed0c-42bc-b161-01f3c06fbd77 · inbound

GM-PRM: A Generative Multimodal Process Reward Model for Multimodal Mathematical Reasoning cites this paper.

GM-PRM: A Generative Multimodal Process Reward Model for Multimodal Mathematical Reasoning VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T00:59:54.314068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:59:54.314068Z digest=sha256:71cdda911951fcad3074748b663d29faf86e1a942997c4ec757659f69bf84607

Observation f4f44dfe-ee49-448d-9142-681886b99854 · inbound

MASPRM: Multi-Agent System Process Reward Model cites this paper.

MASPRM: Multi-Agent System Process Reward Model VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T07:56:47.849322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:56:47.849322Z digest=sha256:7e95e5d3f8426ac778de4d26bcb1d3e4e19f0954850752f131a7937eb0ac6dc5

Observation 7aa50ad5-9f73-46db-b0a6-7f1e5a6f9ccd · inbound

Verifiable Counterfactual Supervision for Process Reward Models cites this paper.

Verifiable Counterfactual Supervision for Process Reward Models VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T05:55:30.057650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T19:24:18.022593Z digest=sha256:9867bdcc944a504bc75dce107832e443466ec75d10b52a41c5e8719eec06598b

Observation 0c86b742-d9fc-4919-9abf-94e64605482e · inbound

PDCR: Perception-Decomposed Confidence Reward for Vision-Language Reasoning cites this paper.

PDCR: Perception-Decomposed Confidence Reward for Vision-Language Reasoning VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:22:50.520864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T19:20:32.435135Z digest=sha256:88c81b5f671a73516c342bec2780ed328baefaa46f8e25e1d981516d20835a8b

Observation e3a156cf-aa5f-482c-a9b6-bb04f1e91f38 · inbound

STRIDE: Learnable Stepwise Language Feedback for LLM Reasoning cites this paper.

STRIDE: Learnable Stepwise Language Feedback for LLM Reasoning VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-20T20:49:00.768167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T20:47:16.236629Z digest=sha256:baa522a688f7bfad7a1db88df0599817fcd31dc74cad3f5a594fc7a85d42c6f6

Observation a32409a8-59b5-4a82-a65f-d99e872ff374 · inbound

Test-Time Scaling for Small VLMs on Multilingual Visual MCQ cites this paper.

Test-Time Scaling for Small VLMs on Multilingual Visual MCQ VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-13T03:00:51.318412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:00:51.318412Z digest=sha256:cfafc258921debe8eb233c3c6b42b7547937a11ce83e4caec7d4b6752d4486f6