Pith. sign in

Paper Citation Record · LEDGER

Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces

As of 19 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 0 inbound Pith citation observations for arXiv:2608.12585.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.12585 v1

Coverage vector

measured 60 of 60 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:10:31.113654Z

measured 60 of 60 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

60 of 60 outbound references displayed

  • verified exact0
  • verified fuzzy36
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f21e2386-3b3c-42cf-bd06-8a7f791376da · outbound

This paper cites optimize_anything: A Universal API for Optimizing any Text Parameter.

Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces optimize_anything: A Universal API for Optimizing any Text Parameter

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-08-16T00:10:31.155688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:10:30.845169Z digest=sha256:0d64157020a1f110854f983044763175a3505e4ec9ce9e8ab9cb17427cda9216

Observation d59f0149-2ae2-4964-8da2-993b1d04fc2e · outbound

This paper cites Claude opus 4.6 system card.

Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces Claude opus 4.6 system card

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:10:32.200939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:10:30.851238Z digest=sha256:8e1566d6dfbc668cc8ea06ddfbc88b6b3363ae13e11c6e3172f1337827dec6e3

Observation 718f155c-2323-4307-9406-a9031faefdf9 · outbound

This paper cites Claude sonnet 4.6 system card.

Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces Claude sonnet 4.6 system card

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:10:32.186714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:10:30.857010Z digest=sha256:33454cc2681a78bbac3f3c72e3113b796ec1cc3c14b71854d1c7aa9ffb033e01

Observation df871948-4df4-4ede-9831-4217e1e92d9b · outbound

This paper cites Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation.

Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T00:10:30.862323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:10:30.862323Z digest=sha256:82ad3b7c54c20bf2b3c51a51bbbce76549ab830953bc8a4b69eac094faf1ac89

Observation 1d7c96cf-3105-44b8-9885-9c6309d84dc6 · outbound

This paper cites Nudging the boundaries of llm reasoning.

Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces Nudging the boundaries of llm reasoning

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:10:32.173193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:10:30.867093Z digest=sha256:2c950537eb67dcb495a80a7dc10e36ca8bdfac73b7f21147e2f0087510284089

Observation 15ce2d86-1359-42a7-807e-5a850116d3e7 · outbound

This paper cites Stop summation: Min-form credit assignment is all process reward model 17 needs for reasoning.

Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces Stop summation: Min-form credit assignment is all process reward model 17 needs for reasoning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T00:10:30.871471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:10:30.871471Z digest=sha256:8adc59c626b663a8e18d37d159975dd379ed0631750b268b5806e7170b4a182d

Observation 2f2c8ee3-ce2b-41a5-a398-852395b08c39 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces Training Verifiers to Solve Math Word Problems

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T00:10:30.876413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:10:30.876413Z digest=sha256:0c0430671dbdfbe0a33214ae386c7dec8096760d9a34d259436a57f7ae2d5215

Observation 921030ce-a718-4456-8d8e-d5a8ec94c920 · outbound

This paper cites DeepSeek-V4-Pro model card.

Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces DeepSeek-V4-Pro model card

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:10:32.158523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:10:30.880998Z digest=sha256:013f3c054d86823bad3e68d61ee3d8ff702a04fba8d6fb246076eee6bc4b3ea9

Observation b5839ea3-7483-4d51-aa0c-81c9046863e4 · outbound

This paper cites Beyond Benchmarks: MathArena as an Evaluation Platform for Mathematics with LLMs.

Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces Beyond Benchmarks: MathArena as an Evaluation Platform for Mathematics with LLMs

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T00:10:30.885694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:10:30.885694Z digest=sha256:253ed60efe5c4e3655d773c8de9fcdb2b15149f7b25644df59ecc27e78adfc8a

Observation 6fcde8c6-5046-47a9-a28e-b47072f85a81 · outbound

This paper cites Tenenbaum, and Igor Mordatch.

Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces Tenenbaum, and Igor Mordatch

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T00:10:30.890190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:10:30.890190Z digest=sha256:3ce03efc0365fb0bdb9f25c697fe754fe2102267a1d67a2922bad22d910fd03e

Observation 608c3d74-1f40-4785-bd73-e78d67664c51 · outbound

This paper cites Gemma 4 Technical Report.

Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces Gemma 4 Technical Report

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T00:10:30.894249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:10:30.894249Z digest=sha256:b1d23bc09df7f4e20041465cf375fda23d83f5001a3fc8bc66a2030e61d38cb0

Observation d9ca9d02-271b-4223-957b-f5d9acf8d133 · outbound

This paper cites GLM-5: from Vibe Coding to Agentic Engineering.

Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces GLM-5: from Vibe Coding to Agentic Engineering

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T00:10:30.898657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:10:30.898657Z digest=sha256:72d2d5304147e183841d18bc085395e111184441638aa009ce26eb40cc425173

Observation 355378b9-a272-4d95-a592-0eccf02c49b4 · outbound

This paper cites Gemini 3.1 pro model card.

Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces Gemini 3.1 pro model card

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:10:32.134679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:10:30.903491Z digest=sha256:e68e2d30733257794b0d6b1ec8775b11ed0e5c90afd74b2e7e124908bb96824e

Observation 049203f7-36ae-4393-98ab-9b58fd1be231 · outbound

This paper cites A Survey on LLM-as-a-Judge.

Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces A Survey on LLM-as-a-Judge

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T00:10:30.907639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:10:30.907639Z digest=sha256:42158dceae32f2d7c03d5c49cb00ab499189e7630fc09436c59f617bfdd87828

Observation 071ccec1-df26-4011-aa0d-39ed69a4c597 · outbound

This paper cites an unresolved cited work.

Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-16T00:10:32.120673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:10:30.912743Z digest=sha256:90457cf91096c5fdac318aac35e2aa235ec4d55609aaa91eecf72754482ec14b

Observation e94f6ab1-8757-4ae8-b04d-001db078b1d0 · outbound

This paper cites Reinforcement learning via self-distillation.

Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces Reinforcement learning via self-distillation

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:10:32.106361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:10:30.917621Z digest=sha256:8ba431c96327967c26ca3b70ff3bb04c2b6bab0655e3b5671a955fc767d06d3d

Observation b71ddfae-7498-4b82-9ef8-911b1c200313 · outbound

This paper cites Let’s verify step by step.International Conference on Learning Representations (ICLR), 2024.

Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces Let’s verify step by step.International Conference on Learning Representations (ICLR), 2024

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:10:32.092394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:10:30.921912Z digest=sha256:0c4bedf38896413e506766b1427826c28ab1b43acb5990451f2ef134184a1e9f

Observation d41f5618-2d4b-4012-95d0-7a734068d122 · outbound

This paper cites MiniMax-M2.7 model card.

Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces MiniMax-M2.7 model card

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:10:32.077650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:10:30.926091Z digest=sha256:043ce1c343a528904ec758d808c5ac1c146a2bf3daaaec433a264a4a84f4ce1e

Observation 6c738d28-fb2d-4b1d-b2a5-44516d3b4401 · outbound

This paper cites MiniMax-M3 model card.

Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces MiniMax-M3 model card

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:10:32.061828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:10:30.930191Z digest=sha256:9c3b8c8b2aebbbbd6a3a4bab39939a830a5bd51d29d5ed444ec46c6ef4d6cd02

Observation f5098b60-f70e-42da-ad53-0e5297b5902f · outbound

This paper cites Kimi K2.6 model card.

Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces Kimi K2.6 model card

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:10:32.046099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:10:30.934393Z digest=sha256:b138ad88dd23a67aa71fa835e617759b7d1d969c33e5a4a43854af0bdaaff32c

Observation c74fd16b-adaf-426c-bd8a-791b7a00c02a · outbound

This paper cites NVIDIA Nemotron 3: Efficient and Open Intelligence.

Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces NVIDIA Nemotron 3: Efficient and Open Intelligence

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T00:10:30.938490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:10:30.938490Z digest=sha256:743772f4a6583e8ea986d047dc7b65943234c0612357235be4cce6aae3bef5f8

Observation 652f0cc2-6db9-434e-ae6f-6a3ff8a1d8e8 · outbound

This paper cites gpt-oss-120b & gpt-oss-20b Model Card.

Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces gpt-oss-120b & gpt-oss-20b Model Card

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T00:10:30.942860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:10:30.942860Z digest=sha256:ba2ef69e472056d2dd222ac9940307c5b1eefa06e2dcc0a6048f73fea939720a

Observation 1c554cdf-237e-4df1-9077-32cdc09277ed · outbound

This paper cites Introducing gpt-5.4.

Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces Introducing gpt-5.4

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:10:32.032103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:10:30.947282Z digest=sha256:3e19e1668bc3ca74483fadbf5cc0a6b2f19c0defe042dc0f3b8439c5d74c79cf

Observation c56f2cf5-2033-48c3-8e7b-4a34d83f6e04 · outbound

This paper cites gpt-5.4 model.

Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces gpt-5.4 model

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:10:32.017031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:10:30.951744Z digest=sha256:7f51c04f647f7497a0819b608509cff690e366775519ed25dc153ac166aacd64

Observation 45587909-2574-45cd-b9f3-600680787b51 · outbound

This paper cites Hard2Verify: A step-level verification benchmark for open-ended frontier math.

Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces Hard2Verify: A step-level verification benchmark for open-ended frontier math

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:10:32.001905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:10:30.956022Z digest=sha256:f7d5f867a28e8b0a5adfceab323f9c7a7fa18410bd056fffca9080b42d3ada81

Observation abb66afa-3208-48ab-85c5-018c01dfbd2f · outbound

This paper cites Qwen3.6-27B: Flagship-level coding in a 27B dense model, April 2026.

Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces Qwen3.6-27B: Flagship-level coding in a 27B dense model, April 2026

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T00:10:30.961108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:10:30.961108Z digest=sha256:35d601933e906797deb4d8f729df1ffc2f26bd18ad00e3c772da41dbee03d0f2

Observation b02523f1-3375-4498-8fe1-e96d5ccf6d4b · outbound

This paper cites PRISM: Pushing the frontier of deep think via process reward model-guided inference.arXiv preprint arXiv:2603.02479, 2026.

Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces PRISM: Pushing the frontier of deep think via process reward model-guided inference.arXiv preprint arXiv:2603.02479, 2026

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T00:10:30.965881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:10:30.965881Z digest=sha256:6fdf8671a9e23c199c38725bb1f9adbdaf5196b04707c9fcdbcaa903d8e28ddc

Observation 8846d2bd-3617-4b6e-8b5c-954153477856 · outbound

This paper cites Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Models.

Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T00:10:30.971538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:10:30.971538Z digest=sha256:18b15322e0b0ca40d73326ac0d773a4da7e52a09cb8792bc0cb1f8a7d1a94125

Observation 58dcdbb4-0c40-44ca-85a1-1dedf7b3057b · outbound

This paper cites GLM-5.2-FP8modelcard.

Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces GLM-5.2-FP8modelcard

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:10:31.979041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:10:30.975946Z digest=sha256:eadfab070692526964ee345cf31a51696796e3acb29bb43671e0ec128729d030

Observation ed9d793f-42e4-49e3-ae45-4ed544b5e97c · outbound

This paper cites Scaf-grpo: Scaffolded group relative policy optimization for enhancing llm reasoning, 2026.

Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces Scaf-grpo: Scaffolded group relative policy optimization for enhancing llm reasoning, 2026

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T00:10:30.980189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:10:30.980189Z digest=sha256:01a0d94d3b90dc078f0d0c534d759af1b0ae7d877762c4abfd2f96aa09f1d451

Observation 1f52b0c4-b61d-44cd-b11d-8b6f1e243439 · outbound

This paper cites ProcessBench: Identifying Process Errors in Mathematical Reasoning.

Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces ProcessBench: Identifying Process Errors in Mathematical Reasoning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T00:10:30.984349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:10:30.984349Z digest=sha256:99ccd65c24b89a166bccb024aa3aef6d22f0ff83cbe08e44315ccd35d4d91b67

Observation 9394466a-8ec3-4354-901f-c5ea6c9bc3bb · outbound

This paper cites Xing, Hao Zhang, Joseph E.

Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces Xing, Hao Zhang, Joseph E

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:10:31.965569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:10:30.988727Z digest=sha256:0a31a66cf98fa76bbce95fd7f10abc8e36224df69da52b658853759eb6e9e36c

Observation a971ac6c-432b-472b-bd3b-4cf6c2779531 · outbound

This paper cites statement_refs.

Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces statement_refs

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T00:10:30.993024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:10:30.993024Z digest=sha256:30528916fd9234bea0f5573a7909144c185bd2464b3b1c995f8915488b432f20

Observation b0df11bf-611a-4a02-a9e6-94b7f7d283ec · outbound

This paper cites an unresolved cited work.

Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-16T00:10:31.951962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:10:30.997123Z digest=sha256:e75c3c7f00e6bef65a7e10154d74b70d43159fadac9e0ea5bfc6e3332b8a11d1

Observation 2f69b421-adae-41b0-911c-8847c0d5b4c1 · outbound

This paper cites The trace is divided into segments, each prefixed with [STEP-x] where x is an integer indicating the ordinal position of each segment in the trace.

Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces The trace is divided into segments, each prefixed with [STEP-x] where x is an integer indicating the ordinal position of each segment in the trace

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:10:31.937454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:10:31.001397Z digest=sha256:c40107ab5831e5b01e8833a176ac53a6717d2f6040f9739a85e1fde5d2b88d05

Observation 80654fb4-1e54-4d9e-aefa-180d3ce11061 · outbound

This paper cites ## Primary objective 20 Judge the *weaknesses* of the provided reasoning trace by pointing to **specific bad reasoning moves**.

Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces ## Primary objective 20 Judge the *weaknesses* of the provided reasoning trace by pointing to **specific bad reasoning moves**

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:10:31.922093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:10:31.005951Z digest=sha256:de9c5f267c8bd12eeafc6b0fc4a54f95722b7ef62e42ce885f1019fb822108c4

Observation efa9f52b-58b2-4615-8f97-81dfbab87af5 · outbound

This paper cites In [STEP-14], the trace states.

Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces In [STEP-14], the trace states

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:10:31.908517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:10:31.010544Z digest=sha256:e13ad3163232084dab7c6f3533b14a03b8c850a9e350eeb9682e4ab8fecba90b

Observation 10587715-23da-49ee-9bd5-391da0d2b361 · outbound

This paper cites an unresolved cited work.

Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-16T00:10:31.894262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:10:31.014658Z digest=sha256:119a5c6d1df385b1b443294c0e4ccb82b49be584a6b751c70be3a51a36bccf3c

Observation f0004429-d15e-4fee-b61c-9d99b35cc430 · outbound

This paper cites Could this comment apply to a totally different problem with no edits?.

Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces Could this comment apply to a totally different problem with no edits?

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:10:31.880632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:10:31.019716Z digest=sha256:14edae456118c83dfe5e5c928ddec45c270579aa4c970053b8244153e0e343bb

Observation ece76177-47c8-4cfd-862a-7c0bb2726d7e · outbound

This paper cites Adopt it only if the trace itself supports the claim.

Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces Adopt it only if the trace itself supports the claim

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:10:31.867247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:10:31.024221Z digest=sha256:5e535618d514f93cc54411ed7b2d33bba6b795a09fb54489c816aa6f8a1035a7

Observation a16b8b3f-a662-4fc7-853d-55e52820cac0 · outbound

This paper cites Supplement, don’t average: fold in supporting evidence from the other auditors, but never replace a specific claim with a vaguer paraphrase.

Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces Supplement, don’t average: fold in supporting evidence from the other auditors, but never replace a specific claim with a vaguer paraphrase

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:10:31.852478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:10:31.028309Z digest=sha256:54028ced39e9186b76f19fe94aa3dd717dc390686c95c229c304fca8a6bd1761

Observation db132aaf-4ef7-4c85-925e-40affeafe5a8 · outbound

This paper cites Include genuine defects that NO auditor raised.

Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces Include genuine defects that NO auditor raised

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:10:31.838266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:10:31.032482Z digest=sha256:f33151efa1c8357ef02af15f3e5dccc821626cb6ff5dc4b88ab1bbe0130b73ec

Observation 3df25981-5aaa-4751-b576-3a4f028053aa · outbound

This paper cites an unresolved cited work.

Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-16T00:10:31.824963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:10:31.037610Z digest=sha256:804025c6710a53d0fbac8c616673fbb98ebf7f7e72964908e2c4059256838253

Observation 5b170482-5faa-4775-bc88-9cad5c57ccc2 · outbound

This paper cites Output format.

Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces Output format

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:10:31.811759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:10:31.042042Z digest=sha256:4431b1f469f204745db33cd3f2346be5327dc1d56aaed6bef257dd2fa28d8317

Observation 337f6346-afd6-439b-b515-d92f0224747c · outbound

This paper cites - **Do NOT abstract away specifics.** When multiple judges describe the same issue at different levels of detail, use the MOST SPECIFIC description as the base.

Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces - **Do NOT abstract away specifics.** When multiple judges describe the same issue at different levels of detail, use the MOST SPECIFIC description as the base

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:10:31.797545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:10:31.046268Z digest=sha256:3945fdadc22c8f75a78e94534cb62026a5e500a2e63cbf434e2bc5a1ffd23b77

Observation 2fd8e08a-991f-4841-ab96-43a427df2997 · outbound

This paper cites Use that as the canonical wording.

Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces Use that as the canonical wording

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:10:31.782210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:10:31.050736Z digest=sha256:55c62b569e85bdc30b6a48329d40c3227bb5064b94082571fa1cc59f3609c1f7

Observation 85b107e2-6636-445a-8184-abcbdfbef682 · outbound

This paper cites an unresolved cited work.

Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-16T00:10:31.768133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:10:31.055480Z digest=sha256:bfac31b74a42fb7c6377e94dc76d4a72f802d679da8e200b47b819cc61b9a464

Observation 5839ec36-68d2-434e-96ce-f06906d53a6f · outbound

This paper cites an unresolved cited work.

Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-16T00:10:31.754766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:10:31.059503Z digest=sha256:d0bf07e5b7682ba34564a5946d5d5a49a16fbbe76feb1c95ae4bb08994cb995f

Observation 2fc9325b-0855-46ca-9872-418300376d7c · outbound

This paper cites Could this description apply to a totally different problem with no edits?.

Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces Could this description apply to a totally different problem with no edits?

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:10:31.740655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:10:31.063645Z digest=sha256:9054bce70e15b756fdcb010a5df270b49a9225bc58a9b579a5f0cbe415662a4e

Observation f3ac5ff6-e02a-4227-84b4-8998b9a25e5b · outbound

This paper cites - Prioritise panelists involved in unresolved disagreements.

Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces - Prioritise panelists involved in unresolved disagreements

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:10:31.725592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:10:31.068235Z digest=sha256:fe2213673965bbd3b44c91274b35f2571a5acce8a753ee6d47352655c3e0587c

Observation 320f0778-34bf-490d-ae79-59d40cf701ae · outbound

This paper cites - Ask them to clarify, defend, or concede specific points raised by others.

Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces - Ask them to clarify, defend, or concede specific points raised by others

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:10:31.710823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:10:31.073066Z digest=sha256:85469175a91066ec862157bbe14526ac923d574c2dba57395c6914c372afcf5f

Observation bba92a13-aa3a-4b88-9dc8-e5560c6efb42 · outbound

This paper cites should_terminate.

Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces should_terminate

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:10:31.695060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:10:31.077316Z digest=sha256:e826860c5363d599aa61b48e566ceee3a3d5242f5fe0ef8f8ffaea35024b998a

Observation a0395b97-0c0a-4144-8d1a-ce242be9406b · outbound

This paper cites consensus_state.

Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces consensus_state

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:10:31.680081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:10:31.081932Z digest=sha256:e37b6ec8fd1e7053b283b2bce3caa360a8346539518db3efe434e361718dde58

Observation 5f1c410a-8d86-43cc-b243-68e660292f29 · outbound

This paper cites an unresolved cited work.

Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-08-16T00:10:31.666002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:10:31.086343Z digest=sha256:782a3cdadd16bbf619679fb11bff37e0559a4dd787dcf65606988a886b11eb3b

Observation a581ee97-b37d-4741-9b93-859143dbcbd6 · outbound

This paper cites Additional rules.

Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces Additional rules

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:10:31.651348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:10:31.090654Z digest=sha256:a6a14a72e379f0d339ebe74a59ff7a36e427a8a30d68da87c9fab2c4713aae19

Observation 1fe3810e-4cc8-43f0-a3c4-b3960c8c83bf · outbound

This paper cites The per-problem analysis addresses dependence among traces from the same problem, but broader generalization remains untested.

Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces The per-problem analysis addresses dependence among traces from the same problem, but broader generalization remains untested

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:10:31.636380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:10:31.097102Z digest=sha256:12a1535936b1379ca9245360593c93c3523789e26258f2b4c0b8971607c15557

Observation fdeeea0b-efc2-45aa-bdb8-3b90e1f13d4d · outbound

This paper cites Some improvement over the original answers may therefore come from making a second attempt rather than from the feedback itself.

Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces Some improvement over the original answers may therefore come from making a second attempt rather than from the feedback itself

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:10:31.621376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:10:31.101215Z digest=sha256:8a7e993b18684620744bfcf2c0429c1611146e94828987c3ebaf53cfc4da845b

Observation befefdbc-7159-4b2c-93fc-5c439c999781 · outbound

This paper cites The experiment consequently measures the complete rich-feedback package; it does not isolate the effect of either field or distinguish explanation from a supplied fix.

Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces The experiment consequently measures the complete rich-feedback package; it does not isolate the effect of either field or distinguish explanation from a supplied fix

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:10:31.606485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:10:31.105377Z digest=sha256:39136e890a68cde154c5a2204335cf41b06cc0f4d0c295d4796886ee1b47d61c

Observation 2a3720f9-c740-45ae-a478-c519493437a1 · outbound

This paper cites Failures caused by omissions may also escape detection.

Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces Failures caused by omissions may also escape detection

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:10:31.591533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:10:31.109511Z digest=sha256:b5cfe2c19834181fb90a3710b1913dcd8c087f3b8a16f0fd8c6803a5f7fa5ea6

Observation 3cd5c726-b836-42e8-a4ef-2e0cf111611e · outbound

This paper cites The external AIME answer key validates the correctness of retry outcomes, not the factual accuracy or severity of individual defect descriptions.

Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces The external AIME answer key validates the correctness of retry outcomes, not the factual accuracy or severity of individual defect descriptions

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:10:31.576459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:10:31.113654Z digest=sha256:caa545671b02beecfe5cce7c5c7e45fda574d2ac4bda25be1579270822857aec

Pith citing papers

No inbound Pith citation observations are available.