Pith. sign in

Paper Citation Record · LEDGER

When the Judge Should Not Decide: Evidence-Locked, Non-Compensatory Selection Bounds LLM-Judge Failure in Reasoning Pipelines

As of 16 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 0 inbound Pith citation observations for arXiv:2608.07813.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.07813 v1

Coverage vector

measured 16 of 16 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T04:18:11.112009Z

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

16 of 16 outbound references displayed

  • verified exact0
  • verified fuzzy2
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cfb53ac5-7f8d-4fd9-bf39-5e8b2690071c · outbound

This paper cites Let's Verify Step by Step.

When the Judge Should Not Decide: Evidence-Locked, Non-Compensatory Selection Bounds LLM-Judge Failure in Reasoning Pipelines Let's Verify Step by Step

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T04:18:11.048811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:18:11.048811Z digest=sha256:f1cc8418015d91be7c14b753c89a67152b61848fc3443fd399e89511079c6809

Observation 32bcac5d-3cb4-4e06-86ba-4e576eae9dcb · outbound

This paper cites Factscore: Fine-grained atomic evaluation of factual precision in long form text generation.

When the Judge Should Not Decide: Evidence-Locked, Non-Compensatory Selection Bounds LLM-Judge Failure in Reasoning Pipelines Factscore: Fine-grained atomic evaluation of factual precision in long form text generation

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:18:11.385320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T04:18:11.059931Z digest=sha256:341100ab2d40394745a3872402c49fc2ce6370c6227ce623c9b476e85841ea11

Observation fdf26fe2-76bc-4506-95a7-fb8b3621f7d1 · outbound

This paper cites doi: 10.18653/v1/2023.emnlp-main.741.

When the Judge Should Not Decide: Evidence-Locked, Non-Compensatory Selection Bounds LLM-Judge Failure in Reasoning Pipelines doi: 10.18653/v1/2023.emnlp-main.741

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T04:18:11.065092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:18:11.065092Z digest=sha256:294940be4e0e3164e26c45818324a1e4a20d746664ccc9240d1219734d171d4a

Observation 5ead05c0-3feb-4beb-8912-c0bcb3e264be · outbound

This paper cites WebGPT: Browser-assisted question-answering with human feedback.

When the Judge Should Not Decide: Evidence-Locked, Non-Compensatory Selection Bounds LLM-Judge Failure in Reasoning Pipelines WebGPT: Browser-assisted question-answering with human feedback

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T04:18:11.070300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:18:11.070300Z digest=sha256:ecd2e0e741e3a412191f89576da29b315c51f29437e185dac8bd7814e1d546c6

Observation 928b9da8-5f40-40aa-971a-76b986046821 · outbound

This paper cites LLM Evaluators Recognize and Favor Their Own Generations.

When the Judge Should Not Decide: Evidence-Locked, Non-Compensatory Selection Bounds LLM-Judge Failure in Reasoning Pipelines LLM Evaluators Recognize and Favor Their Own Generations

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T04:18:11.076461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:18:11.076461Z digest=sha256:0ab27eddee7ca5ce5a886692e4e474496c5993b73cfe31592db76dfadf798d8e

Observation a61de5b6-5a26-4e55-84ce-97b33d266fbe · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

When the Judge Should Not Decide: Evidence-Locked, Non-Compensatory Selection Bounds LLM-Judge Failure in Reasoning Pipelines DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T04:18:11.081607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:18:11.081607Z digest=sha256:14d4cca4421404aeb56f3c076d38a64d2e96736f4249df5cdcd3f28313c54b55

Observation f5e30d92-2f1f-40ae-8c06-6fc1ad964346 · outbound

This paper cites Large Language Models are not Fair Evaluators.

When the Judge Should Not Decide: Evidence-Locked, Non-Compensatory Selection Bounds LLM-Judge Failure in Reasoning Pipelines Large Language Models are not Fair Evaluators

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T04:18:11.086784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:18:11.086784Z digest=sha256:b20193c0988a6957727724de43e190d85be9717ee6d48391965b0c680f0b31f1

Observation 3c0b3d1d-a3da-4811-8d8c-9bf99f87963d · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

When the Judge Should Not Decide: Evidence-Locked, Non-Compensatory Selection Bounds LLM-Judge Failure in Reasoning Pipelines Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T04:18:11.091647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:18:11.091647Z digest=sha256:02b961029d9b6956c7317b27183e72c38aee66f77f919639b90ca94f44c18699

Observation 15392555-41bc-488d-9c08-e7e00bd8c9fb · outbound

This paper cites Hotpotqa: A dataset for diverse, explainable multi-hop question answering.

When the Judge Should Not Decide: Evidence-Locked, Non-Compensatory Selection Bounds LLM-Judge Failure in Reasoning Pipelines Hotpotqa: A dataset for diverse, explainable multi-hop question answering

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:18:11.368928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T04:18:11.097139Z digest=sha256:2b1e5bce0d0a37d7858e89cd4df9ba56b601aafe8da6b70e7d3bd72ebdb7e402

Observation 12235552-9bb1-4cdf-aa37-c5782f4aaa01 · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

When the Judge Should Not Decide: Evidence-Locked, Non-Compensatory Selection Bounds LLM-Judge Failure in Reasoning Pipelines ReAct: Synergizing Reasoning and Acting in Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T04:18:11.107248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:18:11.107248Z digest=sha256:085c0d378f03887cd5d08ecdc56d2f3584afce44da4bd062197dbad1823f6584

Observation 4b695ca7-2953-4cd4-8267-83efb4937583 · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

When the Judge Should Not Decide: Evidence-Locked, Non-Compensatory Selection Bounds LLM-Judge Failure in Reasoning Pipelines Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T04:18:11.112009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:18:11.112009Z digest=sha256:9d818d9944bba9820163905718f880dfa8e0ea734c59bebf2ce5530eebbe84e5

Observation 741af49e-e7b8-48f2-a4c7-fa7288e9f527 · outbound

This paper cites doi: 10.18653/v1/D18-1259.

When the Judge Should Not Decide: Evidence-Locked, Non-Compensatory Selection Bounds LLM-Judge Failure in Reasoning Pipelines doi: 10.18653/v1/D18-1259

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-11T04:18:11.102124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:18:11.102124Z digest=sha256:d6f39adf23c2fea6cf70078ced3b6aa938e32d9cb0550a3d1d1505cdf96974a2

Observation c6f116b2-f598-4d8d-aa97-87f9135693f8 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

When the Judge Should Not Decide: Evidence-Locked, Non-Compensatory Selection Bounds LLM-Judge Failure in Reasoning Pipelines Training Verifiers to Solve Math Word Problems

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-11T04:18:11.043369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:18:11.043369Z digest=sha256:4b45cc2111d952725cff225c94331f51060bab6b6575af8815bf6062516586bd

Observation 50a2fe24-2a3a-43d7-b97e-4cd1651ea685 · outbound

This paper cites FacTool: Factuality Detection in Generative AI -- A Tool Augmented Framework for Multi-Task and Multi-Domain Scenarios.

When the Judge Should Not Decide: Evidence-Locked, Non-Compensatory Selection Bounds LLM-Judge Failure in Reasoning Pipelines FacTool: Factuality Detection in Generative AI -- A Tool Augmented Framework for Multi-Task and Multi-Domain Scenarios

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-11T04:18:11.038075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:18:11.038075Z digest=sha256:a664a9fd69623b962667f777c17b43ce9a9489d81ad8857d02b05a6fff25d4d9

Observation 705fdad2-3aba-444a-9f4c-cd2385e72fc3 · outbound

This paper cites Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection.

When the Judge Should Not Decide: Evidence-Locked, Non-Compensatory Selection Bounds LLM-Judge Failure in Reasoning Pipelines Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-11T04:18:11.032070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:18:11.032070Z digest=sha256:09f0a940315391f948a8089270dc1afe69389ecfd3254832715dc94807b256e8

Observation d7be8218-0e2a-4828-bc27-b9e00949bb21 · outbound

This paper cites SWIRL: A Staged Workflow for Interleaved Reinforcement Learning in Mobile GUI Control.

When the Judge Should Not Decide: Evidence-Locked, Non-Compensatory Selection Bounds LLM-Judge Failure in Reasoning Pipelines SWIRL: A Staged Workflow for Interleaved Reinforcement Learning in Mobile GUI Control

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-11T04:18:11.054622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:18:11.054622Z digest=sha256:929a976281ce95dfcd2087ecf44a5ea98d0e6a5767f6954479b65591cab3d97e

Pith citing papers

No inbound Pith citation observations are available.