Pith. sign in

Paper Citation Record · LEDGER

Likelihood scoring for continuations of mathematical text: a self-supervised benchmark with tests for shortcut vulnerabilities

As of 6 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 0 inbound Pith citation observations for arXiv:2605.10810.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.10810 v2

Coverage vector

measured 50 of 50 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-11T11:50:26.030339Z

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

50 of 50 outbound references displayed

  • verified exact37
  • verified fuzzy6
  • unresolved2
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch4

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 61a5cee9-16e9-43e0-bac9-47ee9a802aa8 · outbound

This paper cites Tülu 3: Pushing Frontiers in Open Language Model Post-Training.

Likelihood scoring for continuations of mathematical text: a self-supervised benchmark with tests for shortcut vulnerabilities Tülu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T17:22:42.172403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T17:20:39.969447Z digest=sha256:09c4ad72663ab824229e7d85b58f2be7daadcbf77f0d22ea2f9304584b860b78

Observation ecdb4662-0742-48f9-925d-de1e0267996e · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Likelihood scoring for continuations of mathematical text: a self-supervised benchmark with tests for shortcut vulnerabilities DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-19T17:22:41.954792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T17:20:39.969447Z digest=sha256:d182979ace035ddc4173ea2bb331f772c8614ce447efd8761f8b0e2cbbfcfdce

Observation d840f19f-a558-438d-bdcc-3e49d58a35f0 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Likelihood scoring for continuations of mathematical text: a self-supervised benchmark with tests for shortcut vulnerabilities DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-19T17:22:41.936034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T17:20:39.969447Z digest=sha256:21e5575ea8dbe41e5ac2d72bb6afa20c3769b69041bbc552f7a75a8b7fe44162

Observation 3ead951c-6459-4ee8-a072-3b0502df03ae · outbound

This paper cites LiveBench: A Challenging, Contamination-Limited LLM Benchmark.

Likelihood scoring for continuations of mathematical text: a self-supervised benchmark with tests for shortcut vulnerabilities LiveBench: A Challenging, Contamination-Limited LLM Benchmark

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-19T17:22:42.004197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T17:20:39.969447Z digest=sha256:cb9c712fa2e1d2937af8522fa7af4829a7ee48ad55c24d56a56af59459da5717

Observation d5124c3e-c4ec-4b63-bc29-25be03d5747a · outbound

This paper cites AntiLeakBench: Preventing Data Contamination by Automatically Con- structing Benchmarks with Updated Real-World Knowledge.

Likelihood scoring for continuations of mathematical text: a self-supervised benchmark with tests for shortcut vulnerabilities AntiLeakBench: Preventing Data Contamination by Automatically Con- structing Benchmarks with Updated Real-World Knowledge

Reference 5

Resolution
verified exact
doi, observed 2026-05-19T17:22:41.656914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T17:20:39.969447Z digest=sha256:6bfd5b101cc84d6c6cd0ef05a139579b0cae510b54d299d0e3d4351de299c248

Observation efed21f2-50f5-4909-9a85-fe123c82a212 · outbound

This paper cites Louis, G.

Likelihood scoring for continuations of mathematical text: a self-supervised benchmark with tests for shortcut vulnerabilities Louis, G

Reference 6

Resolution
malformed identifier
doi_truncated, observed 2026-05-19T17:22:41.650016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T17:20:39.969447Z digest=sha256:483325d6436567584b9fa1b9346cbccc87fe3bb640e0668f635fbef0285ac68d

Observation cf2b9bf9-1871-443b-93d0-c6a495e08fc7 · outbound

This paper cites Scaling Laws for Reward Model Overoptimization.

Likelihood scoring for continuations of mathematical text: a self-supervised benchmark with tests for shortcut vulnerabilities Scaling Laws for Reward Model Overoptimization

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-19T17:22:41.958533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T17:20:39.969447Z digest=sha256:15a2cd4df299e5b30eb97ddcf5dc93b3891216f93b994a9535062715ac627ac7

Observation 8e0d6ab5-b1f0-4275-b9e3-bc374b552458 · outbound

This paper cites LLMs Gaming Verifiers: RLVR can Lead to Reward Hacking.

Likelihood scoring for continuations of mathematical text: a self-supervised benchmark with tests for shortcut vulnerabilities LLMs Gaming Verifiers: RLVR can Lead to Reward Hacking

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-19T17:22:41.950802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T17:20:39.969447Z digest=sha256:0d94e4dcae47c2b51b945df802f938429558a00df2c8106ef892d866d12cabed

Observation b0608133-b245-4ab7-b1bb-08a425e706b2 · outbound

This paper cites Learning to Reason for Long-Form Story Generation.

Likelihood scoring for continuations of mathematical text: a self-supervised benchmark with tests for shortcut vulnerabilities Learning to Reason for Long-Form Story Generation

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T17:22:42.175048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T17:20:39.969447Z digest=sha256:1c84e2902aaacb7baf521c6e7ffb170cb202ca72bc377afd562412b10953d7bf

Observation 3c6377b9-8581-42b0-8f6a-2e27cdac3496 · outbound

This paper cites BOW: Training Language Models to Reason Over Plausible Next Words.

Likelihood scoring for continuations of mathematical text: a self-supervised benchmark with tests for shortcut vulnerabilities BOW: Training Language Models to Reason Over Plausible Next Words

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-08-05T02:57:26.390855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T17:20:39.969447Z digest=sha256:725fac07754f9cc4e0e1ac9eca65f98c05c506702ebf402aa6552d5818d91f3f

Observation 58624433-75e6-4cca-b616-e9c57a00ed2f · outbound

This paper cites Goodman.Learning to Simulate Human Dialogue.

Likelihood scoring for continuations of mathematical text: a self-supervised benchmark with tests for shortcut vulnerabilities Goodman.Learning to Simulate Human Dialogue

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T17:22:42.170264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T17:20:39.969447Z digest=sha256:b6421e8bb34456e205e72b25c5ffbf40931100b4733dc9dbbb0a466a75d6ee3d

Observation cc87e6fa-65c6-4ae4-917b-ff593ee951e2 · outbound

This paper cites Learning to simulate human dialogue.

Likelihood scoring for continuations of mathematical text: a self-supervised benchmark with tests for shortcut vulnerabilities Learning to simulate human dialogue

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-19T17:22:41.974017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T17:20:39.969447Z digest=sha256:a6f4f0e12831ce6dcc91826196b868135ceea79b0b892a5d8c8e1dcfee092942

Observation faf3e3ba-c466-40e4-940d-44333b642945 · outbound

This paper cites Quiet-STaR: Language Models Can Teach Themselves to Think Before Speaking.

Likelihood scoring for continuations of mathematical text: a self-supervised benchmark with tests for shortcut vulnerabilities Quiet-STaR: Language Models Can Teach Themselves to Think Before Speaking

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-19T17:22:41.967822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T17:20:39.969447Z digest=sha256:806ce981aa73890ee929fe3f649f97746ecd14b1dbcb74a6f1c707f437b83adc

Observation 461999dc-ad6a-4e77-a4ae-6c712e72e1bb · outbound

This paper cites Reinforcement Pre-Training.

Likelihood scoring for continuations of mathematical text: a self-supervised benchmark with tests for shortcut vulnerabilities Reinforcement Pre-Training

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-19T17:22:41.994710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T17:20:39.969447Z digest=sha256:1f8dd44b9f22a4c7701a57f16068326229937715a40dc0f5d712112110408c0b

Observation 14f06e40-d62f-4722-9953-a81ef7d994ad · outbound

This paper cites Reinforcement learning on pre-training data.arXiv preprint arXiv:2509.19249.

Likelihood scoring for continuations of mathematical text: a self-supervised benchmark with tests for shortcut vulnerabilities Reinforcement learning on pre-training data.arXiv preprint arXiv:2509.19249

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-19T17:22:41.980815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T17:20:39.969447Z digest=sha256:018137be6bad6f7b4e2e9605c859bd479979485f44e174f9882c1d32758b8381

Observation 52d0e59a-c204-41fb-87cb-89abf8f4502d · outbound

This paper cites RLP: Reinforcement as a Pretraining Objective.

Likelihood scoring for continuations of mathematical text: a self-supervised benchmark with tests for shortcut vulnerabilities RLP: Reinforcement as a Pretraining Objective

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-19T17:22:41.947707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T17:20:39.969447Z digest=sha256:a6870ac91ca6cd2bdccbef09cbc9c8430eca3da62fdace6067b8c419678f3141

Observation 4b577c64-f3c9-487b-a38f-4266168dba15 · outbound

This paper cites Benchmarking LLMs' Judgments with No Gold Standard.

Likelihood scoring for continuations of mathematical text: a self-supervised benchmark with tests for shortcut vulnerabilities Benchmarking LLMs' Judgments with No Gold Standard

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-19T17:22:41.932179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T17:20:39.969447Z digest=sha256:5593cf57e6c823a08e86ecbc34c39488da0d8067d1b8172c390edd9a9b3e9038

Observation ed3cf325-6311-46df-86db-1fe326369ea3 · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

Likelihood scoring for continuations of mathematical text: a self-supervised benchmark with tests for shortcut vulnerabilities Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-19T17:22:41.907325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T17:20:39.969447Z digest=sha256:8f2eb4d037dca6ad664d9c4378c0d445346989abe8fec04f3fe487fd8a0d6a41

Observation a2461b85-a270-4cc7-8fe6-478af047a7a1 · outbound

This paper cites LLMs as Narcissistic Evaluators: When Ego Inflates Evaluation Scores.

Likelihood scoring for continuations of mathematical text: a self-supervised benchmark with tests for shortcut vulnerabilities LLMs as Narcissistic Evaluators: When Ego Inflates Evaluation Scores

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-19T17:22:41.940048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T17:20:39.969447Z digest=sha256:d7edb120af770bc9177d01e3c3759be709e7523aee0070289e76ebb99b639138

Observation c8c56dd2-ae08-41c1-a880-a984a517a4e2 · outbound

This paper cites LLM Evaluators Recognize and Favor Their Own Generations.

Likelihood scoring for continuations of mathematical text: a self-supervised benchmark with tests for shortcut vulnerabilities LLM Evaluators Recognize and Favor Their Own Generations

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-22T18:44:28.909456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T17:20:39.969447Z digest=sha256:56a751b62064e89fac845601d9c7a974da2c0a2c55effba46af04fe1970e566f

Observation 8d4139fe-d65f-4338-a535-7e6fd2ff8b87 · outbound

This paper cites One Token to Fool LLM-as-a-Judge.

Likelihood scoring for continuations of mathematical text: a self-supervised benchmark with tests for shortcut vulnerabilities One Token to Fool LLM-as-a-Judge

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-06-12T02:08:19.458599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T17:20:39.969447Z digest=sha256:ec039c622cb2007d5350d3343b353fec9cc372359d733c0dc6d66c1a22a154ce

Observation e606b530-0887-4636-bc31-b5cbd5bd3382 · outbound

This paper cites Qwen3 Technical Report.

Likelihood scoring for continuations of mathematical text: a self-supervised benchmark with tests for shortcut vulnerabilities Qwen3 Technical Report

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-19T17:22:41.920948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T17:20:39.969447Z digest=sha256:590155279aed809f64f60bc7ba7663a383aa08107e9d67aa658df909964fe694

Observation fa42aa84-eff7-4c62-9ce6-2c2702b6391c · outbound

This paper cites an unresolved cited work.

Likelihood scoring for continuations of mathematical text: a self-supervised benchmark with tests for shortcut vulnerabilities Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-05-19T17:22:42.167950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T17:20:39.969447Z digest=sha256:f69c84ab662a9fe18ae77044a890320d515bd40ff0841e86bcb02bb786ce5ceb

Observation b02aba94-8097-4662-8b40-85ac683a0cee · outbound

This paper cites Kimi K2: Open Agentic Intelligence.

Likelihood scoring for continuations of mathematical text: a self-supervised benchmark with tests for shortcut vulnerabilities Kimi K2: Open Agentic Intelligence

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-19T17:22:41.914260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T17:20:39.969447Z digest=sha256:e3a128d39c9b29259369b159cb0a8daf91725d36265db01056ca64fdfd9268b1

Observation 1e19df9f-2a4b-4d80-b51f-ee26f6c7c996 · outbound

This paper cites https://deploymentsafety.openai.com/gpt-5-5/gpt-5- 5.pdf.

Likelihood scoring for continuations of mathematical text: a self-supervised benchmark with tests for shortcut vulnerabilities https://deploymentsafety.openai.com/gpt-5-5/gpt-5- 5.pdf

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T17:22:42.164288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T17:20:39.969447Z digest=sha256:fe2e8468280b8b82783c7211fa3e5955395ad19ecc188814b68e0f289160489a

Observation 610333d7-3b26-44d5-a96c-a44f2ea0b0c5 · outbound

This paper cites https://anthropic.com/claude- opus- 4- 7- system-card.

Likelihood scoring for continuations of mathematical text: a self-supervised benchmark with tests for shortcut vulnerabilities https://anthropic.com/claude- opus- 4- 7- system-card

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T17:22:42.161863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T17:20:39.969447Z digest=sha256:57ff724ae92714fc59bf356832cbb0547a1fc625f10b6f69f03e41ceb31b3a54

Observation 6d19877f-fbb5-4d58-a92f-ba3ed10dc329 · outbound

This paper cites an unresolved cited work.

Likelihood scoring for continuations of mathematical text: a self-supervised benchmark with tests for shortcut vulnerabilities Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-05-19T17:22:42.166164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T17:20:39.969447Z digest=sha256:f96aced7080267c1fbd3271b5ac6358782a96014fd025801f1c8f18f7de646f7

Observation 44500651-9eda-4dcb-9ed6-403dcb4beb60 · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Likelihood scoring for continuations of mathematical text: a self-supervised benchmark with tests for shortcut vulnerabilities Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-19T17:22:41.928675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T17:20:39.969447Z digest=sha256:e17016876f18ae89f32ef855450b1498f94923bc311c3d5b8f9d94d9368e9e11

Observation 965dae13-ce52-4aa5-8420-e16417529331 · outbound

This paper cites Wichmann.

Likelihood scoring for continuations of mathematical text: a self-supervised benchmark with tests for shortcut vulnerabilities Wichmann

Reference 29

Resolution
metadata mismatch
doi, observed 2026-05-19T17:22:41.683792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T17:20:39.969447Z digest=sha256:1aef337df33b020d14c584951305e50fb2fdce9cfef33659f9db43f32545ad93

Observation ab477633-0f8b-42a5-bf17-d76af8ba216b · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Likelihood scoring for continuations of mathematical text: a self-supervised benchmark with tests for shortcut vulnerabilities LoRA: Low-Rank Adaptation of Large Language Models

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-19T17:22:41.991528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T17:20:39.969447Z digest=sha256:308c2a57b49376b2da8a893f566b216ada6c5203249e8d491ecc97a8092c90f9

Observation 663c3c43-ec16-4717-b034-328e589a3bb1 · outbound

This paper cites The LAMBADA dataset: Word prediction requiring a broad discourse context.

Likelihood scoring for continuations of mathematical text: a self-supervised benchmark with tests for shortcut vulnerabilities The LAMBADA dataset: Word prediction requiring a broad discourse context

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-19T17:22:41.688057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:42e49285610bbe48b51f5ace6e7c628bbcba155d44e9ae40cfa76c6a2d160278

Observation 8936e89d-57d2-40b9-b95e-fb3cfb7b122f · outbound

This paper cites HellaSwag: Can a Machine Really Finish Your Sentence?.

Likelihood scoring for continuations of mathematical text: a self-supervised benchmark with tests for shortcut vulnerabilities HellaSwag: Can a Machine Really Finish Your Sentence?

Reference 32

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T17:22:41.680582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:54ee3cb172b4afd2706b8354df4cdc01b7096a9147000de4909233315eca582b

Observation 387ffd55-39a0-462d-9f19-0f6b0f8e56ed · outbound

This paper cites Eliciting Expertise without Verification.

Likelihood scoring for continuations of mathematical text: a self-supervised benchmark with tests for shortcut vulnerabilities Eliciting Expertise without Verification

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-19T17:22:41.660724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T17:20:39.969447Z digest=sha256:a25ab4fca38c16bbd38575f265862b7a7898346cbce6a34549d773688d5934dc

Observation db21fb2e-80e3-4708-94cd-29ed59ea279b · outbound

This paper cites Eliciting Informative Text Evaluations with Large Language Models.

Likelihood scoring for continuations of mathematical text: a self-supervised benchmark with tests for shortcut vulnerabilities Eliciting Informative Text Evaluations with Large Language Models

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-19T17:22:41.654523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T17:20:39.969447Z digest=sha256:771608ffb1fc022ff46c949d844bdc90abc3c33eef692702424da4f75b843d26

Observation 10970b81-9a0b-4435-a71a-ef1def50328c · outbound

This paper cites From raw corpora to domain benchmarks: Automated evaluation of LLM domain expertise.

Likelihood scoring for continuations of mathematical text: a self-supervised benchmark with tests for shortcut vulnerabilities From raw corpora to domain benchmarks: Automated evaluation of LLM domain expertise

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-19T17:22:41.918111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T17:20:39.969447Z digest=sha256:63b31666c1a2ca39ede26bdb311f140c048f3620eeccb8b1a7edab9f3b05f0bd

Observation 4f3fdcbb-411b-4dd3-b2fa-138136a6315d · outbound

This paper cites Beyond Verifiable Rewards: Scaling Reinforcement Learning for Language Models to Unverifiable Data.

Likelihood scoring for continuations of mathematical text: a self-supervised benchmark with tests for shortcut vulnerabilities Beyond Verifiable Rewards: Scaling Reinforcement Learning for Language Models to Unverifiable Data

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-19T17:22:41.904395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T17:20:39.969447Z digest=sha256:8221c9eff1d94a5d82b439a41401b0026e07ee599998ab65a9cfef9f0a594cc9

Observation c380f5f3-a7fd-4121-a9d9-989a848e34bc · outbound

This paper cites Training Chain-of-Thought via Latent-Variable Inference.

Likelihood scoring for continuations of mathematical text: a self-supervised benchmark with tests for shortcut vulnerabilities Training Chain-of-Thought via Latent-Variable Inference

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-19T17:22:42.011138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T17:20:39.969447Z digest=sha256:5ef9dea3bf7470607a5b6c2d0bf537532f1c7931f131b0374ececba8ecddbe58

Observation da76e2c7-a966-4617-992f-297c9dbfd6de · outbound

This paper cites Amortizing intractable inference in large language models.

Likelihood scoring for continuations of mathematical text: a self-supervised benchmark with tests for shortcut vulnerabilities Amortizing intractable inference in large language models

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T17:22:41.987390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T17:20:39.969447Z digest=sha256:69b34769d30cae74aa076c0c565a3a02b6e22a2e3743ae4af0ae3f71a36d0baa

Observation 90816fac-5eb9-4ded-8ebd-bd67a07304c1 · outbound

This paper cites NOVER: Incentive Training for Language Models via Verifier-Free Rein- forcement Learning.

Likelihood scoring for continuations of mathematical text: a self-supervised benchmark with tests for shortcut vulnerabilities NOVER: Incentive Training for Language Models via Verifier-Free Rein- forcement Learning

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T17:22:42.159617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T17:20:39.969447Z digest=sha256:e534e547ef7050203db95932336ad9b99ec2bebcde6482cbce7e4b397ac9d873

Observation 3984fa88-075b-41d5-be28-a20913e35529 · outbound

This paper cites Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks.

Likelihood scoring for continuations of mathematical text: a self-supervised benchmark with tests for shortcut vulnerabilities Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-19T17:22:41.983841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T17:20:39.969447Z digest=sha256:3bc2c9f8c77726a54050b1802304e3911f7ae9bd5442f820b3b9038464721719

Observation 8727f605-db0e-467b-8e17-1e63a94ffd2b · outbound

This paper cites Reinforcing General Reasoning without Verifiers.

Likelihood scoring for continuations of mathematical text: a self-supervised benchmark with tests for shortcut vulnerabilities Reinforcing General Reasoning without Verifiers

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-19T17:22:41.894869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T17:20:39.969447Z digest=sha256:b53459f57c08cb46c96d0408fd15b2c42da6e1af62e3569da103827f75f6fba9

Observation 72cb19f3-c96d-4734-b47f-68a9b6dacd23 · outbound

This paper cites RLPR: Extrapolating RLVR to General Domains without Verifiers.

Likelihood scoring for continuations of mathematical text: a self-supervised benchmark with tests for shortcut vulnerabilities RLPR: Extrapolating RLVR to General Domains without Verifiers

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-19T17:22:41.998128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T17:20:39.969447Z digest=sha256:b0e9a4ca7c6ff7e393589ddde488e91ea4d69a97027752a508c2472ec5447e07

Observation aa6dc1e9-1f3f-4ec2-8fcc-0db2082d1529 · outbound

This paper cites Likelihood- based reward designs for general llm reasoning.

Likelihood scoring for continuations of mathematical text: a self-supervised benchmark with tests for shortcut vulnerabilities Likelihood- based reward designs for general llm reasoning

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-19T17:22:42.014458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T17:20:39.969447Z digest=sha256:203181b584d9744a67f6247c887b8b7269db83dfcad2087477c7c946ef799a08

Observation 9f47c5dd-87c3-4803-a923-2bb0cd99653c · outbound

This paper cites Let's Verify Step by Step.

Likelihood scoring for continuations of mathematical text: a self-supervised benchmark with tests for shortcut vulnerabilities Let's Verify Step by Step

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-19T17:22:41.970558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T17:20:39.969447Z digest=sha256:03f478ece45a9831d708f771fb2e172e75741cec0459fb9e447b0c4ac4dd3d42

Observation 935fe3b7-6a3a-4a89-a513-58cd0581c024 · outbound

This paper cites Generative Verifiers: Reward Modeling as Next-Token Prediction.

Likelihood scoring for continuations of mathematical text: a self-supervised benchmark with tests for shortcut vulnerabilities Generative Verifiers: Reward Modeling as Next-Token Prediction

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-19T17:22:41.977768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T17:20:39.969447Z digest=sha256:531edf802765a9dfc01fb167ec69b1f3e5eb443abf1732d4e0b8797114c84177

Observation 6cbfbf02-1ba1-421c-8ad5-e79267653787 · outbound

This paper cites Prover-Verifier Games improve legibility of LLM outputs.

Likelihood scoring for continuations of mathematical text: a self-supervised benchmark with tests for shortcut vulnerabilities Prover-Verifier Games improve legibility of LLM outputs

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-19T17:22:41.910605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T17:20:39.969447Z digest=sha256:41005ac340e7c8a94380fbaf449204dbefd62857a2ec4f6a58888ba56fbbd78e

Observation 6970f9e5-0ad5-4862-aae5-490007e2211c · outbound

This paper cites Variation in Verification: Understanding Verification Dynamics in Large Language Models.

Likelihood scoring for continuations of mathematical text: a self-supervised benchmark with tests for shortcut vulnerabilities Variation in Verification: Understanding Verification Dynamics in Large Language Models

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-05-19T17:22:41.963957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T17:20:39.969447Z digest=sha256:095fdac52b49ba5389ce3219937f4942c634414f56b100490aa6096d5f39cb83

Observation 437854ed-ec01-456b-a1ee-d7a6194df3c7 · outbound

This paper cites Reward under attack: Analyzing the robustness and hackability of process reward models.

Likelihood scoring for continuations of mathematical text: a self-supervised benchmark with tests for shortcut vulnerabilities Reward under attack: Analyzing the robustness and hackability of process reward models

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-19T17:22:41.897933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T17:20:39.969447Z digest=sha256:5d71809075cdadfb8e658d491a8bac9410d4c4f63e38d4cc1444a08caf9c50cc

Observation c85fdfea-fd70-4bc8-876d-ca122f1d053c · outbound

This paper cites Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms.

Likelihood scoring for continuations of mathematical text: a self-supervised benchmark with tests for shortcut vulnerabilities Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms

Reference 49

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T17:22:41.901099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T17:20:39.969447Z digest=sha256:4ca93064566b20beeff37b09dfa8038de767a7e7a7ac605823e5f50af7be891f

Observation 5ab4af3b-05cb-4a8c-8778-5f03d9347430 · outbound

This paper cites Inference-time reward hacking in large language models.arXiv preprint arXiv:2506.19248.

Likelihood scoring for continuations of mathematical text: a self-supervised benchmark with tests for shortcut vulnerabilities Inference-time reward hacking in large language models.arXiv preprint arXiv:2506.19248

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-19T17:22:41.924264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T17:20:39.969447Z digest=sha256:40358d8fa18d70bc4e88158ffaecd2aa7b59d21c28b4c964b0e3d2e58e9c06f3

Pith citing papers

No inbound Pith citation observations are available.