Pith. sign in

Paper Citation Record · LEDGER

CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning

As of 15 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 0 inbound Pith citation observations for arXiv:2608.09324.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.09324 v1

Coverage vector

measured 35 of 35 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T19:31:40.971330Z

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

35 of 35 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved32
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f30a8904-d07d-4dbf-95f9-3d25598e0620 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning Training Verifiers to Solve Math Word Problems

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:40.801100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:31:40.801100Z digest=sha256:a9acb7ac9a03fcb5f5c77dae343f524d733c106256a80a4075512340b6564b91

Observation 2154c012-1a09-493b-9037-3a2dcf974eea · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:40.819331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:31:40.819331Z digest=sha256:3de894ff71c9a276ca7624bee62e28df021d8a6f9e2d2c4c4c724bb82880b260

Observation b4c865d0-7b3f-4aab-9ca8-d369bf9b13a4 · outbound

This paper cites OpenAI o1 System Card.

CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning OpenAI o1 System Card

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:40.828258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:31:40.828258Z digest=sha256:1e22e166fdf8fd457846adef354c343d7bc899118080126f9e6cccc9ec565262

Observation 3eb025ee-b9be-4d07-b61a-2a2976290f80 · outbound

This paper cites Mistral 7B.

CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning Mistral 7B

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:40.832494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:31:40.832494Z digest=sha256:3e1e47e06476ee0b29483a82ea663474a727d6ad8340ca56bd59ddf1a8cd6035

Observation 67e570e3-ef33-40e8-8ab1-57cd9d5de482 · outbound

This paper cites Language Models (Mostly) Know What They Know.

CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning Language Models (Mostly) Know What They Know

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:40.837316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:31:40.837316Z digest=sha256:07fd346d89a96ccc8172b1b02a290a9258b942c136e3982565e6b77cb3ed69c5

Observation 3763d5af-c35b-40e8-8efa-1acead1f0eaa · outbound

This paper cites Let’s verify step by step.

CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning Let’s verify step by step

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:31:41.893764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T19:31:40.852283Z digest=sha256:1d3e5aada86cea9c47a7044b1e3adde0b29d1a5f747ce38b3d12982f838fb5ca

Observation 0ca7ff60-b193-43ea-80f3-3c4a1b3e7ec9 · outbound

This paper cites Teaching Models to Express Their Uncertainty in Words.

CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning Teaching Models to Express Their Uncertainty in Words

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:40.858193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:31:40.858193Z digest=sha256:4b2803b2161de3b078b15791475bee24fdfb71c17b04c978e47b1e2891ba5a80

Observation ecf3ec00-7a5e-45db-8cae-bce95f165443 · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning Understanding R1-Zero-Like Training: A Critical Perspective

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:40.864651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:31:40.864651Z digest=sha256:60be20273cc7cdada7609a3f04a7fc52918393472df98f532e7fb072fddba7b8

Observation 9918b9b5-9c6c-477e-8735-df3179ae6635 · outbound

This paper cites The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models.

CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:40.875842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:31:40.875842Z digest=sha256:18ee80c93b42f67909bf4797b253e0501a44c7377ff30f8066fddf3d8b617421

Observation c8f32949-0c2d-416b-b2ca-d565972e4262 · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:40.885338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:31:40.885338Z digest=sha256:58efabf8dd7a345d6fc6226d22a3d0c6121d3c384ef47c88e362ab9cf9579a9b

Observation 972f5e9b-f015-43b8-ad1d-a35d8b32f25f · outbound

This paper cites Spurious Rewards: Rethinking Training Signals in RLVR.

CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning Spurious Rewards: Rethinking Training Signals in RLVR

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:40.900230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:31:40.900230Z digest=sha256:122aea4210b86517767c00631d82feed22a1020efebdc9767f94726b7d8b5d38

Observation 4a0a4733-879f-44ea-9178-5e4fdd7908e8 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:40.904736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:31:40.904736Z digest=sha256:6b188c44ed6edfff67636af67b552057a0bf32da5717da0c86bd8543395748ac

Observation 106e879b-f55d-4cd7-96c1-d27ab1b02433 · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:40.909954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:31:40.909954Z digest=sha256:afb230c015b802e6f80146569336f76fa97d449fdb7ba881ef2831ead2905dc2

Observation ed71015f-8220-4ee6-bd25-0d48f95b0b1d · outbound

This paper cites Solving math word problems with process- and outcome-based feedback.

CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning Solving math word problems with process- and outcome-based feedback

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:40.922992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:31:40.922992Z digest=sha256:e64411ad05db3ab99969118f77689a2d14bbf5ce18377d76be2fd3e11d70d23f

Observation 6f47bff3-84ae-434f-b9e6-1be1c7a93777 · outbound

This paper cites Tent: Fully Test-time Adaptation by Entropy Minimization.

CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning Tent: Fully Test-time Adaptation by Entropy Minimization

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:40.928529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:31:40.928529Z digest=sha256:c20abd9d9f033863d9b5da5b6f736d3238370d82bb9b84ffa0bd2112ee20b5de

Observation da00f497-998e-4b43-ac65-8843f7d92cf8 · outbound

This paper cites Self-training with Noisy Student improves ImageNet classification.

CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning Self-training with Noisy Student improves ImageNet classification

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:40.939920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:31:40.939920Z digest=sha256:12ea479677cf724fecd778dcb8c821e42667fd36bda5b5ecf70043dfa30f6b8b

Observation db26dd7a-caae-4063-b2eb-074cbbf4155f · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:40.945421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:31:40.945421Z digest=sha256:e88921ed31faed1c87c3760747ca5d2ba30948523c547fc22e9e74fd694c652f

Observation c3a74c6b-e452-4117-890a-58820d90bbdf · outbound

This paper cites Qwen3 Technical Report.

CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning Qwen3 Technical Report

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:40.949492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:31:40.949492Z digest=sha256:55b939e72c02011250755bd6d3e7e9ed2095ea2d3ecf3d01cb89ad6f553f3ee8

Observation cb06f9fc-655a-4d51-8dd5-4a8d5a17b7ac · outbound

This paper cites Self-Rewarding Language Models.

CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning Self-Rewarding Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:40.953612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:31:40.953612Z digest=sha256:75ade6833b05f52a314aa6192efe3f2a83f82d4f11294d67b701bf7d7c7fdb47

Observation 9294ac53-e399-4e84-abc6-f6ad3434aec7 · outbound

This paper cites Learning to Reason without External Rewards.

CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning Learning to Reason without External Rewards

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:40.961404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:31:40.961404Z digest=sha256:3c7ce5b4e844bf15012d8ac7c569a4a5badbc86123431efd339109f48e0490e5

Observation 5ababe61-88b4-4790-87a4-14418dfe98ec · outbound

This paper cites B Use of Large Language Models Large language models were used only to improve grammar and clarity in author-written text.

CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning B Use of Large Language Models Large language models were used only to improve grammar and clarity in author-written text

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:31:41.871053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T19:31:40.965726Z digest=sha256:7b18e33bf7935ba09357ab614924a8e6e1c85bb647c8b8516c9cefa867ff641e

Observation 3b0560b6-dc2c-4861-89c7-b1767232d739 · outbound

This paper cites role": "user.

CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning role": "user

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:31:41.848806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T19:31:40.971330Z digest=sha256:e1c8545cd7ff0506e5ce41bce68e6fa9120d95a8d6146db44ca71eee8b54b037

Observation b1e51fd3-abc4-48c4-bc3e-34771d4744c9 · outbound

This paper cites Training language models to follow instructions with human feedback.

CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning Training language models to follow instructions with human feedback

Reference 1965

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:40.870949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:31:40.870949Z digest=sha256:e1dbdd57b3ce348e727165dd7074a48a539ea2ba1139bc65c9a925b6c5df9bbb

Observation 439abf01-d528-4a3e-b250-662f26ec332d · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 1978

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:40.915743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:31:40.915743Z digest=sha256:c2888e829adab2dbc5a93910c7db689bbc891b413b4beb30c42a35b0510ad5ff

Observation b567eace-03b0-4e05-a22e-705ae91c44b9 · outbound

This paper cites Proximal Policy Optimization Algorithms.

CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning Proximal Policy Optimization Algorithms

Reference 1988

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:40.890115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:31:40.890115Z digest=sha256:4c843c6adfd9aef4f6662431ab974fa6d82ae4dbd7942cf0e2d40c97f0d3d37e

Observation d5224e1c-488d-4bd6-a3a4-87d6c2f7adf2 · outbound

This paper cites Large Language Monkeys: Scaling Inference Compute with Repeated Sampling.

CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning Large Language Monkeys: Scaling Inference Compute with Repeated Sampling

Reference 1997

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:40.794957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:31:40.794957Z digest=sha256:4cfddd8185fd3589e4fe605e98a5489ad88cf186d0a25cdfaea73c712ee4d871

Observation 8155807b-da34-401a-897e-8159a418b84f · outbound

This paper cites Maximizing Confidence Alone Improves Reasoning.

CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning Maximizing Confidence Alone Improves Reasoning

Reference 2007

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:40.880140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:31:40.880140Z digest=sha256:fca14ea4032e7295ad55aa4b302d420b097bc0e607089cdf41599f863e2b024b

Observation 2b76d1b3-2956-4e92-b2d1-d42c60fc789a · outbound

This paper cites Can large reasoning models self-train?arXiv preprint arXiv:2505.21444,.

CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning Can large reasoning models self-train?arXiv preprint arXiv:2505.21444,

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:40.895438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:31:40.895438Z digest=sha256:1cc03f75913194be0e79e5a3d8b2efe21ea48fd16e3f0eb7bda32dcc0a6a94a0

Observation a8dc9f2c-f800-4540-8084-e3f16997dd95 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:40.934963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:31:40.934963Z digest=sha256:3f3561a5804f33bbd19125447bfc95a4653096ffedeee34142ff7b325b397976

Observation ceb19587-0ff7-4c21-a5df-33c004284db0 · outbound

This paper cites The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models.

CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:40.806378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:31:40.806378Z digest=sha256:bce25348b17cfe0f7f754c07047b512f1fff220483dcc887a951dcfa62a39691

Observation 21b7eeaa-3378-4cf0-9f30-494b318e321e · outbound

This paper cites ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models.

CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:40.843022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:31:40.843022Z digest=sha256:16c77f75e5e933ea96b68569700e0c682df2851c6bc10883668db5a52fa26459

Observation 2c133695-b7c3-40d1-b1fc-7bb7dea348f8 · outbound

This paper cites The Llama 3 Herd of Models.

CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning The Llama 3 Herd of Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:40.814245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:31:40.814245Z digest=sha256:63d3cae7777e1fe7fc9d2c42d1377922bcd38b2938aa36dd72200a9492bd9946

Observation f64ae17d-0407-4c4c-917c-e2cc6e5b07d0 · outbound

This paper cites Concrete Problems in AI Safety.

CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning Concrete Problems in AI Safety

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:40.788731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:31:40.788731Z digest=sha256:22838b6d1e714c207170517f35015701fad342d1bf7c7ebe363cc5b596bfc0fc

Observation 2169df09-6c62-4dc3-81d5-3b3baee10156 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning Measuring Mathematical Problem Solving With the MATH Dataset

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:40.824204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:31:40.824204Z digest=sha256:2a7846b7dc0f9317c90dc3544b2da463aa0f4ddbf8db2fbb3db6136fd7132c1b

Observation b466cfc6-3bbf-4847-b6d6-ea2c260e53cd · outbound

This paper cites Co-rewarding: Stable self-supervised rl for eliciting reasoning in large language models.

CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning Co-rewarding: Stable self-supervised rl for eliciting reasoning in large language models

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:40.957387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:31:40.957387Z digest=sha256:010eab89c8dd09bb592e18e1f693092f7649104679d0933189cd3eacb5fb240f

Pith citing papers

No inbound Pith citation observations are available.