Pith. sign in

Paper Citation Record · LEDGER

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text

As of 2 August 2026, this Paper Citation Record lists 78 of 78 outbound references and 1 inbound Pith citation observation for arXiv:2604.20051.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.20051 v2

Coverage vector

measured 78 of 78 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T01:51:00.913166Z

measured 79 of 79 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-02T06:30:47.504484+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-30T18:33:28.486944Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

78 of 78 outbound references displayed

  • verified exact23
  • verified fuzzy41
  • unresolved10
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c16066c7-4b69-46b8-80e6-d3f021fd5220 · outbound

This paper cites HealthBench: Evaluating Large Language Models Towards Improved Human Health.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text HealthBench: Evaluating Large Language Models Towards Improved Human Health

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:18:21.444858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:49439f2216b34eed6dfc93bfab66ff3b15e351c3a1679e0252393df39d5d6493

Observation b8b22f47-daa5-455e-80bb-bc43f4b79248 · outbound

This paper cites SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:21:18.201366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:a79963efc68bae6228d191dbd12baf4ef8debfb45cc1f9c6ee1200fe16a8c553

Observation 9b305eee-49a1-4e18-b7ad-5f8008fd23ad · outbound

This paper cites Self-playing Adversarial Language Game Enhances LLM Reasoning.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Self-playing Adversarial Language Game Enhances LLM Reasoning

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:21:16.065828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:c8a3fdc0e7377ad60b1fdcfd1702e1713752f964c12a1dab94f106dc21af0d39

Observation cad5ea11-a550-4ee0-88cf-28b1382f298d · outbound

This paper cites Scaling Instruction-Finetuned Language Models.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Scaling Instruction-Finetuned Language Models

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T01:23:37.463337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:ba62d53268a2795aae15cfb4ed42ad46f19619bc5e3e235b10add586b2d0b952

Observation 5ed87529-ca00-4062-963b-8737ce276527 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Training Verifiers to Solve Math Word Problems

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-11T13:21:16.791554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:7ffd8b8ca6e36e6383b5a77cb5e5353b62053f6e94e99eb445b34d7dcfef5837

Observation ee773845-5881-47a3-aa5b-3c7013d6cb71 · outbound

This paper cites DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-11T13:21:16.865950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:e8c8aa547cc8b3c46f0216fa59782d5981578247d3f3de4495d2f429b1ef4bca

Observation 1b4f9982-637e-4fc4-b90c-afdd0c0ea026 · outbound

This paper cites Qa-lign: Aligning llms through constitutionally decomposed qa.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Qa-lign: Aligning llms through constitutionally decomposed qa

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:52:24.332412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:98df4ca51c6f35de658050f5a8d80c9f1420194c74a9eed3a4b67c684b116ab4

Observation 13ee7d65-ed37-46ce-85f8-ba1621619c42 · outbound

This paper cites Openwebtext corpus.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Openwebtext corpus

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:52:24.335087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:bf3b9dafb29bb100502734ddd8742fc1c543130b333f6d5f176d147d3b0fc329

Observation 4b045794-9b96-4880-9d6f-2330be434a01 · outbound

This paper cites Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:07:56.884714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:a1f461077c50c3c87dfba4ab0227cf0c8f17b8b09e0ef87dd9e45c48b9a1bf4b

Observation a50147df-d2eb-4762-bf22-a40cc296e96f · outbound

This paper cites Lighteval: A lightweight framework for llm evaluation.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Lighteval: A lightweight framework for llm evaluation

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:52:24.356439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:4cd80e404303985160cbf81aa9028d068f8171548885a794fa8a0c5420afaf4f

Observation f6e52eaa-d0ce-4952-b903-f137865b1e89 · outbound

This paper cites arXiv preprint arXiv:2511.10507 , year=.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text arXiv preprint arXiv:2511.10507 , year=

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:21:17.432193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:43d16d7ccdf19093c6ed6137a3e4ecd3644b050e5054a0913d7ad4d4b8cddcae

Observation ead073ee-97ef-44fa-9525-ef81d57423d6 · outbound

This paper cites Measuring mathematical problem solving with the math dataset.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Measuring mathematical problem solving with the math dataset

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:52:24.320248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:022f8af41fdd37f02f9ab9bb90bb1a12c38eff98c81123bfdc76524fa2f1eab7

Observation 16ef9c3e-12c6-4542-bc13-dc746c5cc050 · outbound

This paper cites R-Zero: Self-Evolving Reasoning LLM from Zero Data.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text R-Zero: Self-Evolving Reasoning LLM from Zero Data

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:23:29.765778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:db5817a58ebfda4ae683e9e75e0a6bdc8659fb721e017b89c3a90bf94f2aa9be

Observation 35977edc-4327-404e-9af5-ab41d386ea9c · outbound

This paper cites Dcrm: A heuristic to measure response pair quality in preference optimization.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Dcrm: A heuristic to measure response pair quality in preference optimization

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:52:24.314439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:ba17470e0c5ba1b722b94f66308475e27961b347b2b17175401af2fcc07d9b3d

Observation e6e2648b-8406-4f60-a171-efc2ff33a10e · outbound

This paper cites Large language models can self-improve.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Large language models can self-improve

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:52:24.317630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:193d47b94eebb4a3df2c3f89dd116dbbafb7908d4d4280cb44c3f8e6789cfe27

Observation df2ee1d2-1693-4e1f-85b2-8b40ec090896 · outbound

This paper cites Reinforcement Learning with Rubric Anchors.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Reinforcement Learning with Rubric Anchors

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:21:15.766370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:b62f12b653e8bf9456582cf802b2a2f2fc3c8cb6bef409c7dd7a8e176c92ffc1

Observation 46953652-3dbc-4c3c-b4ce-d0cec873f233 · outbound

This paper cites TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:00:28.883814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:c2c667b579f2a2efff3ce8c22950d9192e0f1cc9915441f04000aa0b9cac896a

Observation a918c201-3f10-4d45-90eb-574ba0ab1a87 · outbound

This paper cites Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:52:24.322811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:8a449ffa8b6a0422847a16037240643fc7f6588c8ad7e5dbb37084a2416c296c

Observation 13561926-7329-4bec-a0c7-7d0a9ae73aef · outbound

This paper cites Truthfulqa: Measuring how models mimic human falsehoods.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Truthfulqa: Measuring how models mimic human falsehoods

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:52:24.306823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:efa4b7b1f122a3a411037b4d5d9b15e14e6b987e9587f76a1d6d9c1143f93984

Observation 44782e97-3381-4c8d-bcce-60eb2d38ca20 · outbound

This paper cites Spice: Self-play in corpus environments improves reasoning.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Spice: Self-play in corpus environments improves reasoning

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:21:16.426077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:f404d0741800da1cb21dbb9c48e2f4fdc7199a91b382d5f22bd18db0bc2c151c

Observation 0cebd866-f376-4158-9a07-dd5a5d0450fd · outbound

This paper cites Decoupled weight decay regularization.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Decoupled weight decay regularization

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:52:24.297173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:da5b2b8b4e8413c96af91ec1d49681e1420b1cef527a2e78546283878331390f

Observation 23004da3-0573-4573-80d6-05c6ebcc779f · outbound

This paper cites Advanced version of gemini with deep think officially achieves gold-medal standard at the international mathematical olympiad.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Advanced version of gemini with deep think officially achieves gold-medal standard at the international mathematical olympiad

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:52:24.354128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:b377572f964fb37345b854bc63a918d6ebfbad1e54a188cdf48b0034f5393d01

Observation 6cc02ad7-93fe-46f2-a834-5e82891d9f13 · outbound

This paper cites Building trust in clinical llms: Bias analysis and dataset transparency.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Building trust in clinical llms: Bias analysis and dataset transparency

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:52:24.289794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:a857532d37beaf56e8e1ab1e9b28440f82aeffec8d3c7c3c76b7c04c06d823ee

Observation f58203d7-0612-46ac-beab-f7a969ec5519 · outbound

This paper cites Eq-bench creative writing benchmark v3.https://github.com/EQ-bench/ creative-writing-bench.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Eq-bench creative writing benchmark v3.https://github.com/EQ-bench/ creative-writing-bench

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:52:24.368160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:1eb7b942e066c9454a444d0b3afd6d9d0b52b6633f41f6019e223d2b982b6dd7

Observation 87aef807-ba0a-4b09-9a75-1a2b97409d78 · outbound

This paper cites Medmcqa: A large-scale multi-subject multi-choice dataset for medical domain question answering.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Medmcqa: A large-scale multi-subject multi-choice dataset for medical domain question answering

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:52:24.294895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:5d36db43ccab8ab89cb9b6e41d9a7407cb0d98d37e99fc9b229ce98d3cf0f2b5

Observation ecc7d154-0fe6-434c-b2ec-717878a2d5df · outbound

This paper cites van Duijn, Niki Stein, Mike Preuss, Peter van der Putten, and Kees Joost Batenburg.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text van Duijn, Niki Stein, Mike Preuss, Peter van der Putten, and Kees Joost Batenburg

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T13:21:15.039504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:973363a02c76927d4d7d7896fcc23f36ac8f46845e4b2d83441855b5d6f181c8

Observation a9194573-f9b1-4726-a93e-d20276e60934 · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 28

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T13:21:15.477936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:f07dfce2eec8a03658bed9ce06d532abd15bb67aa1b4a25f2a52428bcad27a06

Observation b906ade4-0446-4b02-979d-2b8156a451d0 · outbound

This paper cites an unresolved cited work.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-05-23T01:52:24.299211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:4c6e9bad3ba8aeaf7b56f39ebf69c4417c04aae482f12ce4cee0066d53078640

Observation 3cf4977f-485d-41f4-aaa5-3b550ffc7f41 · outbound

This paper cites Sutherland.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Sutherland

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:52:24.351536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:d1bd07f68470aa8b30a524afc9156084860214b6d1719d0119f68d7dce7c8748

Observation b446829e-512a-4c08-b734-14fa0f269c19 · outbound

This paper cites Karl: Knowledge agentsvial reinforcement learning.arXiv preprint.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Karl: Knowledge agentsvial reinforcement learning.arXiv preprint

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:21:15.049063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:f2e8e5f92b9c5843fc1898504952353354739c91c606da6302f5e8a5b3b3f475

Observation baea06cd-7924-41b9-9015-7b0e5ff2c439 · outbound

This paper cites Online rubrics elicitation from pairwise comparisons.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Online rubrics elicitation from pairwise comparisons

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:21:17.015919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:0317822e0aed957dad18d1abb8c9da8b219dabcd1bf68ab9825728d0253636a8

Observation fcba67ff-a40e-41c0-a2c7-7c2a4c532907 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Proximal Policy Optimization Algorithms

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-11T13:21:17.128690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:2a5ef1316705573b9db361e7bc3fd4e8d131fa020b58d09e981e35078c19434c

Observation 8f0b2e61-a066-4b38-b789-48f668c93a99 · outbound

This paper cites DR Tulu: Reinforcement Learning with Evolving Rubrics for Deep Research.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text DR Tulu: Reinforcement Learning with Evolving Rubrics for Deep Research

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-20T00:00:28.487879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:28c07f6098618148fcc8cb838d416e7194992ea72d01914417e5a3618cb1c0d0

Observation 0fdb52c6-4825-4dc1-bab5-1f2408888556 · outbound

This paper cites v 1: Unifying generation and self-verification for parallel reasoners.arXiv preprint arXiv:2603.04304, 2026a.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text v 1: Unifying generation and self-verification for parallel reasoners.arXiv preprint arXiv:2603.04304, 2026a

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:21:15.655944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:79788ac7c65c37186d65086d96b4e170f252424d9fdd0aba3d02a1f05474e52b

Observation 882aaab8-7a17-4de6-8f88-86d8b69a5d45 · outbound

This paper cites Book titles and abstracts.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Book titles and abstracts

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:52:24.345305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:097c2ba3ea69a37acab5a8895d5fe8f752917dc6568c64a6450bbd4c6f637db5

Observation 004af272-0f62-4965-88bb-d0c37e4a3a2b · outbound

This paper cites Mind the gap: Examining the self-improvement capabilities of large language models.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Mind the gap: Examining the self-improvement capabilities of large language models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:52:24.278955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:91401dc275c7a9a9ca23e6af7c5defe03336562a1f9f8d47390dbf8b41d35433

Observation a16b9864-e86d-454b-844b-7afae12100a8 · outbound

This paper cites Understanding the performance gap between online and offline alignment algorithms.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Understanding the performance gap between online and offline alignment algorithms

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:21:16.758349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:d87d38bb6a1e5ca2ff6441ec80ffb6a226348672eb286c306edeee7118cce172

Observation d3c19b3e-ed8b-411a-8b86-303893dc19c8 · outbound

This paper cites GPT-4 Technical Report.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text GPT-4 Technical Report

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-11T13:21:17.585101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:5193b78743802a92bf2a41c33b49c5a334c937c2e05deee4942e908f404d0d05

Observation 2d8c22e1-6318-4674-b612-7b551211262f · outbound

This paper cites Qwen2.5 Technical Report.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Qwen2.5 Technical Report

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-11T13:21:17.816751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:e5f4a424b3e54a2c472928f35b13eebd4ebf44eb1883eaaf36f8404404a4511f

Observation 181960ab-b1c2-485c-8966-4aba5b9046ab · outbound

This paper cites Will we run out of data? limits of llm scaling based on human-generated data.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Will we run out of data? limits of llm scaling based on human-generated data

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:52:24.340812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:1f3396055b4b6095804c59cf9af6feb067f83d06d169bf60860f05285ecb9682

Observation eb4e54e2-abdf-414b-b079-7117811a630e · outbound

This paper cites Checklists are better than reward models for aligning language models.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Checklists are better than reward models for aligning language models

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:52:24.338244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:88ab2330ab22e743fe0cc0c7753429673795ef516a67fa808ee986ecc2ddead8

Observation f2bd45f3-8fbb-42a1-912d-b190b80ed32e · outbound

This paper cites Mmlu-pro: A more robust and challenging multi-task language understanding benchmark.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Mmlu-pro: A more robust and challenging multi-task language understanding benchmark

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:52:24.276123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:fa387a0817267db108e3d95fdfdeffabca5787982087b23e5059fe03cb136cc5

Observation c0593d53-32db-40ea-8e65-eccbeecc140f · outbound

This paper cites Self-rewarding language models.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Self-rewarding language models

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:52:24.281802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:ac968fcbfca04715e13c9168d84301748d3a862ed38b4992a3ffd3de8816e9dd

Observation fdfe0c9d-e0e2-414b-9b12-156e970004f7 · outbound

This paper cites Better llm reasoning via dual-play.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Better llm reasoning via dual-play

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:21:17.500054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:02a0b4b093287e31d73b559d0216a64f36af699b6b7cce9575ae1504cb676d5c

Observation 26d16d03-9b4c-430a-b486-f55d7916c973 · outbound

This paper cites Absolute Zero: Reinforced Self-play Reasoning with Zero Data.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Absolute Zero: Reinforced Self-play Reasoning with Zero Data

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-13T18:23:09.397455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:5e21024d145c52ce11474384c98633ac821b90700f8f1f825126931f21b76ea3

Observation d22ca3be-823c-4c42-84d8-f51a07a433b9 · outbound

This paper cites Xing, Hao Zhang, Joseph E.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Xing, Hao Zhang, Joseph E

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:52:24.284299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:78fbb91e853d959bae5432203a3a724a18d15fd4be76ba8faf9aca320c6f1e59

Observation c794d79d-057f-49d1-ad6a-e9b15a81eba8 · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Instruction-Following Evaluation for Large Language Models

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-05-11T13:21:16.695019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:827f37ddf86ecf4685501d1a2dfda334f8375aa1a44d0bb383f5572b157b9e66

Observation c8b42532-f4e1-44ed-84cc-8cc5bb2445ca · outbound

This paper cites Breaking the exploration bottleneck: Rubric-scaffolded reinforcement learning for general llm reasoning.arXiv preprint arXiv:2508.16949.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Breaking the exploration bottleneck: Rubric-scaffolded reinforcement learning for general llm reasoning.arXiv preprint arXiv:2508.16949

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:21:15.325923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:384e073650a2dd3d69e3600d99b5255cdf7b53aa371715089000a5303cfa3599

Observation ec11e069-6aa6-4de0-9353-d988fd8028ef · outbound

This paper cites Self-Challenging Language Model Agents.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Self-Challenging Language Model Agents

Reference 50

Resolution
malformed identifier
arxiv_id, observed 2026-05-11T13:21:15.160379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:58a78f0d58b083f1963cb4dd01d6825146af43eef0ab39e7cb6cf4780bed2cfe

Observation 8e7bbc34-6935-49cc-83fd-12c915292c90 · outbound

This paper cites an unresolved cited work.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-05-23T01:52:24.271410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:92078b61b735fe89154769e730b4333dc241ee8653f62d879d175501ff64486b

Observation 400b7898-0135-480e-9c81-e19e4b02b72c · outbound

This paper cites an unresolved cited work.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-05-23T01:52:24.359068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:740f8e741f4f152f1b01b204d51cbe7f712b526013f442bdc621149099756edc

Observation b089ffea-c2bf-482c-ab3c-5383421a50ec · outbound

This paper cites an unresolved cited work.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-05-23T01:52:24.273495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:650afa91ab51873847c72f350540cb4b3e6ae1de7a5007867a686948b02a4331

Observation 5383c140-1ba4-46a1-a067-85939b894728 · outbound

This paper cites Enzyme Identification.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Enzyme Identification

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:52:24.287023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:df12f80cd994c4437991375c434ddc71b8fd6560b8b83cbf69b525c1e2863627

Observation 924c6ae8-f303-4138-932e-be8ec1d40f06 · outbound

This paper cites an unresolved cited work.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-05-23T01:52:24.304236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:c482feaee5f88686c47927f56052d75697ad857f0f3cf39b190d1785466f3897

Observation f72c0fbb-70ca-4011-b8f8-c017c65e93d4 · outbound

This paper cites Figure 44: Query (Instruction Following).

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Figure 44: Query (Instruction Following)

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:52:24.325189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:83471addd6b2c5c6b3fb953625028e7d932fb6a184eabcc7ed076df5de41dd06

Observation 1feccc5b-73fa-4e8e-950f-4968331f9809 · outbound

This paper cites an unresolved cited work.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-05-23T01:52:24.292393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:c7f7faa6506c37e400e7838260e2b53c4237cb55159355fdad83b787bcf4ccd1

Observation c7a45f16-28bc-4b44-b3bb-50d36e1f2d88 · outbound

This paper cites criterion_1.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text criterion_1

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:52:24.301430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:edcfe02f1a2144108cd5566e9ae183b359247647d47e4d1201e92d1fb647da9d

Observation 2a659855-aabb-49c3-bad4-a3e3de8d1645 · outbound

This paper cites He stated that he would continue to work within the framework of the Philippine constitution and existing governmental structures.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text He stated that he would continue to work within the framework of the Philippine constitution and existing governmental structures

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:52:24.311963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:85a29e9cc43444e0882c21cd6a0a0799763fa0d43723124e948079c648883499

Observation bea8f7a6-a5f0-47e3-a56c-49d776582e49 · outbound

This paper cites name": "Correctness of the first part of the answer.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text name": "Correctness of the first part of the answer

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:52:24.266993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:c5a74beaef66c9c1662a9464ba56a88e8318f07ecbd4d4329d6047ad192c5cdb

Observation 4c7e9617-7256-4cd4-88ef-a583aebe27bc · outbound

This paper cites He stated that the declaration was made under duress due to threats from opposing groups and to protect national interest.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text He stated that the declaration was made under duress due to threats from opposing groups and to protect national interest

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:52:24.348055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:3b0e13ba93a4192ab074015a87a870f24c124587d3d10362e7f2e2894638b791

Observation abfbb882-59a3-4b30-b198-13856b5b3055 · outbound

This paper cites name": "Correctness of the first part of the answer.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text name": "Correctness of the first part of the answer

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:52:24.384418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:a2ae1ed3801dea49f3c592fdca6cb0cc86e222d103b4fde1f37d7442c3dd34db

Observation e573b348-c5b5-4967-8abb-f7271f4bf106 · outbound

This paper cites an unresolved cited work.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-05-23T01:52:24.262169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:c229bf056c3c606206383b79cccc4d2064e23de1938963399d661b385a7002c9

Observation c6c0fe7b-fe19-416e-a416-8d538ddf34da · outbound

This paper cites Ensure that it is necessary and sufficient to use these facts to derive the correct answer to the problem.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Ensure that it is necessary and sufficient to use these facts to derive the correct answer to the problem

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:52:24.382216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:e50255b85476ee083a4af17868575bd191b7443a0e05cc02a096259b1c707013

Observation b908ea69-d599-4d05-a49e-8de5f4d22450 · outbound

This paper cites * Provide a reference answer within <answer>...</answer> tags.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text * Provide a reference answer within <answer>...</answer> tags

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:52:24.255746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:d474fd59c6e970db7caa5b401a6ac74a8d473e30a80121f3b5689ca50313ec17

Observation a32a39af-36e9-4b1e-8309-4265a1689059 · outbound

This paper cites an unresolved cited work.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Unresolved cited work

Reference 67

Resolution
unresolved
raw_fallback, observed 2026-05-23T01:52:24.259907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:7ad2b610fc769caa268fcd98f96476b33487d672b432250d7738a2daef9ba344

Observation ef5d8dc3-4013-404c-941c-8e7f960e1109 · outbound

This paper cites ## Format * Enclose the question statement within <problem>...</problem> tags.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text ## Format * Enclose the question statement within <problem>...</problem> tags

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:52:24.264469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:57b6da18b148beb47b254dea6ee27bd823ee98fede2dfb7fc29f3cae1189605a

Observation 40fc4e96-66de-4203-b52b-5e63518f6c62 · outbound

This paper cites according to the text sample.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text according to the text sample

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:52:24.330038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:fc431a6f90b16861ac8dd99ef3e5121f9f4d8d9b9f6d69ae3a5669bae5205962

Observation 8af7f8f5-6931-4834-a013-b780d491fdc4 · outbound

This paper cites an unresolved cited work.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Unresolved cited work

Reference 70

Resolution
unresolved
raw_fallback, observed 2026-05-23T01:52:24.363704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:9f85158e8fd8415ef6b64db527b28eae31474b7f0162642f61d7b2edf022ad13

Observation 9ea86125-3e8a-4cca-893d-2c0ad75b6d66 · outbound

This paper cites Length: 1000 words.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Length: 1000 words

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:52:24.361507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:c38fa81077fda038488942c696e2b69a1c4364cba6568ccda2400af26e7a84df

Observation 4134d31f-b407-4820-a9b3-136d3806fb2b · outbound

This paper cites according to the knowledge.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text according to the knowledge

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:52:24.365723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:66af1656152ad2671d8c38d2d8e68ef16d968cbf8c28204d3ad0a63f0e435a62

Observation d778a6bd-ac8f-48d8-9a0b-f68d07aa19fb · outbound

This paper cites criterion_1.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text criterion_1

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:52:24.386631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:002554d3a53dd95c604a3016bc49bf9456058be2193d93e8d4d84b62e59932bc

Observation bf01da50-c8c1-499b-8678-5d123539d64e · outbound

This paper cites In general, the reference answer should have a high quality compared to the candidate answers, but this is not always true.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text In general, the reference answer should have a high quality compared to the candidate answers, but this is not always true

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:52:24.379607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:0b47108cf8e95e722ff6a6fc59e2794cfe465735266c08178069a5af2de5aa4d

Observation 994a32f2-05a2-41e0-8a3f-2c3120e3b373 · outbound

This paper cites an unresolved cited work.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Unresolved cited work

Reference 75

Resolution
unresolved
raw_fallback, observed 2026-05-23T01:52:24.372941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:1ad3418103ef02ff4dc4fa78affb909f5d481b182b5b8b3d2b52a9d6f36a26cd

Observation a7f4a737-b37d-4a43-80c2-08d93a67021d · outbound

This paper cites factuality.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text factuality

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:52:24.269116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:db55bece6f1a8ee4241097a4a3b218ba4539a2e529e6ce577e065ead183d8107

Observation e54c8d39-e205-41e6-98bf-944ee822c9c0 · outbound

This paper cites YYY" of.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text YYY" of

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:52:24.327410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:ce173d62e395def6d2ef92de25b49c578af1da14570026e1d1f36c14e722b8ee

Observation 97e0a20b-77e7-4aca-ad10-aece5e24e217 · outbound

This paper cites Not applicable.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Not applicable

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:52:24.370694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:70d729d31cde309ecc1b862b6236678609ad5ccb39c701b454c969e2a5e9cab2

Observation ccd96c6c-c2a5-401b-8cbe-6ffd1ce1abf2 · outbound

This paper cites XXX". To describe it, instead of saying.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text XXX". To describe it, instead of saying

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:52:24.375045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:af38cf5032eeaa68372b98729781150102ce18de8d67384d90f94f028f75086c

Observation e80b7bcd-7cdf-42d8-bdbd-e244093b742c · outbound

This paper cites criterion_1.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text criterion_1

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T01:52:24.377227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:c21d6afb681218d271be3d52730faac902f1c3a8657605a4a450af2f44e36250

Pith citing papers

Observation cbeba9fc-2602-4817-898f-741f2802f596 · inbound

SERPO: Self-Evolving Rubric Policy Optimization for Open-Ended Test-Time Reinforcement Learning cites this paper.

SERPO: Self-Evolving Rubric Policy Optimization for Open-Ended Test-Time Reinforcement Learning Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text

Reference 53

Resolution
unresolved
no resolver link, observed 2026-07-30T18:33:28.486944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T18:33:28.486944Z digest=sha256:81b639501379370d6a7222ac9e222b602be2c628abc4166b3cacfd2a34ab4731