Pith. sign in

Paper Citation Record · LEDGER

Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

As of 5 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 33 inbound Pith citation observations for arXiv:2405.01535.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2405.01535 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 33 of 33 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 33 of 33 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T22:34:14.479640Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-11T00:47:43.114558Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 590c7677-cd90-4db9-bdf5-23bcfbdf10ff · inbound

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods cites this paper.

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 114

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:08:36.997032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T23:08:34.312466Z digest=sha256:1c3f654622b056d7270503c60de324cf891970b00ee5f596e405d35c26e18cbc

Observation fdcb2bc2-df0e-490e-af8c-b4179bcfc620 · inbound

Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains cites this paper.

Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:07:56.779378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T06:07:56.678339Z digest=sha256:b45a18c96fa476a3712ea1f9e02ac2fb45de3861b29d9b5cc77f63db1a6cb773

Observation 36a0b4e8-50ac-4506-929d-f69b460d6c98 · inbound

Multi-Modal Requirements Data-based Acceptance Criteria Generation using LLMs cites this paper.

Multi-Modal Requirements Data-based Acceptance Criteria Generation using LLMs Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T22:34:14.479640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:34:14.479640Z digest=sha256:96d0465ec6700cae56e90e8aaf81489609d964d40c050addd2de8745f1ffa0f8

Observation 53653002-d5c0-4ce3-8d90-b996e49ae7ba · inbound

FHIR-RAG-MEDS: Integrating HL7 FHIR with Retrieval-Augmented Large Language Models for Enhanced Medical Decision Support cites this paper.

FHIR-RAG-MEDS: Integrating HL7 FHIR with Retrieval-Augmented Large Language Models for Enhanced Medical Decision Support Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-04T21:52:09.143146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:52:09.143146Z digest=sha256:3bfc7b950356e872d6623c6a51b1c208d127fe1da5e1cdf65a301ad76cf44128

Observation d647bb2f-22f3-4969-abb0-3bb9932d4832 · inbound

RLBFF: Binary Flexible Feedback to bridge between Human Feedback & Verifiable Rewards cites this paper.

RLBFF: Binary Flexible Feedback to bridge between Human Feedback & Verifiable Rewards Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T22:10:42.067732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-21T22:09:47.649346Z digest=sha256:90d366174f1156f9fe0e9f9f9df351f8b690df9730a209248caf270dcdf5718c

Observation 5b2270f2-4b7d-494f-997f-0e7fed373c85 · inbound

On the Shelf Life of Fine-Tuned LLM-Judges: Future-Proofing, Backward-Compatibility, and Question Generalization cites this paper.

On the Shelf Life of Fine-Tuned LLM-Judges: Future-Proofing, Backward-Compatibility, and Question Generalization Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T12:56:24.515871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T12:53:45.767341Z digest=sha256:a3591a35bb9cf97ebafbd9e3718adccc3af73ecf0ad04ccefae2f9c9d137ecd0

Observation d43b8ff0-58d4-4439-988b-4f5774d5b637 · inbound

Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process cites this paper.

Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T19:58:22.721897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T19:57:03.999154Z digest=sha256:2153ac156be504645d21ceecc8ba750bcd1f8bab0472f3145d6dcb58a679acc1

Observation a2bd1b92-28a1-409f-8155-49f11e65098e · inbound

Prompt Optimization Is a Coin Flip: Diagnosing When It Helps in Compound AI Systems cites this paper.

Prompt Optimization Is a Coin Flip: Diagnosing When It Helps in Compound AI Systems Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T11:20:10.362429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T11:18:09.127345Z digest=sha256:7acaee442bf51597baff6cc3a5db8292e654787fe4b57be513237a0177240230

Observation 5d883154-baf0-4768-a1f5-d9bc29f1ce91 · inbound

Beyond Verifiable Rewards: Rubric-Based GRM for Reinforced Fine-Tuning SWE Agents cites this paper.

Beyond Verifiable Rewards: Rubric-Based GRM for Reinforced Fine-Tuning SWE Agents Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:25:35.707687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T12:22:13.551709Z digest=sha256:643dbbae07989529af4d37ac4c28f19cedb6c3840413f5e1214f57d6ec9260e7

Observation c7ba134f-680e-41bc-9b2f-bf76c8548b26 · inbound

KnowPilot: Your Knowledge-Driven Copilot for Domain Tasks cites this paper.

KnowPilot: Your Knowledge-Driven Copilot for Domain Tasks Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:16:20.733143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:15:44.360621Z digest=sha256:39a01680f271d6820daa83e4ece0fd0ddcd2b7f0ea9e23dfcde4c266c373b9df

Observation e028541d-1397-4718-bef4-e099a2e7c074 · inbound

Judging the Judges: A Systematic Evaluation of Bias Mitigation Strategies in LLM-as-a-Judge Pipelines cites this paper.

Judging the Judges: A Systematic Evaluation of Bias Mitigation Strategies in LLM-as-a-Judge Pipelines Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:41:14.183428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T08:14:18.535385Z digest=sha256:f6490cebd320b43cb5e88806cb4df34860cf2923cfe001638d85717ce55797f6

Observation 796152a9-b5b4-44b6-9ca0-b548b920c51e · inbound

HalluScan: A Systematic Benchmark for Detecting and Mitigating Hallucinations in Instruction-Following LLMs cites this paper.

HalluScan: A Systematic Benchmark for Detecting and Mitigating Hallucinations in Instruction-Following LLMs Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-25T06:50:27.967610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-25T06:49:16.755597Z digest=sha256:7c554d0452cc341c85678071b165fdb26d87ae726775b6b9bb312537be27bf99

Observation 4762d278-be0f-4bdf-9748-d04d9f55649b · inbound

AgentTrust: Runtime Safety Evaluation and Interception for AI Agent Tool Use cites this paper.

AgentTrust: Runtime Safety Evaluation and Interception for AI Agent Tool Use Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:26:06.983575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T17:34:11.020584Z digest=sha256:68d09234cc93955ee2e319fd8be4573583103f44c2f2eed81008c55e5ad37049

Observation 7c31b613-fbc1-462f-8110-d7ce0056ad2e · inbound

Rubric-Grounded RL: Structured Judge Rewards for Generalizable Reasoning cites this paper.

Rubric-Grounded RL: Structured Judge Rewards for Generalizable Reasoning Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T03:45:58.155751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T02:18:33.057569Z digest=sha256:70cf9e3867d2efbf4aa7759a866d8b1aced34e2c5b65da4093c350a81ded9db4

Observation f47f4383-3c86-4a3a-b47e-78f3f7bcb3cd · inbound

TRACE: A taxonomy-grounded synthetic dataset for teaching-program generation and session interpretation in Applied Behavior Analysis cites this paper.

TRACE: A taxonomy-grounded synthetic dataset for teaching-program generation and session interpretation in Applied Behavior Analysis Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T12:04:38.868040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-30T12:00:02.323398Z digest=sha256:ae25373ac867091ffe6456226bd3b952dd1415132f387b88af6228ac853b94dc

Observation 14bebf32-d519-45d4-b38e-77762109313d · inbound

DeepSurvey: Enhancing Analytical Depth and Citation Reliability in Automated Survey Generation cites this paper.

DeepSurvey: Enhancing Analytical Depth and Citation Reliability in Automated Survey Generation Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T07:13:16.182183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T07:12:13.559728Z digest=sha256:ee98e4fa5923a0cf63da1cd461d8b6b4e609adee63cdfba5cc7e741827a5cb9c

Observation 0c724aff-f1b6-4fe7-b24e-7e310d35c157 · inbound

CoEval: Ranking Language Models for Custom Tasks Without Labeled Data or Trustworthy Benchmarks cites this paper.

CoEval: Ranking Language Models for Custom Tasks Without Labeled Data or Trustworthy Benchmarks Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:36:27.430493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T10:46:24.554332Z digest=sha256:aeda81bc1d7c7a5f1f17faf9a180080e431e2268fc2c95e2a992edc7e07d5eab

Observation 0f471fc0-fa12-414d-96d3-8c79df5dad8f · inbound

Organizational Control Layer: Governance Infrastructure at the Execution Boundary of LLM Agent Systems cites this paper.

Organizational Control Layer: Governance Infrastructure at the Execution Boundary of LLM Agent Systems Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-02T11:16:53.801408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-28T04:17:18.483765Z digest=sha256:e0f55ef13ec2a795c65b95f01a97f53ba984603bc2b7c1113537f37684ed027b

Observation 611005d7-db0d-4c40-8716-3a1256b735de · inbound

Beyond Rubrics: Exploration-Guided Evaluation Skills for Reward Modeling cites this paper.

Beyond Rubrics: Exploration-Guided Evaluation Skills for Reward Modeling Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T17:07:12.476685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T22:14:11.926333Z digest=sha256:813f68809efd0fa9b169eec073dd20827a997239b2cbf437dc9b0ea34dee4cdc

Observation 223b9bde-65a6-47db-a9bc-fc0ff814a6b7 · inbound

PoQ-Judge: A Multi-Architecture Evaluation Framework for Cost-Aware Proof-of-Quality in Decentralized LLM Inference cites this paper.

PoQ-Judge: A Multi-Architecture Evaluation Framework for Cost-Aware Proof-of-Quality in Decentralized LLM Inference Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-07-05T16:31:16.126945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-05T16:23:12.903749Z digest=sha256:a45c6a16e252144496aa7e735b679c84244b10442c4f65b0d581fcebcd9a76fe

Observation 9fe2139e-23b4-4d39-86c4-6d1aa8555d4b · inbound

Quantifying and Auditing LLM Evaluation via Positive--Unlabeled Learning cites this paper.

Quantifying and Auditing LLM Evaluation via Positive--Unlabeled Learning Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-04T02:49:24.811672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T19:04:45.062426Z digest=sha256:951d123c7f1a5d8936e0241815ac5e97ad527d6bdfcca56db7bc589f3321a320

Observation 00cfc5bf-e014-4d94-bb5c-8fabface31f9 · inbound

CourseBlueprint: A Structured Pipeline for Adaptive Pedagogical Video Generation Grounded in Course Corpora cites this paper.

CourseBlueprint: A Structured Pipeline for Adaptive Pedagogical Video Generation Grounded in Course Corpora Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T15:24:50.063709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T15:19:30.276155Z digest=sha256:faa3676d036b2ea52cff5d3bde6fb97965ef9798225bc1d1a67cb65a7f21eb38

Observation b5da3bcb-dd04-41a4-94ed-70f9892a995f · inbound

Evaluation Awareness Is Not One Capability: Evidence from Open Language Models cites this paper.

Evaluation Awareness Is Not One Capability: Evidence from Open Language Models Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:49:46.884857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-26T08:23:20.122338Z digest=sha256:3b4ef275cb2b6084e5da628494621ba146176c08908ce0e7208fa75251957ef1

Observation 3d8ef08f-a664-4ec4-9cfc-1e7fa7f9c790 · inbound

Open Problems in Constitutional Preference Reconstruction cites this paper.

Open Problems in Constitutional Preference Reconstruction Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-06-30T06:54:20.507786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T06:49:24.255622Z digest=sha256:809ec5a93d65c41ca1b81e84044bfe34f1aa69919e55cd64c34a1cadb7d6d932

Observation f7fdd0ba-10c1-494e-b0cf-9c03b7fb6d45 · inbound

RoPoLL: Robust Panel of LLM Judges cites this paper.

RoPoLL: Robust Panel of LLM Judges Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-01T12:55:44.495867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T01:34:31.158023Z digest=sha256:84e1632d0e6dd7882e1cb129e2e1cc2cf504186b62a952af21e67851b61a64db

Observation bac534d8-fbc6-46c3-8931-61fe4213d9de · inbound

Healthier LLMs: Retrieval-Augmented Generation for Public Health Question Answering cites this paper.

Healthier LLMs: Retrieval-Augmented Generation for Public Health Question Answering Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-07-11T00:47:43.136596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-11T00:44:36.322629Z digest=sha256:a53fbb096cd2bd9781e5e4588a0cdca1ea4960a4b65a2d47b75d01eba5e3d29d

Observation b8e4f9ee-1810-4140-8569-9b7ccc42b505 · inbound

When the Judge Changes, So Does the Measurement: Auditing LLM-as-Judge Reliability cites this paper.

When the Judge Changes, So Does the Measurement: Auditing LLM-as-Judge Reliability Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-10T05:56:50.556375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-10T05:47:32.670216Z digest=sha256:b8deba975e951e53f8f6168d08e3557206f20ce90446c92019aaaae8c78abba0

Observation b90ccc66-222a-4d14-ade8-1e5bd0942a48 · inbound

Autoregressive Modeling of Film with Applications in Video Montage cites this paper.

Autoregressive Modeling of Film with Applications in Video Montage Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:19.642469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:34:19.642469Z digest=sha256:e5b773998c19ec6dd7761a2b8487bca95d5247280ab78a316f8e4bbc7245600d

Observation 4b2a61fe-472f-44f8-84cd-293b4924161f · inbound

BACON: Budgeted Human Calibration for Modeling and Evaluation with Multiple AI Judges cites this paper.

BACON: Budgeted Human Calibration for Modeling and Evaluation with Multiple AI Judges Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T10:02:37.926198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T10:02:37.926198Z digest=sha256:cf968c8624274c698c724bbc156e24baf5c38430710136354be8b36e5a8d71df

Observation da6dcea5-c533-4d7a-bd14-53e25c6e8103 · inbound

BACON: Budgeted Human Calibration for Modeling and Evaluation with Multiple AI Judges cites this paper.

BACON: Budgeted Human Calibration for Modeling and Evaluation with Multiple AI Judges Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-02T10:02:41.214150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T10:02:41.214150Z digest=sha256:c53251c1bd6ccde9bf8be19678321a0cac1b78a7b3b398e40a1d4d23e49730e7

Observation 21b9f970-9a5d-43b6-932e-bd504a854b9a · inbound

Summary of DCASE 2026 Task 5: Audio-Dependent Question Answering cites this paper.

Summary of DCASE 2026 Task 5: Audio-Dependent Question Answering Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T14:37:34.810212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:37:34.810212Z digest=sha256:6c725a871af7956f1b8913174d7b86fd0fcf70f7af2c68274ab9f19731d3d57c

Observation db556d63-3042-4177-93f2-9f33928644b8 · inbound

SERPO: Self-Evolving Rubric Policy Optimization for Open-Ended Test-Time Reinforcement Learning cites this paper.

SERPO: Self-Evolving Rubric Policy Optimization for Open-Ended Test-Time Reinforcement Learning Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-07-30T18:33:28.502791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T18:33:28.502791Z digest=sha256:845bdc6914599ef20f6cc9945d7206f995a7db4e569f8fdbae69c564506a1dc3

Observation a06887da-3a93-419c-980d-a3cfc432c04b · inbound

SERPO: Self-Evolving Rubric Policy Optimization for Open-Ended Test-Time Reinforcement Learning cites this paper.

SERPO: Self-Evolving Rubric Policy Optimization for Open-Ended Test-Time Reinforcement Learning Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T01:47:19.245261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T01:47:19.245261Z digest=sha256:62eb4b8b042e4b884fabe7cf1d0deedcb3a34ad457eb61882344fcd6d5aab6e1