Pith. sign in

Paper Citation Record · LEDGER

Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 65 inbound Pith citation observations for arXiv:2405.01535.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2405.01535 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 65 of 65 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 65 of 65 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:31:57.822140Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-11T00:47:43.114558Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e798bd68-e6a6-4124-8cd8-31a2e27dc89b · inbound

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set cites this paper.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:57.390500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:57.390500Z digest=sha256:18b9c7e0f60a1615286981204d5e4c7c0f2330e606018ca56332da44e8525007

Observation 4952cc09-d83e-42c5-8f0c-60f732d50cf1 · inbound

Self-Generated Critiques Boost Reward Modeling for Language Models cites this paper.

Self-Generated Critiques Boost Reward Modeling for Language Models Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T12:58:29.634019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:58:29.634019Z digest=sha256:7d3a6940f3d888b91b2ef1be1c9b0c48ed56dcdd34ac28b58470b5a4856726c7

Observation ab1e1182-0883-4f44-824b-3f94b995133e · inbound

Know Your RAG: Dataset Taxonomy and Generation Strategies for Evaluating RAG Systems cites this paper.

Know Your RAG: Dataset Taxonomy and Generation Strategies for Evaluating RAG Systems Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T05:59:37.098468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:59:37.098468Z digest=sha256:a680182fb782799fcbe0755f63a6c245726efe5af441ece1d38794cfffd4f659

Observation 590c7677-cd90-4db9-bdf5-23bcfbdf10ff · inbound

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods cites this paper.

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 114

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:08:36.997032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T23:08:34.312466Z digest=sha256:d000a70e8fe16000b80a2c8e3f9e826b05ba3dd597ae9ecc236ad7e195e5f71a

Observation f3e320c6-e7ae-4fc5-b5b9-607cc2e7ea25 · inbound

Copyright-Protected Language Generation via Adaptive Model Fusion cites this paper.

Copyright-Protected Language Generation via Adaptive Model Fusion Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T19:34:17.652399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:34:17.652399Z digest=sha256:8c6a2313ba00d8c2fd0fc8820efe58c77d924aa9bb1e89a9c8a52ba10d0cf7eb

Observation b0784781-3e87-41b0-8bd2-37d74582ca1c · inbound

CoPrUS: Consistency Preserving Utterance Synthesis towards more realistic benchmark dialogues cites this paper.

CoPrUS: Consistency Preserving Utterance Synthesis towards more realistic benchmark dialogues Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T18:50:02.236431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T18:50:02.236431Z digest=sha256:2bf20afdd5db7992948a3f67007e5c1105e11c480c10a51af2a9fb6df8b31eaa

Observation 8349407d-6001-4c04-8785-882126d54c9b · inbound

Disentangling Preference Representation and Text Generation for Efficient Individual Preference Alignment cites this paper.

Disentangling Preference Representation and Text Generation for Efficient Individual Preference Alignment Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:26.722279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:26.722279Z digest=sha256:c137d8a13bdffbd9115d812ab58893d6b3785121e5ead21acfef05eea952990f

Observation c065bc01-aca6-499e-9fdd-dfc70ec4dbf6 · inbound

Influences on LLM Calibration: A Study of Response Agreement, Loss Functions, and Prompt Styles cites this paper.

Influences on LLM Calibration: A Study of Response Agreement, Loss Functions, and Prompt Styles Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T21:45:55.313121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:45:55.313121Z digest=sha256:fd08a6c822e4c7dcf5f2535ebbb7ce6e7014034e39c512cd25036d0958b11533

Observation dc6fb207-9096-4997-a5f1-528bfcb8ecfc · inbound

SedarEval: Automated Evaluation using Self-Adaptive Rubrics cites this paper.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T14:13:21.837440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:13:21.837440Z digest=sha256:ab644d63fbc21ff27806aed28b8db91e54b75c3156bd029c03be623ae4ad9574

Observation 79e9433a-0aff-4a2c-b88b-8e9db9716d59 · inbound

Atla Selene Mini: A General Purpose Evaluation Model cites this paper.

Atla Selene Mini: A General Purpose Evaluation Model Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T13:47:45.802686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:47:45.802686Z digest=sha256:a963114c17f735f9a46f3a727bf7f58a918796de76e19baeb4bb2e054f6d35bc

Observation 683d7c4b-09d5-45cd-aac8-4830e0839131 · inbound

Can LLMs Replace Human Evaluators? An Empirical Study of LLM-as-a-Judge in Software Engineering cites this paper.

Can LLMs Replace Human Evaluators? An Empirical Study of LLM-as-a-Judge in Software Engineering Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T16:31:38.573942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:31:38.573942Z digest=sha256:1a31e2c1997cf467f95cf8b64d42d8bc61978550d2b839ed63bc59a378f082eb

Observation 5904f9a0-4020-4755-8ee1-a246be78e0f0 · inbound

Automatic Evaluation of Healthcare LLMs Beyond Question-Answering cites this paper.

Automatic Evaluation of Healthcare LLMs Beyond Question-Answering Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T14:46:29.642901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:46:29.642901Z digest=sha256:c75a28e63074afebdb237be8bf4758d74a952df2066360a0972d5abe9627c334

Observation 5a5e4f0d-2c60-4e5b-a850-f6e3a454a01e · inbound

Salamandra Technical Report cites this paper.

Salamandra Technical Report Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-08T04:58:32.673443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:58:32.673443Z digest=sha256:344e91182411c71139095c0c182dceebb5f503bfb47c745aaccc2879dcdd3713

Observation 1909eda2-2b26-40df-8dff-e0c87713161f · inbound

Persona-judge: Personalized Alignment of Large Language Models via Token-level Self-judgment cites this paper.

Persona-judge: Personalized Alignment of Large Language Models via Token-level Self-judgment Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T12:31:57.822140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:31:57.822140Z digest=sha256:d7879355c9180bd58d1b8f567062b72aa5de5c638cee59501682138161c91136

Observation f0e28339-24c1-4110-8269-34221fde98fb · inbound

PROMPTEVALS: A Dataset of Assertions and Guardrails for Custom Production Large Language Model Pipelines cites this paper.

PROMPTEVALS: A Dataset of Assertions and Guardrails for Custom Production Large Language Model Pipelines Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T11:45:45.715779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:45:45.715779Z digest=sha256:70568ae73f4839a0261f214d58fa54537253ed496ea3fbae8fa40244c350da2a

Observation 3598cd01-4af1-4f62-99a6-b6773f0a2feb · inbound

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators cites this paper.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.408060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.408060Z digest=sha256:b50ea214b7d6bb11c0eb6a6b8c2860593bb2c711032d96f43336e6058afce743

Observation 6f389968-28ad-466b-bb34-36db4df252fc · inbound

Trillion 7B Technical Report cites this paper.

Trillion 7B Technical Report Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T11:32:05.945961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:32:05.945961Z digest=sha256:be84db48fa25120e2d68dce867a4ebd0df2d4292919e5d66aaf7404e5096960b

Observation ff496347-eb8e-41e4-b2b0-c892c5a0654b · inbound

Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models cites this paper.

Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T20:32:25.158010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:32:25.158010Z digest=sha256:8bf47c88fea5867d41a56e273877701f4e50f627dda80383329f1e01925c3a6e

Observation a95a7f08-ce38-4ffc-bf63-5094d36e1b5a · inbound

Safety Degradation in AI Agents cites this paper.

Safety Degradation in AI Agents Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T15:41:06.288324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:41:06.288324Z digest=sha256:0c5c951c6d5212a139513566916116c0f77a5ba993fc43e53412d56283865f96

Observation 512b44bf-1a37-4f4e-a85a-8f1394205af5 · inbound

YESciEval: Robust LLM-as-a-Judge for Scientific Question Answering cites this paper.

YESciEval: Robust LLM-as-a-Judge for Scientific Question Answering Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T15:41:09.659621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:41:09.659621Z digest=sha256:e9901d2c70f75165ce1f891dca14fb072c6d046390ea38b5be5bfdb46da66ed6

Observation 1c1e8b91-a8e6-435b-96cf-4d46f5630f26 · inbound

DECASTE: Unveiling Caste Stereotypes in Large Language Models through Multi-Dimensional Bias Analysis cites this paper.

DECASTE: Unveiling Caste Stereotypes in Large Language Models through Multi-Dimensional Bias Analysis Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T15:30:41.924851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:30:41.924851Z digest=sha256:a0220d59c9643073c8c63db8a0180067256a93aa3f1f8ede3f3a498930f8465c

Observation 58ccd800-b8e4-404c-93cb-c281b0a29015 · inbound

Improving Fairness of Large Language Models in Multi-document Summarization cites this paper.

Improving Fairness of Large Language Models in Multi-document Summarization Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T05:41:12.789750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:41:12.789750Z digest=sha256:506c7b5e3987b8ca53341094d724e2723abb16f0456f5ee776585a83c9f9da5c

Observation 50201047-13f5-4e3b-a1ed-20cad6b23555 · inbound

Outside Knowledge Conversational Video (OKCV) Dataset -- Dialoguing over Videos cites this paper.

Outside Knowledge Conversational Video (OKCV) Dataset -- Dialoguing over Videos Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:43.101253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:43.101253Z digest=sha256:6627ccbed2bf077cfe66e6da9cbe99b758b63b5abf6ffd26b62bfdbed6d2c2ef

Observation 584a9fc1-5e42-42ae-911f-6acb49934e14 · inbound

A Survey of Automatic Evaluation Methods on Text, Visual and Speech Generations cites this paper.

A Survey of Automatic Evaluation Methods on Text, Visual and Speech Generations Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 146

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:46.727694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:46.727694Z digest=sha256:032fd30400e2e0b8ed8ee424bab6ad89bb86161e7c2b82b16ed858d057a34a00

Observation efa151a7-b8e1-4f0f-80bf-09e365e4def7 · inbound

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation cites this paper.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:47.156442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:47.156442Z digest=sha256:642f36dd7778ea5ff7ece5805f08d8c5f42c97f06b54432bff4a5ab6b4d444f9

Observation d51b3700-0f16-4a5f-b047-83286e9cbf86 · inbound

MinosEval: Distinguishing Factoid and Non-Factoid for Tailored Open-Ended QA Evaluation with LLMs cites this paper.

MinosEval: Distinguishing Factoid and Non-Factoid for Tailored Open-Ended QA Evaluation with LLMs Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T19:45:00.322748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:45:00.322748Z digest=sha256:0b5aab0175977e25e0f8aed8535f3f3607e0db543b6f8413e4f6bd3fd2754dd9

Observation ff9758d9-04ae-4b3e-9200-95eb543cfa8a · inbound

FairI Tales: Evaluation of Fairness in Indian Contexts with a Focus on Bias and Stereotypes cites this paper.

FairI Tales: Evaluation of Fairness in Indian Contexts with a Focus on Bias and Stereotypes Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T21:54:43.729141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:54:43.729141Z digest=sha256:5183ce94583ee9baa018c6e505298ee69578f103fe903122635b0d1af0be7605

Observation 1655dd3d-0ca5-4a26-92e2-4b204c58f7e4 · inbound

Hierarchical Memory Organization for Wikipedia Generation cites this paper.

Hierarchical Memory Organization for Wikipedia Generation Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:57.815566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:48:57.815566Z digest=sha256:f59bfb8c0eee862d4f9d0e95c1c916d5b017277ebfa1701934f901e41df1bfca

Observation f5c69d9e-df55-40ca-bd3d-93e7e01a6879 · inbound

Interpretable Mnemonic Generation for Kanji Learning via Expectation-Maximization cites this paper.

Interpretable Mnemonic Generation for Kanji Learning via Expectation-Maximization Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T19:37:07.993176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:37:07.993176Z digest=sha256:60e9b396266a3b719de8153e6a4b5a5a97c5e815f2b09f4fc82de5d694e1f836

Observation fdcb2bc2-df0e-490e-af8c-b4179bcfc620 · inbound

Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains cites this paper.

Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:07:56.779378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-13T06:07:56.678339Z digest=sha256:edb936a2650e17feaa82eada89da30032cec252c8bd572d1374ef26147e3bdf7

Observation 36a0b4e8-50ac-4506-929d-f69b460d6c98 · inbound

Multi-Modal Requirements Data-based Acceptance Criteria Generation using LLMs cites this paper.

Multi-Modal Requirements Data-based Acceptance Criteria Generation using LLMs Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T22:34:14.479640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:34:14.479640Z digest=sha256:732639ac919c3039c798ea9665567cee475545b77a1ea1941922769d29bc6b4c

Observation 53653002-d5c0-4ce3-8d90-b996e49ae7ba · inbound

FHIR-RAG-MEDS: Integrating HL7 FHIR with Retrieval-Augmented Large Language Models for Enhanced Medical Decision Support cites this paper.

FHIR-RAG-MEDS: Integrating HL7 FHIR with Retrieval-Augmented Large Language Models for Enhanced Medical Decision Support Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-04T21:52:09.143146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:52:09.143146Z digest=sha256:b211b430c62dc57f266c947b413cdc76850cfe2cc9094f26b2d2c3046ddc248e

Observation d647bb2f-22f3-4969-abb0-3bb9932d4832 · inbound

RLBFF: Binary Flexible Feedback to bridge between Human Feedback & Verifiable Rewards cites this paper.

RLBFF: Binary Flexible Feedback to bridge between Human Feedback & Verifiable Rewards Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T22:10:42.067732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-21T22:09:47.649346Z digest=sha256:c851398dbfa2586e934ad4b8c29d5ae9e30ad1a82d900eecd082d620a46f7ec6

Observation 5b2270f2-4b7d-494f-997f-0e7fed373c85 · inbound

On the Shelf Life of Fine-Tuned LLM-Judges: Future-Proofing, Backward-Compatibility, and Question Generalization cites this paper.

On the Shelf Life of Fine-Tuned LLM-Judges: Future-Proofing, Backward-Compatibility, and Question Generalization Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T12:56:24.515871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-18T12:53:45.767341Z digest=sha256:78b7a6e3fc245602c18d5b13392d790323409f9b7a3c9601c0c4dc3dc55c1024

Observation d43b8ff0-58d4-4439-988b-4f5774d5b637 · inbound

Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process cites this paper.

Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T19:58:22.721897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T19:57:03.999154Z digest=sha256:3a5ba708fde4727e9799e23d2148f109038ff92d942963742164e5830dbc34fc

Observation a2bd1b92-28a1-409f-8155-49f11e65098e · inbound

Prompt Optimization Is a Coin Flip: Diagnosing When It Helps in Compound AI Systems cites this paper.

Prompt Optimization Is a Coin Flip: Diagnosing When It Helps in Compound AI Systems Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T11:20:10.362429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T11:18:09.127345Z digest=sha256:1223e7e28b36c0629ddc098ea9460c3c7a7e900e8a28a70535964e57154642b7

Observation 5d883154-baf0-4768-a1f5-d9bc29f1ce91 · inbound

Beyond Verifiable Rewards: Rubric-Based GRM for Reinforced Fine-Tuning SWE Agents cites this paper.

Beyond Verifiable Rewards: Rubric-Based GRM for Reinforced Fine-Tuning SWE Agents Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:25:35.707687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-15T12:22:13.551709Z digest=sha256:3f6f6fb56aaba45042ce7e413dad0f735a6b635cb76e0673576fefec55fab34f

Observation c7ba134f-680e-41bc-9b2f-bf76c8548b26 · inbound

KnowPilot: Your Knowledge-Driven Copilot for Domain Tasks cites this paper.

KnowPilot: Your Knowledge-Driven Copilot for Domain Tasks Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:16:20.733143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T06:15:44.360621Z digest=sha256:bbfd602b9d9e62746637f97506afab6ded13d7710e532fc645e3f5c2e40f380e

Observation e028541d-1397-4718-bef4-e099a2e7c074 · inbound

Judging the Judges: A Systematic Evaluation of Bias Mitigation Strategies in LLM-as-a-Judge Pipelines cites this paper.

Judging the Judges: A Systematic Evaluation of Bias Mitigation Strategies in LLM-as-a-Judge Pipelines Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:41:14.183428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-08T08:14:18.535385Z digest=sha256:55e3da4819ed377c2ddf91d90fe881a17a31a1e23a58d0e76407cc7cf37411a8

Observation 796152a9-b5b4-44b6-9ca0-b548b920c51e · inbound

HalluScan: A Systematic Benchmark for Detecting and Mitigating Hallucinations in Instruction-Following LLMs cites this paper.

HalluScan: A Systematic Benchmark for Detecting and Mitigating Hallucinations in Instruction-Following LLMs Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-25T06:50:27.967610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-25T06:49:16.755597Z digest=sha256:ee321f8a0c153497054ad455a93c69ce2b20a0e5832c76a5818d338e9a77e2ec

Observation 4762d278-be0f-4bdf-9748-d04d9f55649b · inbound

AgentTrust: Runtime Safety Evaluation and Interception for AI Agent Tool Use cites this paper.

AgentTrust: Runtime Safety Evaluation and Interception for AI Agent Tool Use Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:26:06.983575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-08T17:34:11.020584Z digest=sha256:8fc1999af49a053c44787ddb304bd8b832f9f8592048683ba3e2b9198b1c7c17

Observation 7c31b613-fbc1-462f-8110-d7ce0056ad2e · inbound

Rubric-Grounded RL: Structured Judge Rewards for Generalizable Reasoning cites this paper.

Rubric-Grounded RL: Structured Judge Rewards for Generalizable Reasoning Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T03:45:58.155751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T02:18:33.057569Z digest=sha256:68b67c849694f8304d6256c74652ef8872bd8e9ce80678e711c8f1fd1d96014b

Observation f47f4383-3c86-4a3a-b47e-78f3f7bcb3cd · inbound

TRACE: A taxonomy-grounded synthetic dataset for teaching-program generation and session interpretation in Applied Behavior Analysis cites this paper.

TRACE: A taxonomy-grounded synthetic dataset for teaching-program generation and session interpretation in Applied Behavior Analysis Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T12:04:38.868040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-30T12:00:02.323398Z digest=sha256:e48556f0bbe5eda2ec9e26398cdd28e9387de2f0653328b5e088afbeef4d0d25

Observation 14bebf32-d519-45d4-b38e-77762109313d · inbound

DeepSurvey: Enhancing Analytical Depth and Citation Reliability in Automated Survey Generation cites this paper.

DeepSurvey: Enhancing Analytical Depth and Citation Reliability in Automated Survey Generation Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T07:13:16.182183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T07:12:13.559728Z digest=sha256:cf3b8fdfc5910eb91f909a8d7cf40aacf90859146237a349ad962a6da62225eb

Observation 0c724aff-f1b6-4fe7-b24e-7e310d35c157 · inbound

CoEval: Ranking Language Models for Custom Tasks Without Labeled Data or Trustworthy Benchmarks cites this paper.

CoEval: Ranking Language Models for Custom Tasks Without Labeled Data or Trustworthy Benchmarks Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:36:27.430493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T10:46:24.554332Z digest=sha256:9a7bbd7d7f4f014e305c06f878a5e02d04b2cf238bcba72203c0219b27f2c561

Observation 0f471fc0-fa12-414d-96d3-8c79df5dad8f · inbound

Organizational Control Layer: Governance Infrastructure at the Execution Boundary of LLM Agent Systems cites this paper.

Organizational Control Layer: Governance Infrastructure at the Execution Boundary of LLM Agent Systems Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-02T11:16:53.801408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-28T04:17:18.483765Z digest=sha256:8423dc0da00d85659c397939cfb7394fb72038a5f1ee5911e374605e483c476a

Observation 611005d7-db0d-4c40-8716-3a1256b735de · inbound

Beyond Rubrics: Exploration-Guided Evaluation Skills for Reward Modeling cites this paper.

Beyond Rubrics: Exploration-Guided Evaluation Skills for Reward Modeling Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T17:07:12.476685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-27T22:14:11.926333Z digest=sha256:e0c3a12b3cb0b1c0d10fc07d34d199ce040cdf8766e944f86659f38e857e36ab

Observation 223b9bde-65a6-47db-a9bc-fc0ff814a6b7 · inbound

PoQ-Judge: A Multi-Architecture Evaluation Framework for Cost-Aware Proof-of-Quality in Decentralized LLM Inference cites this paper.

PoQ-Judge: A Multi-Architecture Evaluation Framework for Cost-Aware Proof-of-Quality in Decentralized LLM Inference Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-07-05T16:31:16.126945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-05T16:23:12.903749Z digest=sha256:23ac2be8ea532ec8cc93539efb666b98b5f1d5061c655cc14e2c46c6975709fd

Observation 9fe2139e-23b4-4d39-86c4-6d1aa8555d4b · inbound

Quantifying and Auditing LLM Evaluation via Positive--Unlabeled Learning cites this paper.

Quantifying and Auditing LLM Evaluation via Positive--Unlabeled Learning Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-04T02:49:24.811672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-26T19:04:45.062426Z digest=sha256:9824fa386f296025ead4d43775bd33718ae12d5bad1e96596613a276f6d4a847

Observation 00cfc5bf-e014-4d94-bb5c-8fabface31f9 · inbound

CourseBlueprint: A Structured Pipeline for Adaptive Pedagogical Video Generation Grounded in Course Corpora cites this paper.

CourseBlueprint: A Structured Pipeline for Adaptive Pedagogical Video Generation Grounded in Course Corpora Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T15:24:50.063709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-30T15:19:30.276155Z digest=sha256:e3c1b9ecd3cbb92227aed5a840b85c6c165bb7a51abd505c743c0d40f8d69633

Observation b5da3bcb-dd04-41a4-94ed-70f9892a995f · inbound

Evaluation Awareness Is Not One Capability: Evidence from Open Language Models cites this paper.

Evaluation Awareness Is Not One Capability: Evidence from Open Language Models Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:49:46.884857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-26T08:23:20.122338Z digest=sha256:aa2e1a07e01b47de3bdf262892830a7dda1ac80d3eb38da86a22f9ad48cdc0bf

Observation 3d8ef08f-a664-4ec4-9cfc-1e7fa7f9c790 · inbound

Open Problems in Constitutional Preference Reconstruction cites this paper.

Open Problems in Constitutional Preference Reconstruction Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-06-30T06:54:20.507786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-30T06:49:24.255622Z digest=sha256:0ed697eca125ddca2768d521203c60c104835384904ba42b839d4a87f4cfd4e5

Observation f7fdd0ba-10c1-494e-b0cf-9c03b7fb6d45 · inbound

RoPoLL: Robust Panel of LLM Judges cites this paper.

RoPoLL: Robust Panel of LLM Judges Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-01T12:55:44.495867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T01:34:31.158023Z digest=sha256:62248790b56747da79cf5edffba1fe4133050cd1f4753121107dc70039f65aea

Observation bac534d8-fbc6-46c3-8931-61fe4213d9de · inbound

Healthier LLMs: Retrieval-Augmented Generation for Public Health Question Answering cites this paper.

Healthier LLMs: Retrieval-Augmented Generation for Public Health Question Answering Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-07-11T00:47:43.136596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-11T00:44:36.322629Z digest=sha256:190c9a18cc9e3d4c8347006977a093fd338fcae17ce90e4a6d21638d3a5077cc

Observation b8e4f9ee-1810-4140-8569-9b7ccc42b505 · inbound

When the Judge Changes, So Does the Measurement: Auditing LLM-as-Judge Reliability cites this paper.

When the Judge Changes, So Does the Measurement: Auditing LLM-as-Judge Reliability Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-10T05:56:50.556375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-10T05:47:32.670216Z digest=sha256:6be2de8c2fd63a57c450a1f1e593d7ce0f7f2c0b58e4d52d4cf9ead0d356aef4

Observation b90ccc66-222a-4d14-ade8-1e5bd0942a48 · inbound

Autoregressive Modeling of Film with Applications in Video Montage cites this paper.

Autoregressive Modeling of Film with Applications in Video Montage Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:19.642469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:34:19.642469Z digest=sha256:51aeeef2a45b2adad860b5fb25a8b8f5084b1570e9219dee98704ef567cc4913

Observation 4b2a61fe-472f-44f8-84cd-293b4924161f · inbound

BACON: Budgeted Human Calibration for Modeling and Evaluation with Multiple AI Judges cites this paper.

BACON: Budgeted Human Calibration for Modeling and Evaluation with Multiple AI Judges Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T10:02:37.926198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T10:02:37.926198Z digest=sha256:6ab8f82548d33c918bd3e5550104d531418480f8997ac15d227cc38cfd45cb6e

Observation da6dcea5-c533-4d7a-bd14-53e25c6e8103 · inbound

BACON: Budgeted Human Calibration for Modeling and Evaluation with Multiple AI Judges cites this paper.

BACON: Budgeted Human Calibration for Modeling and Evaluation with Multiple AI Judges Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-02T10:02:41.214150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T10:02:41.214150Z digest=sha256:ea203162593ba611a962928075b4f972e71682a8044dc140fc3bf627511d99c7

Observation 21b9f970-9a5d-43b6-932e-bd504a854b9a · inbound

Summary of DCASE 2026 Task 5: Audio-Dependent Question Answering cites this paper.

Summary of DCASE 2026 Task 5: Audio-Dependent Question Answering Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T14:37:34.810212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:37:34.810212Z digest=sha256:e4bf8228db6adfc1663a8330da4cf59a5cfc2e0c310041f6d71c5b0adb19ece2

Observation db556d63-3042-4177-93f2-9f33928644b8 · inbound

SERPO: Self-Evolving Rubric Policy Optimization for Open-Ended Test-Time Reinforcement Learning cites this paper.

SERPO: Self-Evolving Rubric Policy Optimization for Open-Ended Test-Time Reinforcement Learning Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-07-30T18:33:28.502791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T18:33:28.502791Z digest=sha256:839e5060e9e0f742727884641e887695ea05313aeca70c0bb16f0cd7bd371917

Observation a06887da-3a93-419c-980d-a3cfc432c04b · inbound

SERPO: Self-Evolving Rubric Policy Optimization for Open-Ended Test-Time Reinforcement Learning cites this paper.

SERPO: Self-Evolving Rubric Policy Optimization for Open-Ended Test-Time Reinforcement Learning Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T01:47:19.245261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T01:47:19.245261Z digest=sha256:f3dba59e688f1caab8e46826402e1161ebdfa3d4af438e03c7cd7d03232a4fab

Observation b1849c50-99e6-4357-995d-28168cc08f34 · inbound

TQLite: Multi-LLM Jury Guided Distillation for Real-time MQM Translation Quality Evaluation cites this paper.

TQLite: Multi-LLM Jury Guided Distillation for Real-time MQM Translation Quality Evaluation Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-08T04:27:33.980804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T04:27:33.980804Z digest=sha256:78dc136f411c7133de6429535bc74be30f6d9223adc3f18c1559352f15b7ad8c

Observation 6be1e988-f0e5-4429-bf44-975420dc87b3 · inbound

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution cites this paper.

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-08T10:20:42.039474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T10:20:42.039474Z digest=sha256:69c7455cbc42d8bb8bd7159969a40c3b1f6c9ace6612e8c5db70aaab4c11d4ce

Observation 27dcd247-5330-4510-9713-2c82e7f39459 · inbound

Apodex Discovery: Reality Benchmarks and Environments for Evaluating and Building Discoverative Artificial Intelligence cites this paper.

Apodex Discovery: Reality Benchmarks and Environments for Evaluating and Building Discoverative Artificial Intelligence Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T14:17:46.886273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:17:46.886273Z digest=sha256:2b140130f25bfa15a81cb99b66160fa51c489f9f001898130359632353e9778a

Observation 105ed384-b17f-4051-81df-84aedb235d2d · inbound

Requirements-Augmented Generation for Trustworthy Acceptance Testing of LLM-Based Software cites this paper.

Requirements-Augmented Generation for Trustworthy Acceptance Testing of LLM-Based Software Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T19:36:49.498691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:36:49.498691Z digest=sha256:c0d0355838e03495fb8c79130dff4e9b04173182ce2d652fc90da6f03af06ac6