Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 65 inbound Pith citation observations for arXiv:2405.01535.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T12:31:57.822140Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-11T00:47:43.114558Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation e798bd68-e6a6-4124-8cd8-31a2e27dc89b · inbound
From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4952cc09-d83e-42c5-8f0c-60f732d50cf1 · inbound
Self-Generated Critiques Boost Reward Modeling for Language Models Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab1e1182-0883-4f44-824b-3f94b995133e · inbound
Know Your RAG: Dataset Taxonomy and Generation Strategies for Evaluating RAG Systems Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 590c7677-cd90-4db9-bdf5-23bcfbdf10ff · inbound
LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
Reference 114
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation f3e320c6-e7ae-4fc5-b5b9-607cc2e7ea25 · inbound
Copyright-Protected Language Generation via Adaptive Model Fusion Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0784781-3e87-41b0-8bd2-37d74582ca1c · inbound
CoPrUS: Consistency Preserving Utterance Synthesis towards more realistic benchmark dialogues Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8349407d-6001-4c04-8785-882126d54c9b · inbound
Disentangling Preference Representation and Text Generation for Efficient Individual Preference Alignment Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c065bc01-aca6-499e-9fdd-dfc70ec4dbf6 · inbound
Influences on LLM Calibration: A Study of Response Agreement, Loss Functions, and Prompt Styles Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc6fb207-9096-4997-a5f1-528bfcb8ecfc · inbound
SedarEval: Automated Evaluation using Self-Adaptive Rubrics Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79e9433a-0aff-4a2c-b88b-8e9db9716d59 · inbound
Atla Selene Mini: A General Purpose Evaluation Model Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 683d7c4b-09d5-45cd-aac8-4830e0839131 · inbound
Can LLMs Replace Human Evaluators? An Empirical Study of LLM-as-a-Judge in Software Engineering Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5904f9a0-4020-4755-8ee1-a246be78e0f0 · inbound
Automatic Evaluation of Healthcare LLMs Beyond Question-Answering Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a5e4f0d-2c60-4e5b-a850-f6e3a454a01e · inbound
Salamandra Technical Report Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1909eda2-2b26-40df-8dff-e0c87713161f · inbound
Persona-judge: Personalized Alignment of Large Language Models via Token-level Self-judgment Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0e28339-24c1-4110-8269-34221fde98fb · inbound
PROMPTEVALS: A Dataset of Assertions and Guardrails for Custom Production Large Language Model Pipelines Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3598cd01-4af1-4f62-99a6-b6773f0a2feb · inbound
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f389968-28ad-466b-bb34-36db4df252fc · inbound
Trillion 7B Technical Report Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff496347-eb8e-41e4-b2b0-c892c5a0654b · inbound
Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a95a7f08-ce38-4ffc-bf63-5094d36e1b5a · inbound
Safety Degradation in AI Agents Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 512b44bf-1a37-4f4e-a85a-8f1394205af5 · inbound
YESciEval: Robust LLM-as-a-Judge for Scientific Question Answering Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c1e8b91-a8e6-435b-96cf-4d46f5630f26 · inbound
DECASTE: Unveiling Caste Stereotypes in Large Language Models through Multi-Dimensional Bias Analysis Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58ccd800-b8e4-404c-93cb-c281b0a29015 · inbound
Improving Fairness of Large Language Models in Multi-document Summarization Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50201047-13f5-4e3b-a1ed-20cad6b23555 · inbound
Outside Knowledge Conversational Video (OKCV) Dataset -- Dialoguing over Videos Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 584a9fc1-5e42-42ae-911f-6acb49934e14 · inbound
A Survey of Automatic Evaluation Methods on Text, Visual and Speech Generations Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
Reference 146
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation efa151a7-b8e1-4f0f-80bf-09e365e4def7 · inbound
Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d51b3700-0f16-4a5f-b047-83286e9cbf86 · inbound
MinosEval: Distinguishing Factoid and Non-Factoid for Tailored Open-Ended QA Evaluation with LLMs Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff9758d9-04ae-4b3e-9200-95eb543cfa8a · inbound
FairI Tales: Evaluation of Fairness in Indian Contexts with a Focus on Bias and Stereotypes Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1655dd3d-0ca5-4a26-92e2-4b204c58f7e4 · inbound
Hierarchical Memory Organization for Wikipedia Generation Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5c69d9e-df55-40ca-bd3d-93e7e01a6879 · inbound
Interpretable Mnemonic Generation for Kanji Learning via Expectation-Maximization Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fdcb2bc2-df0e-490e-af8c-b4179bcfc620 · inbound
Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 36a0b4e8-50ac-4506-929d-f69b460d6c98 · inbound
Multi-Modal Requirements Data-based Acceptance Criteria Generation using LLMs Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53653002-d5c0-4ce3-8d90-b996e49ae7ba · inbound
FHIR-RAG-MEDS: Integrating HL7 FHIR with Retrieval-Augmented Large Language Models for Enhanced Medical Decision Support Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d647bb2f-22f3-4969-abb0-3bb9932d4832 · inbound
RLBFF: Binary Flexible Feedback to bridge between Human Feedback & Verifiable Rewards Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 5b2270f2-4b7d-494f-997f-0e7fed373c85 · inbound
On the Shelf Life of Fine-Tuned LLM-Judges: Future-Proofing, Backward-Compatibility, and Question Generalization Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation d43b8ff0-58d4-4439-988b-4f5774d5b637 · inbound
Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation a2bd1b92-28a1-409f-8155-49f11e65098e · inbound
Prompt Optimization Is a Coin Flip: Diagnosing When It Helps in Compound AI Systems Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 5d883154-baf0-4768-a1f5-d9bc29f1ce91 · inbound
Beyond Verifiable Rewards: Rubric-Based GRM for Reinforced Fine-Tuning SWE Agents Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c7ba134f-680e-41bc-9b2f-bf76c8548b26 · inbound
KnowPilot: Your Knowledge-Driven Copilot for Domain Tasks Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation e028541d-1397-4718-bef4-e099a2e7c074 · inbound
Judging the Judges: A Systematic Evaluation of Bias Mitigation Strategies in LLM-as-a-Judge Pipelines Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 796152a9-b5b4-44b6-9ca0-b548b920c51e · inbound
HalluScan: A Systematic Benchmark for Detecting and Mitigating Hallucinations in Instruction-Following LLMs Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 4762d278-be0f-4bdf-9748-d04d9f55649b · inbound
AgentTrust: Runtime Safety Evaluation and Interception for AI Agent Tool Use Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 7c31b613-fbc1-462f-8110-d7ce0056ad2e · inbound
Rubric-Grounded RL: Structured Judge Rewards for Generalizable Reasoning Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation f47f4383-3c86-4a3a-b47e-78f3f7bcb3cd · inbound
TRACE: A taxonomy-grounded synthetic dataset for teaching-program generation and session interpretation in Applied Behavior Analysis Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 14bebf32-d519-45d4-b38e-77762109313d · inbound
DeepSurvey: Enhancing Analytical Depth and Citation Reliability in Automated Survey Generation Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 0c724aff-f1b6-4fe7-b24e-7e310d35c157 · inbound
CoEval: Ranking Language Models for Custom Tasks Without Labeled Data or Trustworthy Benchmarks Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 0f471fc0-fa12-414d-96d3-8c79df5dad8f · inbound
Organizational Control Layer: Governance Infrastructure at the Execution Boundary of LLM Agent Systems Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 611005d7-db0d-4c40-8716-3a1256b735de · inbound
Beyond Rubrics: Exploration-Guided Evaluation Skills for Reward Modeling Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 223b9bde-65a6-47db-a9bc-fc0ff814a6b7 · inbound
PoQ-Judge: A Multi-Architecture Evaluation Framework for Cost-Aware Proof-of-Quality in Decentralized LLM Inference Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 9fe2139e-23b4-4d39-86c4-6d1aa8555d4b · inbound
Quantifying and Auditing LLM Evaluation via Positive--Unlabeled Learning Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 00cfc5bf-e014-4d94-bb5c-8fabface31f9 · inbound
CourseBlueprint: A Structured Pipeline for Adaptive Pedagogical Video Generation Grounded in Course Corpora Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation b5da3bcb-dd04-41a4-94ed-70f9892a995f · inbound
Evaluation Awareness Is Not One Capability: Evidence from Open Language Models Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 3d8ef08f-a664-4ec4-9cfc-1e7fa7f9c790 · inbound
Open Problems in Constitutional Preference Reconstruction Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation f7fdd0ba-10c1-494e-b0cf-9c03b7fb6d45 · inbound
RoPoLL: Robust Panel of LLM Judges Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation bac534d8-fbc6-46c3-8931-61fe4213d9de · inbound
Healthier LLMs: Retrieval-Augmented Generation for Public Health Question Answering Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation b8e4f9ee-1810-4140-8569-9b7ccc42b505 · inbound
When the Judge Changes, So Does the Measurement: Auditing LLM-as-Judge Reliability Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation b90ccc66-222a-4d14-ade8-1e5bd0942a48 · inbound
Autoregressive Modeling of Film with Applications in Video Montage Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b2a61fe-472f-44f8-84cd-293b4924161f · inbound
BACON: Budgeted Human Calibration for Modeling and Evaluation with Multiple AI Judges Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da6dcea5-c533-4d7a-bd14-53e25c6e8103 · inbound
BACON: Budgeted Human Calibration for Modeling and Evaluation with Multiple AI Judges Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21b9f970-9a5d-43b6-932e-bd504a854b9a · inbound
Summary of DCASE 2026 Task 5: Audio-Dependent Question Answering Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db556d63-3042-4177-93f2-9f33928644b8 · inbound
SERPO: Self-Evolving Rubric Policy Optimization for Open-Ended Test-Time Reinforcement Learning Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a06887da-3a93-419c-980d-a3cfc432c04b · inbound
SERPO: Self-Evolving Rubric Policy Optimization for Open-Ended Test-Time Reinforcement Learning Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1849c50-99e6-4357-995d-28168cc08f34 · inbound
TQLite: Multi-LLM Jury Guided Distillation for Real-time MQM Translation Quality Evaluation Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6be1e988-f0e5-4429-bf44-975420dc87b3 · inbound
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27dcd247-5330-4510-9713-2c82e7f39459 · inbound
Apodex Discovery: Reality Benchmarks and Environments for Evaluating and Building Discoverative Artificial Intelligence Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 105ed384-b17f-4051-81df-84aedb235d2d · inbound
Requirements-Augmented Generation for Trustworthy Acceptance Testing of LLM-Based Software Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.