{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2024:OXM3VLQ2HKNL3ED5IMJTJIXTLJ","short_pith_number":"pith:OXM3VLQ2","schema_version":"1.0","canonical_sha256":"75d9baae1a3a9abd907d431334a2f35a7e1ae34e470644d1936cc76c5a573109","source":{"kind":"arxiv","id":"2406.12624","version":6},"attestation_state":"computed","paper":{"title":"Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges","license":"http://creativecommons.org/publicdomain/zero/1.0/","headline":"","cross_cats":["cs.AI"],"primary_cat":"cs.CL","authors_text":"Aman Singh Thakur, Dieuwke Hupkes, Kartik Choudhary, Sankaran Vaidyanathan, Venkat Srinik Ramayapally","submitted_at":"2024-06-18T13:49:54Z","abstract_excerpt":"Offering a promising solution to the scalability challenges associated with human evaluation, the LLM-as-a-judge paradigm is rapidly gaining traction as an approach to evaluating large language models (LLMs). However, there are still many open questions about the strengths and weaknesses of this paradigm, and what potential biases it may hold. In this paper, we present a comprehensive study of the performance of various LLMs acting as judges, focusing on a clean scenario in which inter-human agreement is high. Investigating thirteen judge models of different model sizes and families, judging a"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2406.12624","kind":"arxiv","version":6},"metadata":{"license":"http://creativecommons.org/publicdomain/zero/1.0/","primary_cat":"cs.CL","submitted_at":"2024-06-18T13:49:54Z","cross_cats_sorted":["cs.AI"],"title_canon_sha256":"84dd766717e1c69f2702919e663d515f1b0b05d69b3aab4c660b0ca348d2e99e","abstract_canon_sha256":"998a61e7698b7970653d26b39cff7c2fc0e238244b9eaf0cb0dd9a72138d0fa0"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T11:54:40.224464Z","signature_b64":"FlDYoENNnRQnNkVl7dvjHes7JRy5HRO+++lm02zMb0lv8n+e3oA7M3HG/mtfzLxUdEQufMRsnWI72WLB/JPtBg==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"75d9baae1a3a9abd907d431334a2f35a7e1ae34e470644d1936cc76c5a573109","last_reissued_at":"2026-07-05T11:54:40.224004Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T11:54:40.224004Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges","license":"http://creativecommons.org/publicdomain/zero/1.0/","headline":"","cross_cats":["cs.AI"],"primary_cat":"cs.CL","authors_text":"Aman Singh Thakur, Dieuwke Hupkes, Kartik Choudhary, Sankaran Vaidyanathan, Venkat Srinik Ramayapally","submitted_at":"2024-06-18T13:49:54Z","abstract_excerpt":"Offering a promising solution to the scalability challenges associated with human evaluation, the LLM-as-a-judge paradigm is rapidly gaining traction as an approach to evaluating large language models (LLMs). However, there are still many open questions about the strengths and weaknesses of this paradigm, and what potential biases it may hold. In this paper, we present a comprehensive study of the performance of various LLMs acting as judges, focusing on a clean scenario in which inter-human agreement is high. Investigating thirteen judge models of different model sizes and families, judging a"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2406.12624","kind":"arxiv","version":6},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2406.12624/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2406.12624","created_at":"2026-07-05T11:54:40.224063+00:00"},{"alias_kind":"arxiv_version","alias_value":"2406.12624v6","created_at":"2026-07-05T11:54:40.224063+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2406.12624","created_at":"2026-07-05T11:54:40.224063+00:00"},{"alias_kind":"pith_short_12","alias_value":"OXM3VLQ2HKNL","created_at":"2026-07-05T11:54:40.224063+00:00"},{"alias_kind":"pith_short_16","alias_value":"OXM3VLQ2HKNL3ED5","created_at":"2026-07-05T11:54:40.224063+00:00"},{"alias_kind":"pith_short_8","alias_value":"OXM3VLQ2","created_at":"2026-07-05T11:54:40.224063+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":16,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2606.19714","citing_title":"AURA: Adaptive Uncertainty-aware Refinement for LLM-as-a-Judge Auditing","ref_index":35,"is_internal_anchor":false},{"citing_arxiv_id":"2606.02282","citing_title":"POIROT: Interrogating Agents for Failure Detection in Multi-Agent Systems","ref_index":21,"is_internal_anchor":false},{"citing_arxiv_id":"2605.28158","citing_title":"OR-Space: A Full-Lifecycle Workspace Benchmark for Industrial Optimization Agents","ref_index":34,"is_internal_anchor":false},{"citing_arxiv_id":"2510.00915","citing_title":"Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers","ref_index":28,"is_internal_anchor":false},{"citing_arxiv_id":"2408.15549","citing_title":"WildFeedback: Aligning LLMs With In-situ User Interactions And Feedback","ref_index":35,"is_internal_anchor":false},{"citing_arxiv_id":"2410.20791","citing_title":"From Cool Demos to Production-Ready FMware: Core Challenges and a Technology Roadmap","ref_index":98,"is_internal_anchor":false},{"citing_arxiv_id":"2411.15594","citing_title":"A Survey on LLM-as-a-Judge","ref_index":148,"is_internal_anchor":false},{"citing_arxiv_id":"2502.16942","citing_title":"NUTSHELL: A Dataset for Abstract Generation from Scientific Talks","ref_index":6,"is_internal_anchor":false},{"citing_arxiv_id":"2504.14044","citing_title":"Multi-Stage Retrieval for Operational Technology Cybersecurity Compliance Using Large Language Models: A Railway Casestudy","ref_index":43,"is_internal_anchor":false},{"citing_arxiv_id":"2601.21464","citing_title":"Conversation for Non-verifiable Learning: Self-Evolving LLMs through Meta-Evaluation","ref_index":20,"is_internal_anchor":false},{"citing_arxiv_id":"2603.23448","citing_title":"Code Review Agent Benchmark","ref_index":22,"is_internal_anchor":false},{"citing_arxiv_id":"2605.10639","citing_title":"Navigating the Sea of LLM Evaluation: Investigating Bias in Toxicity Benchmarks","ref_index":27,"is_internal_anchor":false},{"citing_arxiv_id":"2412.05579","citing_title":"LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods","ref_index":223,"is_internal_anchor":false},{"citing_arxiv_id":"2604.24158","citing_title":"Multi-Dimensional Evaluation of Sustainable City Trips with LLM-as-a-Judge and Human-in-the-Loop","ref_index":9,"is_internal_anchor":false},{"citing_arxiv_id":"2604.16706","citing_title":"Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench","ref_index":22,"is_internal_anchor":false},{"citing_arxiv_id":"2605.03147","citing_title":"Effective Performance Measurement: Challenges and Opportunities in KPI Extraction from Earnings Calls","ref_index":54,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/OXM3VLQ2HKNL3ED5IMJTJIXTLJ","json":"https://pith.science/pith/OXM3VLQ2HKNL3ED5IMJTJIXTLJ.json","graph_json":"https://pith.science/api/pith-number/OXM3VLQ2HKNL3ED5IMJTJIXTLJ/graph.json","events_json":"https://pith.science/api/pith-number/OXM3VLQ2HKNL3ED5IMJTJIXTLJ/events.json","paper":"https://pith.science/paper/OXM3VLQ2"},"agent_actions":{"view_html":"https://pith.science/pith/OXM3VLQ2HKNL3ED5IMJTJIXTLJ","download_json":"https://pith.science/pith/OXM3VLQ2HKNL3ED5IMJTJIXTLJ.json","view_paper":"https://pith.science/paper/OXM3VLQ2","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2406.12624&json=true","fetch_graph":"https://pith.science/api/pith-number/OXM3VLQ2HKNL3ED5IMJTJIXTLJ/graph.json","fetch_events":"https://pith.science/api/pith-number/OXM3VLQ2HKNL3ED5IMJTJIXTLJ/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/OXM3VLQ2HKNL3ED5IMJTJIXTLJ/action/timestamp_anchor","attest_storage":"https://pith.science/pith/OXM3VLQ2HKNL3ED5IMJTJIXTLJ/action/storage_attestation","attest_author":"https://pith.science/pith/OXM3VLQ2HKNL3ED5IMJTJIXTLJ/action/author_attestation","sign_citation":"https://pith.science/pith/OXM3VLQ2HKNL3ED5IMJTJIXTLJ/action/citation_signature","submit_replication":"https://pith.science/pith/OXM3VLQ2HKNL3ED5IMJTJIXTLJ/action/replication_record"}},"created_at":"2026-07-05T11:54:40.224063+00:00","updated_at":"2026-07-05T11:54:40.224063+00:00"}