{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2023:T4IFBEK3STFEW26YYS6K76XUGN","short_pith_number":"pith:T4IFBEK3","schema_version":"1.0","canonical_sha256":"9f1050915b94ca4b6bd8c4bcaffaf4335f38003ffc6663982cf7cdb672cfeea6","source":{"kind":"arxiv","id":"2308.11462","version":1},"attestation_state":"computed","paper":{"title":"LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":["cs.AI","cs.CY"],"primary_cat":"cs.CL","authors_text":"Adam Chilton, Aditya Narayana, Alex Chohlas-Wood, Austin Peters, Brandon Waldon, Christopher R\\'e, Daniel E. Ho, Daniel N. Rockmore, Diego Zambrano, Dmitry Talisman, Enam Hoque, Faiz Surani, Frank Fagan, Galit Sarfaty, Gregory M. Dickinson, Haggai Porat, Jason Hegland, Jessica Wu, Joel Niklaus, Joe Nudell, John Nay, Jonathan H. Choi, Julian Nyarko, Kevin Tobia, Margaret Hagan, Megan Ma, Michael Livermore, Neel Guha, Nikon Rasumov-Rahe, Nils Holzenberger, Noam Kolt, Peter Henderson, Sean Rehaag, Shang Gao, Sharad Goel, Spencer Williams, Sunny Gandhi, Tom Zur, Varun Iyer, Zehua Li","submitted_at":"2023-08-20T22:08:03Z","abstract_excerpt":"The advent of large language models (LLMs) and their adoption by the legal community has given rise to the question: what types of legal reasoning can LLMs perform? To enable greater study of this question, we present LegalBench: a collaboratively constructed legal reasoning benchmark consisting of 162 tasks covering six different types of legal reasoning. LegalBench was built through an interdisciplinary process, in which we collected tasks designed and hand-crafted by legal professionals. Because these subject matter experts took a leading role in construction, tasks either measure legal rea"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2308.11462","kind":"arxiv","version":1},"metadata":{"license":"http://creativecommons.org/licenses/by/4.0/","primary_cat":"cs.CL","submitted_at":"2023-08-20T22:08:03Z","cross_cats_sorted":["cs.AI","cs.CY"],"title_canon_sha256":"26f787639fd0c04623ab0974256ce6dd404e9c4b0e313d5b503b816c5dc8ea24","abstract_canon_sha256":"fd227cb7dbeea12553b863694b1e1ae2e2eae09bbf1c1f9a1a9c2aa96e7ba66d"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T06:43:35.260299Z","signature_b64":"o6zvVEvEZxAXx+WtXtyLA6u1IEFFfwU30Xa+j/SK3dRropV+3x2xFtEtLSt5gkJawOXMhMZZCKSs0owdMenGCQ==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"9f1050915b94ca4b6bd8c4bcaffaf4335f38003ffc6663982cf7cdb672cfeea6","last_reissued_at":"2026-07-05T06:43:35.259816Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T06:43:35.259816Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":["cs.AI","cs.CY"],"primary_cat":"cs.CL","authors_text":"Adam Chilton, Aditya Narayana, Alex Chohlas-Wood, Austin Peters, Brandon Waldon, Christopher R\\'e, Daniel E. Ho, Daniel N. Rockmore, Diego Zambrano, Dmitry Talisman, Enam Hoque, Faiz Surani, Frank Fagan, Galit Sarfaty, Gregory M. Dickinson, Haggai Porat, Jason Hegland, Jessica Wu, Joel Niklaus, Joe Nudell, John Nay, Jonathan H. Choi, Julian Nyarko, Kevin Tobia, Margaret Hagan, Megan Ma, Michael Livermore, Neel Guha, Nikon Rasumov-Rahe, Nils Holzenberger, Noam Kolt, Peter Henderson, Sean Rehaag, Shang Gao, Sharad Goel, Spencer Williams, Sunny Gandhi, Tom Zur, Varun Iyer, Zehua Li","submitted_at":"2023-08-20T22:08:03Z","abstract_excerpt":"The advent of large language models (LLMs) and their adoption by the legal community has given rise to the question: what types of legal reasoning can LLMs perform? To enable greater study of this question, we present LegalBench: a collaboratively constructed legal reasoning benchmark consisting of 162 tasks covering six different types of legal reasoning. LegalBench was built through an interdisciplinary process, in which we collected tasks designed and hand-crafted by legal professionals. Because these subject matter experts took a leading role in construction, tasks either measure legal rea"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2308.11462","kind":"arxiv","version":1},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2308.11462/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2308.11462","created_at":"2026-07-05T06:43:35.259872+00:00"},{"alias_kind":"arxiv_version","alias_value":"2308.11462v1","created_at":"2026-07-05T06:43:35.259872+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2308.11462","created_at":"2026-07-05T06:43:35.259872+00:00"},{"alias_kind":"pith_short_12","alias_value":"T4IFBEK3STFE","created_at":"2026-07-05T06:43:35.259872+00:00"},{"alias_kind":"pith_short_16","alias_value":"T4IFBEK3STFEW26Y","created_at":"2026-07-05T06:43:35.259872+00:00"},{"alias_kind":"pith_short_8","alias_value":"T4IFBEK3","created_at":"2026-07-05T06:43:35.259872+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":26,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2606.26346","citing_title":"How Do Tool-Augmented LLM Agents Perform on Real-World Energy Analytics Tasks?","ref_index":8,"is_internal_anchor":false},{"citing_arxiv_id":"2606.22778","citing_title":"HAKARI-Bench: A Lightweight Benchmark for Comparing Retrieval Architectures and Efficiency Settings under Unified Conditions","ref_index":95,"is_internal_anchor":false},{"citing_arxiv_id":"2606.21121","citing_title":"Answer Engineering: Local Trajectory Editing for Protocol-Constrained Decision Making in Large Language Models","ref_index":44,"is_internal_anchor":false},{"citing_arxiv_id":"2606.18158","citing_title":"The Measurement Gap in the Automation of EU Law: Benchmarking Doctrinal Legal Reasoning under the EU AI Act","ref_index":45,"is_internal_anchor":false},{"citing_arxiv_id":"2606.23716","citing_title":"Legal Reasoning Is Not Lawyering: Rethinking Legal Benchmarks for Pro Se Access to Justice","ref_index":1,"is_internal_anchor":false},{"citing_arxiv_id":"2606.18021","citing_title":"LegalHalluLens: Typed Hallucination Auditing and Calibrated Multi-Agent Debate for Trustworthy Legal AI","ref_index":5,"is_internal_anchor":false},{"citing_arxiv_id":"2606.10457","citing_title":"Trace2Policy: From Expert Behavior Traces to Self-Evolving Decision Agents","ref_index":36,"is_internal_anchor":false},{"citing_arxiv_id":"2606.08036","citing_title":"GIScholarBench: Benchmarking LLM Overconfidence in GIS Research","ref_index":4,"is_internal_anchor":false},{"citing_arxiv_id":"2605.07096","citing_title":"Query-efficient model evaluation using cached responses","ref_index":10,"is_internal_anchor":false},{"citing_arxiv_id":"2605.24454","citing_title":"Decompose-and-Refine: Structured Legal Question Answering with Parametric Retrieval","ref_index":1,"is_internal_anchor":false},{"citing_arxiv_id":"2605.25474","citing_title":"TypedCSIP: Typed Counterfactual Pretraining for Chinese Legislative Conflict Classification","ref_index":6,"is_internal_anchor":false},{"citing_arxiv_id":"2605.29170","citing_title":"UA-Legal-Bench: A Benchmark for Evaluating Large Language Models on Ukrainian Legal Reasoning","ref_index":7,"is_internal_anchor":false},{"citing_arxiv_id":"2605.29738","citing_title":"Multi-Legal-Bench: Evaluating LLMs on Legal Reasoning Across Jurisdictions, Languages, and Legal Traditions","ref_index":7,"is_internal_anchor":false},{"citing_arxiv_id":"2606.00898","citing_title":"Citation Grounding: Detecting and Reducing LLM Citation Hallucinations via Legal Citation Graphs","ref_index":9,"is_internal_anchor":false},{"citing_arxiv_id":"2603.06610","citing_title":"CapTrack: Multifaceted Evaluation of Forgetting in LLM Post-Training","ref_index":15,"is_internal_anchor":false},{"citing_arxiv_id":"2511.05501","citing_title":"Towards Real-World Validity in Generative AI Benchmarks: Understanding and Designing Domain-Centered Evaluations for Journalism Practitioners","ref_index":25,"is_internal_anchor":false},{"citing_arxiv_id":"2510.26083","citing_title":"Nirvana: A Specialized Generalist Model With Task-Aware Memory Mechanism","ref_index":12,"is_internal_anchor":false},{"citing_arxiv_id":"2605.09611","citing_title":"Byte-Exact Deduplication in Retrieval-Augmented Generation: A Three-Regime Empirical Analysis Across Public Benchmarks","ref_index":27,"is_internal_anchor":false},{"citing_arxiv_id":"2604.24902","citing_title":"Safety Drift After Fine-Tuning: Evidence from High-Stakes Domains","ref_index":22,"is_internal_anchor":false},{"citing_arxiv_id":"2604.23730","citing_title":"Expert Evaluation of LLM's Open-Ended Legal Reasoning on the Japanese Bar Exam Writing Task","ref_index":7,"is_internal_anchor":false},{"citing_arxiv_id":"2604.23511","citing_title":"Breaking the Secret: Economic Interventions for Combating Collusion in Embodied Multi-Agent Systems","ref_index":42,"is_internal_anchor":false},{"citing_arxiv_id":"2604.20726","citing_title":"Exploiting LLM-as-a-Judge Disposition on Free Text Legal QA via Prompt Optimization","ref_index":6,"is_internal_anchor":false},{"citing_arxiv_id":"2604.18878","citing_title":"LegalBench-BR: A Benchmark for Evaluating Large Language Models on Brazilian Legal Decision Classification","ref_index":3,"is_internal_anchor":false},{"citing_arxiv_id":"2604.10718","citing_title":"SciPredict: Can LLMs Predict the Outcomes of Scientific Experiments in Natural Sciences?","ref_index":15,"is_internal_anchor":false},{"citing_arxiv_id":"2604.19820","citing_title":"KnowPilot: Your Knowledge-Driven Copilot for Domain Tasks","ref_index":5,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/T4IFBEK3STFEW26YYS6K76XUGN","json":"https://pith.science/pith/T4IFBEK3STFEW26YYS6K76XUGN.json","graph_json":"https://pith.science/api/pith-number/T4IFBEK3STFEW26YYS6K76XUGN/graph.json","events_json":"https://pith.science/api/pith-number/T4IFBEK3STFEW26YYS6K76XUGN/events.json","paper":"https://pith.science/paper/T4IFBEK3"},"agent_actions":{"view_html":"https://pith.science/pith/T4IFBEK3STFEW26YYS6K76XUGN","download_json":"https://pith.science/pith/T4IFBEK3STFEW26YYS6K76XUGN.json","view_paper":"https://pith.science/paper/T4IFBEK3","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2308.11462&json=true","fetch_graph":"https://pith.science/api/pith-number/T4IFBEK3STFEW26YYS6K76XUGN/graph.json","fetch_events":"https://pith.science/api/pith-number/T4IFBEK3STFEW26YYS6K76XUGN/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/T4IFBEK3STFEW26YYS6K76XUGN/action/timestamp_anchor","attest_storage":"https://pith.science/pith/T4IFBEK3STFEW26YYS6K76XUGN/action/storage_attestation","attest_author":"https://pith.science/pith/T4IFBEK3STFEW26YYS6K76XUGN/action/author_attestation","sign_citation":"https://pith.science/pith/T4IFBEK3STFEW26YYS6K76XUGN/action/citation_signature","submit_replication":"https://pith.science/pith/T4IFBEK3STFEW26YYS6K76XUGN/action/replication_record"}},"created_at":"2026-07-05T06:43:35.259872+00:00","updated_at":"2026-07-05T06:43:35.259872+00:00"}