Pith. sign in

Paper Citation Record · LEDGER

LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 52 inbound Pith citation observations for arXiv:2308.11462.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2308.11462 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 52 of 52 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:07:46.716010Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

28
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c77353f1-115c-4f9d-8422-8352267a617a · inbound

AIDE: Attribute-Guided MultI-Hop Data Expansion for Data Scarcity in Task-Specific Fine-tuning cites this paper.

AIDE: Attribute-Guided MultI-Hop Data Expansion for Data Scarcity in Task-Specific Fine-tuning LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T20:03:20.597044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:03:20.597044Z digest=sha256:21bff4b94cbfca39f38a795e84e1dc6f1a2b807907ed00e97c30f3570069301c

Observation 9578b812-3808-46ea-82e2-455156ae3030 · inbound

TrimLLM: Progressive Layer Dropping for Domain-Specific LLMs cites this paper.

TrimLLM: Progressive Layer Dropping for Domain-Specific LLMs LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T15:13:23.511117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:13:23.511117Z digest=sha256:fac88e4962e834116a523b74f014bc1bced7d759514ef44c297d1e9f49ac0224

Observation 139bd959-449b-45b8-9680-ded324374c43 · inbound

LegalGuardian: A Privacy-Preserving Framework for Secure Integration of Large Language Models in Legal Practice cites this paper.

LegalGuardian: A Privacy-Preserving Framework for Secure Integration of Large Language Models in Legal Practice LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T18:53:52.471004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:53:52.471004Z digest=sha256:d067278f0bfb20e9796abe1353c0e1c8933282d7dd1fb1c61c09b37c670eb8e3

Observation 8a430616-9e27-4c79-a038-68fdfff5f118 · inbound

FinanceQA: A Benchmark for Evaluating Financial Analysis Capabilities of Large Language Models cites this paper.

FinanceQA: A Benchmark for Evaluating Financial Analysis Capabilities of Large Language Models LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-10T00:53:29.674023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:53:29.674023Z digest=sha256:9a1b8335ae9b4566e6d246a469c631f5137e1be7df8652619db3554ac0f2c897

Observation a2a5231f-a917-4034-8073-9b3cf8d2e29f · inbound

AIMS.au: A Dataset for the Analysis of Modern Slavery Countermeasures in Corporate Statements cites this paper.

AIMS.au: A Dataset for the Analysis of Modern Slavery Countermeasures in Corporate Statements LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T14:04:30.177659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:04:30.177659Z digest=sha256:162c95328f644a2895aa5a936f351cf14d28a38e2f8fdeb6f9b637eedb7fe484

Observation 88b76c17-53f9-4a4b-b8e7-537b9807d8ea · inbound

Continual Pre-Training is (not) What You Need in Domain Adaption cites this paper.

Continual Pre-Training is (not) What You Need in Domain Adaption LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T12:07:46.716010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:07:46.716010Z digest=sha256:0b8bcb6688c2c78a381790e6add67aa8b5d54de165df0037722fb6282ba2ec2e

Observation 970a124b-f547-4234-9c47-272dfd98a0dd · inbound

Transformer-Based Extraction of Statutory Definitions from the U.S. Code cites this paper.

Transformer-Based Extraction of Statutory Definitions from the U.S. Code LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T11:09:43.467530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:09:43.467530Z digest=sha256:354e91e710159bb3cdeb6c9ea243af73cb8ef336c4eda04b6929714ccd44172b

Observation 067804ab-144d-4d0d-9762-7e1ba74c4a27 · inbound

Context Reasoner: Incentivizing Reasoning Capability for Contextualized Privacy and Safety Compliance via Reinforcement Learning cites this paper.

Context Reasoner: Incentivizing Reasoning Capability for Contextualized Privacy and Safety Compliance via Reinforcement Learning LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T15:37:26.528544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:37:26.528544Z digest=sha256:d89cba3aaa168f1fb77c514ab8ebc325a257cb94981361f117ceb6765f6401b3

Observation 4457ddbc-cf62-45a7-bc22-44e191c38f93 · inbound

Towards Large Reasoning Models for Agriculture cites this paper.

Towards Large Reasoning Models for Agriculture LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:21:43.049112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:21:43.049112Z digest=sha256:39218a67d7f7d8eb3f74a47d47dca55637f8cd9c61b0edf11030b3732724fcff

Observation 26293698-8d4e-4e0a-aad9-ac0f2333dea5 · inbound

AIMSCheck: Leveraging LLMs for AI-Assisted Review of Modern Slavery Statements Across Jurisdictions cites this paper.

AIMSCheck: Leveraging LLMs for AI-Assisted Review of Modern Slavery Statements Across Jurisdictions LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:23.083854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:23.083854Z digest=sha256:f839cd59d5cf6084aa32bae15631f9ec31ddfb11c9c91dc683381563696f2b2b

Observation a0539aad-19f8-411e-b7d9-c78ec157b0f0 · inbound

WisWheat: A Three-Tiered Vision-Language Dataset for Wheat Management cites this paper.

WisWheat: A Three-Tiered Vision-Language Dataset for Wheat Management LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:10.295568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:02:10.295568Z digest=sha256:373f3940ac4c84b09d510119cf5e3d717ecf5992fa03c1baaec7342ad87adc51

Observation 0c23ae7a-5ba5-4c5b-8fa2-fd50e0ddb26d · inbound

Enterprise Large Language Model Evaluation Benchmark cites this paper.

Enterprise Large Language Model Evaluation Benchmark LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T22:56:32.246086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:56:32.246086Z digest=sha256:f114056e6da22ff3bedc708f8851311c8a1f90df2247326d4be045852311b182

Observation 3eb6f863-281f-4072-b227-c83fca4980d6 · inbound

An Integrated Framework of Prompt Engineering and Multidimensional Knowledge Graphs for Legal Dispute Analysis cites this paper.

An Integrated Framework of Prompt Engineering and Multidimensional Knowledge Graphs for Legal Dispute Analysis LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.873373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:35:40.873373Z digest=sha256:324c0f52246151971524baf4f027aa2c93e3d4edd245a6a83c4d1ce0fbe997bd

Observation 8eb618d8-e412-41c4-9431-2ba7c5d36302 · inbound

All for law and law for all: Adaptive RAG Pipeline for Legal Research cites this paper.

All for law and law for all: Adaptive RAG Pipeline for Legal Research LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T17:21:00.808328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:21:00.808328Z digest=sha256:faebea01a253ba8b3e14b66a0670fcb81a5ee35929e4dc39697ac115f35ddbcc

Observation 9ce16b5a-141e-4626-be52-158ad33feedf · inbound

AI for Statutory Simplification: A Comprehensive State Legal Corpus and Labor Benchmark cites this paper.

AI for Statutory Simplification: A Comprehensive State Legal Corpus and Labor Benchmark LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T15:51:46.275259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:51:46.275259Z digest=sha256:901c35b035b030ff27a12266baa2ae4253c7f4ad39ba4a2fefc2fef424f34e31

Observation 0efc6177-c0c8-4f41-8fd8-c92df459e153 · inbound

L-MARS: Legal Multi-Agent System with Agentic Search and Citation-Faithfulness Audit cites this paper.

L-MARS: Legal Multi-Agent System with Agentic Search and Citation-Faithfulness Audit LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T13:20:40.339812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:20:40.339812Z digest=sha256:a9a666e8144a7c0b88cddd2dc4a926c9ed181e7913b4c6ca94605b7481f4544c

Observation 19fa69a8-2765-403f-8f13-90135d012582 · inbound

Nirvana: A Specialized Generalist Model With Task-Aware Memory Mechanism cites this paper.

Nirvana: A Specialized Generalist Model With Task-Aware Memory Mechanism LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T03:05:48.097934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-18T03:05:06.069642Z digest=sha256:4f23a5b5f70b64418ed2843d04286c9ffba46ec2afac20daa4f098ea26c793cc

Observation 11d79624-30eb-429a-9918-0e82de232510 · inbound

Towards Real-World Validity in Generative AI Benchmarks: Understanding and Designing Domain-Centered Evaluations for Journalism Practitioners cites this paper.

Towards Real-World Validity in Generative AI Benchmarks: Understanding and Designing Domain-Centered Evaluations for Journalism Practitioners LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T11:01:17.029897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-18T10:59:16.139525Z digest=sha256:13baef53901e682d215f692353fa701a4b526ccaf530a299b350c51090062ae9

Observation 2fb5a109-91ee-4c48-9816-733b21e74e4e · inbound

CapTrack: Multifaceted Evaluation of Forgetting in LLM Post-Training cites this paper.

CapTrack: Multifaceted Evaluation of Forgetting in LLM Post-Training LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-25T06:45:25.582582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-25T06:40:51.046965Z digest=sha256:de8b65880833864410d4ef50f14e8285a03bd85dc6108385261272c0e9a08990

Observation 0770fa1b-4680-4353-8f8d-7a5cd87c9fbc · inbound

SciPredict: Can LLMs Predict the Outcomes of Scientific Experiments in Natural Sciences? cites this paper.

SciPredict: Can LLMs Predict the Outcomes of Scientific Experiments in Natural Sciences? LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:36:03.821732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T15:55:34.768853Z digest=sha256:1aa5e3876f7d5279b89a3f4806658cd0e7983ca2160f4e242532546a5fb1421f

Observation e0c78989-4ba3-4b36-a240-7692ee3934dc · inbound

LegalBench-BR: A Benchmark for Evaluating Large Language Models on Brazilian Legal Decision Classification cites this paper.

LegalBench-BR: A Benchmark for Evaluating Large Language Models on Brazilian Legal Decision Classification LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T11:56:26.755806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T04:27:23.053818Z digest=sha256:bfb7d6c7161c8e0577bc8faecdcc22f745105964babae1ad142e800a9d1f3595

Observation e4704eac-14b1-4ed5-9f66-7916bd338bd6 · inbound

KnowPilot: Your Knowledge-Driven Copilot for Domain Tasks cites this paper.

KnowPilot: Your Knowledge-Driven Copilot for Domain Tasks LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:16:20.746297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T06:15:44.360621Z digest=sha256:a3157310e89b25b80876ef28f4b008bc934cda115cc2ba5b5c8f2470432833be

Observation 398924d6-1683-4640-a4fd-9614f707dddd · inbound

Learning When Not to Decide: A Framework for Overcoming Factual Presumptuousness in AI Adjudication cites this paper.

Learning When Not to Decide: A Framework for Overcoming Factual Presumptuousness in AI Adjudication LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T02:11:57.255230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T02:08:24.770003Z digest=sha256:4f4458b8ed0d1cf51b4b6db3874b8befa384b20f6a3bc6df4575b0e05071bfaf

Observation a7818266-5c70-4bb6-b044-e164685aefa9 · inbound

Exploiting LLM-as-a-Judge Disposition on Free Text Legal QA via Prompt Optimization cites this paper.

Exploiting LLM-as-a-Judge Disposition on Free Text Legal QA via Prompt Optimization LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T13:41:10.502069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T01:10:56.318664Z digest=sha256:7136fa9aac1fbc57cf53b6b5db6222e686da27819e49c34da912c010b159afbb

Observation cf967cc5-c82b-4272-8813-f631e1ac7063 · inbound

Breaking the Secret: Economic Interventions for Combating Collusion in Embodied Multi-Agent Systems cites this paper.

Breaking the Secret: Economic Interventions for Combating Collusion in Embodied Multi-Agent Systems LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T21:16:16.378379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-08T06:13:17.633146Z digest=sha256:72db9291bc0b13b372a5584559de8d4594d1b2916cdd9da27c06174ab226abd1

Observation 1fd275bc-c7b0-44e8-8671-6e4c428bb6cd · inbound

Expert Evaluation of LLM's Open-Ended Legal Reasoning on the Japanese Bar Exam Writing Task cites this paper.

Expert Evaluation of LLM's Open-Ended Legal Reasoning on the Japanese Bar Exam Writing Task LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T21:16:36.010805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-08T05:59:54.438189Z digest=sha256:a7d41cba8837309db00b8d5c986d28daa3771982da9bb801933919b3fb8aca8f

Observation ebcbdeb7-2d0d-4972-9bbf-2ed349ec7aad · inbound

Safety Drift After Fine-Tuning: Evidence from High-Stakes Domains cites this paper.

Safety Drift After Fine-Tuning: Evidence from High-Stakes Domains LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T23:11:19.337354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-07T17:53:57.169962Z digest=sha256:8866e532c815d1b6958408c9a075b18dfe0c9320bb1ea89ce0c8d00bc3edbe7a

Observation 11fe66d3-1f22-46f5-8cfc-2863a2f81a9c · inbound

Query-efficient model evaluation using cached responses cites this paper.

Query-efficient model evaluation using cached responses LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T23:35:07.682388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T23:28:47.530333Z digest=sha256:af7ce47dbc80d5d045ee3bb94395256d20b73bf5cd9cabfb822792188ec1846c

Observation 2d676823-a8f4-4cee-8652-9637cce9eb42 · inbound

Byte-Exact Deduplication in Retrieval-Augmented Generation: A Three-Regime Empirical Analysis Across Public Benchmarks cites this paper.

Byte-Exact Deduplication in Retrieval-Augmented Generation: A Three-Regime Empirical Analysis Across Public Benchmarks LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:26:24.208127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-12T04:18:11.836537Z digest=sha256:c227d246be681bf4378dde66a3f4f8a233191a938022511483d018937102137a

Observation f34661f7-0ff0-484e-953f-b59557e57208 · inbound

Decompose-and-Refine: Structured Legal Question Answering with Parametric Retrieval cites this paper.

Decompose-and-Refine: Structured Legal Question Answering with Parametric Retrieval LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T13:24:39.984446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T13:23:16.122654Z digest=sha256:a11d96029d68e2dc8ff2513de2991af61c587f0625a88ff55f9a9ee84fa6a3c1

Observation 0a20cf91-1c6c-49fb-80ab-640f03d4acf5 · inbound

TypedCSIP: Typed Counterfactual Pretraining for Chinese Legislative Conflict Classification cites this paper.

TypedCSIP: Typed Counterfactual Pretraining for Chinese Legislative Conflict Classification LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:14:00.218650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T22:06:28.247167Z digest=sha256:320daaca69f5597fc1992264dfce0ecb1a91df82d394042f5536bf42f9ea14a5

Observation 762cbd88-088c-4f89-a677-07649586ee0f · inbound

UA-Legal-Bench: A Benchmark for Evaluating Large Language Models on Ukrainian Legal Reasoning cites this paper.

UA-Legal-Bench: A Benchmark for Evaluating Large Language Models on Ukrainian Legal Reasoning LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T12:23:24.527606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T12:15:12.570661Z digest=sha256:0a6f375496388db6c16f2a9b8ee9dded387db6268682d2891e36f4ef14b9508f

Observation 89084092-98aa-40f3-ae4c-8eaade03f412 · inbound

Multi-Legal-Bench: Evaluating LLMs on Legal Reasoning Across Jurisdictions, Languages, and Legal Traditions cites this paper.

Multi-Legal-Bench: Evaluating LLMs on Legal Reasoning Across Jurisdictions, Languages, and Legal Traditions LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T08:33:15.935573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T07:24:08.269519Z digest=sha256:9599d0fb2fca9ff0f5bb8bc9748d5257cb16cb8b6553f175b4feb0c2ece324c5

Observation 0e17b22c-bbd3-4541-94e6-7a937abef3c6 · inbound

Citation Grounding Measures the Oracle: Graph Coverage Determines Reported LLM Hallucination Rates in Law cites this paper.

Citation Grounding Measures the Oracle: Graph Coverage Determines Reported LLM Hallucination Rates in Law LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T20:32:37.522670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-28T18:37:22.299236Z digest=sha256:bf37fc10d10a5070af33e6faf2640a20fb839f0485ee74547a5cb90fb3e09544

Observation c02298e6-473d-41c6-a231-e1f3028cde94 · inbound

GIScholarBench: Benchmarking LLM Overconfidence in GIS Research cites this paper.

GIScholarBench: Benchmarking LLM Overconfidence in GIS Research LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:07:25.908133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-27T19:17:26.825374Z digest=sha256:88305bbf15a5f332cd5d3f13acf4e9c4e73393c35faa59011a63d9b19ca0c24c

Observation 11bcab69-c3de-42e6-b065-a10de6e042fb · inbound

Trace2Policy: From Expert Behavior Traces to Self-Evolving Decision Agents cites this paper.

Trace2Policy: From Expert Behavior Traces to Self-Evolving Decision Agents LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T05:27:39.663246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-27T13:19:10.714343Z digest=sha256:54a7ac421d44e237ce809a92603f897956cde587019aa3afc2868bcb387764e9

Observation 39b03420-626a-435f-bb7b-9451693948dd · inbound

LegalHalluLens: Typed Hallucination Auditing and Calibrated Multi-Agent Debate for Trustworthy Legal AI cites this paper.

LegalHalluLens: Typed Hallucination Auditing and Calibrated Multi-Agent Debate for Trustworthy Legal AI LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-03T21:28:58.612291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-27T00:34:53.489694Z digest=sha256:eab2a1c22a7539a99a69a8be9bb03670b8353f6d7644f39489849d5d7f7feecb

Observation c454b269-996a-4f29-b290-fd070686884d · inbound

The Measurement Gap in the Automation of EU Law: Benchmarking Doctrinal Legal Reasoning under the EU AI Act cites this paper.

The Measurement Gap in the Automation of EU Law: Benchmarking Doctrinal Legal Reasoning under the EU AI Act LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T23:29:02.495106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-26T22:16:25.919667Z digest=sha256:6e0b656dc3cc8dca2de95c80383b308ff1d83c8431ea8ab091ebc51b64d5a645

Observation 5d0442ac-aa2b-455d-935b-99fef5a8c367 · inbound

Answer Engineering: Local Trajectory Editing for Protocol-Constrained Decision Making in Large Language Models cites this paper.

Answer Engineering: Local Trajectory Editing for Protocol-Constrained Decision Making in Large Language Models LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-07-04T06:49:37.576858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-26T14:13:13.678245Z digest=sha256:45fde6f63fedee202ee4e1a59225d0a60f246b5ea1b6bbc4b052248fb5638f5a

Observation 9675f5fb-93ac-48c5-bd49-3d74a374a686 · inbound

HAKARI-Bench: A Lightweight Benchmark for Comparing Retrieval Architectures and Efficiency Settings under Unified Conditions cites this paper.

HAKARI-Bench: A Lightweight Benchmark for Comparing Retrieval Architectures and Efficiency Settings under Unified Conditions LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 95

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T11:59:51.263031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-26T07:22:34.547816Z digest=sha256:d8833628f622fbdc14f02c374ba708d2698763863dc41fa69e6361b4957edba7

Observation 77a56eeb-1020-40e3-b02d-1e0bc4c61849 · inbound

Legal Reasoning Is Not Lawyering: Rethinking Legal Benchmarks for Pro Se Access to Justice cites this paper.

Legal Reasoning Is Not Lawyering: Rethinking Legal Benchmarks for Pro Se Access to Justice LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T23:19:03.985515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-26T22:26:54.239505Z digest=sha256:fd666e64a9d929332e5bc6b8d1d0753dfdd16d25b04806a213a147b2806b5048

Observation 58c35bc4-9e48-487d-bae7-3e457539c130 · inbound

How Do Tool-Augmented LLM Agents Perform on Real-World Energy Analytics Tasks? cites this paper.

How Do Tool-Augmented LLM Agents Perform on Real-World Energy Analytics Tasks? LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T15:29:56.715847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-26T01:34:10.103638Z digest=sha256:29ab0b85641d6e8cdfc58a0886eb3e2c9c3ce0301997bc34f54a4f0c2a78f3c5

Observation b2379fd8-9d26-4202-822d-e816f773e61c · inbound

Constraint-Aware Hierarchical Search for Regulation-Driven Fine-Grained Classification cites this paper.

Constraint-Aware Hierarchical Search for Regulation-Driven Fine-Grained Classification LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-14T10:38:03.549728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T10:38:03.549728Z digest=sha256:6e62c7eb4eb45f3301832c48ff2ba80049cc34c9dfc772681aa3a399eb181c3f

Observation 0f1c2e0f-2e20-40e7-897e-9605cfbc2a4f · inbound

BLAD: A Historically Contextualized, Multilingual Dataset of Bangladeshi Legal Acts (1799 to 2025) cites this paper.

BLAD: A Historically Contextualized, Multilingual Dataset of Bangladeshi Legal Acts (1799 to 2025) LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T19:01:02.854545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T19:01:02.854545Z digest=sha256:8a6a2635209198843dfe4440684e4604aebdc0585568b9e0e841cf8a7f3d7e75

Observation 7971769f-9742-4175-9f19-32ee4b0d2009 · inbound

AILQA: Evaluating AI-Driven Legal Question Answering Systems for the Indian Legal System cites this paper.

AILQA: Evaluating AI-Driven Legal Question Answering Systems for the Indian Legal System LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T14:17:23.403369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:17:23.403369Z digest=sha256:ac257519d2373b6a228f2c003e6f85443ed133cdba9491001203e7f19de9377d

Observation 0b3da14e-438a-45a5-b651-7b16132da10e · inbound

Toward Automated Detection of Documentation Inconsistencies in Electronic Health Records cites this paper.

Toward Automated Detection of Documentation Inconsistencies in Electronic Health Records LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T04:04:13.822723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:04:13.822723Z digest=sha256:8cacb22d4113fd44ebc1e1fea7489285e0429514f283835f44870e05c9a7dd2a

Observation 10731c04-647d-4086-9b51-aa77a094d1c9 · inbound

Confidently Wrong: Exception Chain Collapse in Frontier LLM Rule Evaluation cites this paper.

Confidently Wrong: Exception Chain Collapse in Frontier LLM Rule Evaluation LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 2010

Resolution
unresolved
no resolver link, observed 2026-07-30T23:41:20.109043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T23:41:20.109043Z digest=sha256:fd68743462e2875ef80a13a044a069abcad4d5b0938917941e09a1306b4564e6

Observation 1557386e-86b7-4762-b185-713397cfc408 · inbound

Reasoning Consensus: Structural Ensembling of LLM Reasoning via Weighted DAG Aggregation cites this paper.

Reasoning Consensus: Structural Ensembling of LLM Reasoning via Weighted DAG Aggregation LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-01T01:32:21.858955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:32:21.858955Z digest=sha256:50a291af589ab16924f15d05d4dcf42e96f7e2c199e38eb83ae2ce9e414122e3

Observation 082e6d06-398f-4600-a127-18d04123a176 · inbound

Question Begets Question: Self-Evolving Curriculum for Reinforcement Fine-Tuning on Competition Mathematics cites this paper.

Question Begets Question: Self-Evolving Curriculum for Reinforcement Fine-Tuning on Competition Mathematics LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T00:08:23.331668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:08:23.331668Z digest=sha256:36ea993a14e37e57c62a386292466ea8c3250d0bfdfa664bd17b86f53d787f8c

Observation e9d80172-d4ba-400c-a498-5e85697ba782 · inbound

When Does Trace-Driven Evaluation Mislead MoE Expert Caching? Replay Semantics, Workload Contamination, and Operating Regimes cites this paper.

When Does Trace-Driven Evaluation Mislead MoE Expert Caching? Replay Semantics, Workload Contamination, and Operating Regimes LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T00:48:38.348560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:48:38.348560Z digest=sha256:86ffe66136bf7f52ca069c71044ed181a1a6f036f5f9fb5b510fc1f7368d1294

Observation a249d4c8-fc89-4e53-94c6-30048010c096 · inbound

When Does Trace-Driven Evaluation Mislead MoE Expert Caching? Replay Semantics, Workload Contamination, and Operating Regimes cites this paper.

When Does Trace-Driven Evaluation Mislead MoE Expert Caching? Replay Semantics, Workload Contamination, and Operating Regimes LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-14T04:43:19.334168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:43:19.334168Z digest=sha256:2e9812f72e90bbe6d0555815291c3b406c9fc7efd9a5bc7e01bbf7945f80638b

Observation a5c890be-981e-49d4-b186-c5aec3ffd088 · inbound

V-FiLLM: Verified Financial LLM Reasoning Benchmark cites this paper.

V-FiLLM: Verified Financial LLM Reasoning Benchmark LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:14.710985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:27:14.710985Z digest=sha256:fa008a76ba321992118eef8506b441c34bc87c92fe588317579bbd5be012eed2