Pith. sign in

Paper Citation Record · LEDGER

LegalAgentBench: Evaluating LLM Agents in Legal Domain

As of 14 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 10 inbound Pith citation observations for arXiv:2412.17259.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.17259 v1

Coverage vector

measured 36 of 36 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T05:43:31.849186Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T19:55:14.057214Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-19T06:27:07.348511Z

Reference resolution

36 of 36 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved35
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3f32d2e7-5539-4033-8d72-be73b897a77a · outbound

This paper cites GPT-4 Technical Report.

LegalAgentBench: Evaluating LLM Agents in Legal Domain GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T05:43:30.900150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:43:30.900150Z digest=sha256:08e477dd0faeef72c9773d8ebdcebbfcb4425796580071324e8b3caca8182bdb

Observation 8e1f35d0-69a9-4887-85be-44c0a58b556b · outbound

This paper cites Qwen Technical Report.

LegalAgentBench: Evaluating LLM Agents in Legal Domain Qwen Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T05:43:30.967095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:43:30.967095Z digest=sha256:567ea18d2c9553378e4e96ae8147faafa6f71f7960908214b4588ca0fdfef854

Observation af5b9f70-e61b-4f1a-9a53-62ca0258cae4 · outbound

This paper cites an unresolved cited work.

LegalAgentBench: Evaluating LLM Agents in Legal Domain Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-11T05:43:32.649432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T05:43:30.970796Z digest=sha256:c2da827bf451d5d63acfbad1a8c6b481335ea67ca4cb89bad4f693b11f97ec96

Observation 1be704f5-7426-4fe1-8650-19990bcc37d4 · outbound

This paper cites PRE: A Peer Review Based Large Language Model Evaluator.

LegalAgentBench: Evaluating LLM Agents in Legal Domain PRE: A Peer Review Based Large Language Model Evaluator

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T05:43:30.974094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:43:30.974094Z digest=sha256:c9fb1ec55a6781c45f7130b959e362ae7f1946e3dbbf9f37d69bfad613db8e7a

Observation 9a9be135-ac96-415c-99b6-9c1283155d03 · outbound

This paper cites Chatlaw: A Multi-Agent Legal Assistant based on a Role-Aligned Mixture-of-Experts Architecture.

LegalAgentBench: Evaluating LLM Agents in Legal Domain Chatlaw: A Multi-Agent Legal Assistant based on a Role-Aligned Mixture-of-Experts Architecture

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T05:43:30.977617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:43:30.977617Z digest=sha256:d454d5bba3a1501538bf708b858801b8011b1a91447fd502a6265f21a37c0220

Observation 479e31bf-d38c-498e-a322-cde885df5bab · outbound

This paper cites an unresolved cited work.

LegalAgentBench: Evaluating LLM Agents in Legal Domain Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-11T05:43:32.559261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T05:43:30.981031Z digest=sha256:db329e3b5c9cb675c138058f876ddbd3a7c02bada182416df4b06d5d14fd229f

Observation 02c16e60-00ea-4586-b169-5780ed07ea6c · outbound

This paper cites ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools.

LegalAgentBench: Evaluating LLM Agents in Legal Domain ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T05:43:30.984072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:43:30.984072Z digest=sha256:c3df56f274f8465e452528bc5aaae8be26317cc29165afb9e1a61c0763cecfd5

Observation 7f35c2f7-b91c-4403-98ed-ae7e369eb270 · outbound

This paper cites an unresolved cited work.

LegalAgentBench: Evaluating LLM Agents in Legal Domain Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-11T05:43:32.549373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T05:43:30.987929Z digest=sha256:46c3ce06d560b21a85911ec6df2f826b4c4096dda70a71d4d3d7196968e98b92

Observation 4c483699-7693-414e-bf15-0b1d645a95c1 · outbound

This paper cites MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use.

LegalAgentBench: Evaluating LLM Agents in Legal Domain MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T05:43:30.989852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:43:30.989852Z digest=sha256:08de6e9fb83471bf11457f9b722091876317b6cc7e5faf54f3f0616bc83de3fe

Observation 613f0dff-f7c4-4095-bac5-7cab6bf7fdcc · outbound

This paper cites an unresolved cited work.

LegalAgentBench: Evaluating LLM Agents in Legal Domain Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T05:43:30.993794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:43:30.993794Z digest=sha256:fd72bac654965a45a7ff89249148e449fc64f47846ef3fb9d6b45c7eaed3a410

Observation aea06b71-c147-48bd-b617-ce29827cde7a · outbound

This paper cites BLADE: Enhancing Black-box Large Language Models with Small Domain-Specific Models.

LegalAgentBench: Evaluating LLM Agents in Legal Domain BLADE: Enhancing Black-box Large Language Models with Small Domain-Specific Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T05:43:30.995851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:43:30.995851Z digest=sha256:007c58c1ae836cbad373b309f0c56bf1a6ebe3f6205ddf6573f01a7a50d7f7fc

Observation 801376c8-ac76-41ff-af7c-90b8dfd529bb · outbound

This paper cites DELTA: Pre-train a Discriminative Encoder for Legal Case Retrieval via Structural Word Alignment.

LegalAgentBench: Evaluating LLM Agents in Legal Domain DELTA: Pre-train a Discriminative Encoder for Legal Case Retrieval via Structural Word Alignment

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-08-11T05:43:32.144353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T05:43:31.116045Z digest=sha256:584aa6982e81297eea1d3dd2c7e377f91a0a42f54c53d98de56a925ef08920ed

Observation f1c7cc3d-055b-470a-bedb-57d5d2dd1f0a · outbound

This paper cites CalibraEval: Calibrating Prediction Distribution to Mitigate Selection Bias in LLMs-as-Judges.

LegalAgentBench: Evaluating LLM Agents in Legal Domain CalibraEval: Calibrating Prediction Distribution to Mitigate Selection Bias in LLMs-as-Judges

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T05:43:31.225208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:43:31.225208Z digest=sha256:856e5147a15e6e93c5b682c290c9c609cdc8a602ad0d1ad1d3685c525c0c12da

Observation 9c5c66cb-cc01-46ea-9e5b-6bd0c05b44ee · outbound

This paper cites LexEval: A Comprehensive Chinese Legal Benchmark for Evaluating Large Language Models.

LegalAgentBench: Evaluating LLM Agents in Legal Domain LexEval: A Comprehensive Chinese Legal Benchmark for Evaluating Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T05:43:31.285636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:43:31.285636Z digest=sha256:3d1a686d8792cfa551ffd09e3a228a1d82da504921cf3582cb7e3d9cf7f90cf6

Observation 0b9ce6c4-d32f-4e44-a464-d3e5d793ebd3 · outbound

This paper cites LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods.

LegalAgentBench: Evaluating LLM Agents in Legal Domain LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T05:43:31.292630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:43:31.292630Z digest=sha256:56b5a93713c6fee0c4e4d81ca7b0fcb8f716b905d3ffff08129b895794397405

Observation 07961047-b6ca-474b-91c7-3e61428e48cd · outbound

This paper cites an unresolved cited work.

LegalAgentBench: Evaluating LLM Agents in Legal Domain Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-11T05:43:32.537805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T05:43:31.298664Z digest=sha256:c4e5269953ff027cc9bd72a5102e0c181f1b933031d6fa4fbaafb6d39b476acd

Observation c27ae8dd-0cf6-4b0e-98fb-4e7acd8f05f2 · outbound

This paper cites LLM+P: Empowering Large Language Models with Optimal Planning Proficiency.

LegalAgentBench: Evaluating LLM Agents in Legal Domain LLM+P: Empowering Large Language Models with Optimal Planning Proficiency

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T05:43:31.302040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:43:31.302040Z digest=sha256:e2ca58d06caf60607d25ad42d3b8b90972730df1c4c1480561b4c7650d4eb8ce

Observation 795af71d-aaa9-45b7-9752-8ac5f06ebe76 · outbound

This paper cites AgentBench: Evaluating LLMs as Agents.

LegalAgentBench: Evaluating LLM Agents in Legal Domain AgentBench: Evaluating LLMs as Agents

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T05:43:31.305999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:43:31.305999Z digest=sha256:d29866bae6df639531a7a3d17245457fe311e026eaa4231ef9f3788374abfd32

Observation 6864b6b7-ed6e-4ab2-a60a-c29b73ed629e · outbound

This paper cites AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents.

LegalAgentBench: Evaluating LLM Agents in Legal Domain AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T05:43:31.309052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:43:31.309052Z digest=sha256:2c5951591eaaf35d6af8c1edf1f5c35dcad61bf461003e5ab2651a84d7850f41

Observation d4ab7364-f1db-4d79-b34c-0c267e009bd2 · outbound

This paper cites ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs.

LegalAgentBench: Evaluating LLM Agents in Legal Domain ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T05:43:31.313784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:43:31.313784Z digest=sha256:27836d94d6a86f80a2ff985d0f4721c48b7aa839359f2cae9efc8b7b609716e1

Observation e48f8d9f-ff08-48c1-8db4-fa30368b36ca · outbound

This paper cites an unresolved cited work.

LegalAgentBench: Evaluating LLM Agents in Legal Domain Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T05:43:31.317281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:43:31.317281Z digest=sha256:67265934d7c1dfa3adb170052ca6f18c236a1d505eae17a84dc06c7c0d08b92c

Observation 34e450d2-43ad-483e-953c-ea7761db7115 · outbound

This paper cites an unresolved cited work.

LegalAgentBench: Evaluating LLM Agents in Legal Domain Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T05:43:31.365645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:43:31.365645Z digest=sha256:06b82ed21bf6e7e123e36211968ab14ca3c4bb6869b9f2e59caf6a0f65bf5f6a

Observation ce4df773-884a-4634-90be-360e509a2c28 · outbound

This paper cites an unresolved cited work.

LegalAgentBench: Evaluating LLM Agents in Legal Domain Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-11T05:43:32.384796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T05:43:31.431729Z digest=sha256:56943e95e8568dbe7b63472faa3dec3dcf3229025433c31639cf20d38001f5dd

Observation 1d827b88-f0fb-45cb-893f-d93300fbd135 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

LegalAgentBench: Evaluating LLM Agents in Legal Domain LLaMA: Open and Efficient Foundation Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T05:43:31.503942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:43:31.503942Z digest=sha256:f32b0dfb694b1dabc4079cec4613dafb41a6d3641973fd40d18201499bd1c881

Observation efb23076-f80c-40cf-9209-d6dff8d6bbe7 · outbound

This paper cites an unresolved cited work.

LegalAgentBench: Evaluating LLM Agents in Legal Domain Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T05:43:31.507757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:43:31.507757Z digest=sha256:c166a89bd8b6f4af303a8a11342669dc93068c340bd8348c6d4922c357398ddb

Observation f9ff3bc8-7345-41c1-a60e-0be86a830765 · outbound

This paper cites Plan-and-Solve Prompting: Improving Zero-Shot Chain-of-Thought Reasoning by Large Language Models.

LegalAgentBench: Evaluating LLM Agents in Legal Domain Plan-and-Solve Prompting: Improving Zero-Shot Chain-of-Thought Reasoning by Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T05:43:31.510196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:43:31.510196Z digest=sha256:5ff2ce791a22444314bafc09d0942b7ba9505fecc5df169d20c4cf2c7d84b376

Observation ebecddb3-945d-45e6-a4b1-ea7d571f671c · outbound

This paper cites an unresolved cited work.

LegalAgentBench: Evaluating LLM Agents in Legal Domain Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T05:43:31.513300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:43:31.513300Z digest=sha256:1b523c83d2c7559ba7374e0e734a8297f838defd4a3ff4b2c82bd3241a95df6f

Observation a861edd9-5a31-4791-9722-fbb4a400705c · outbound

This paper cites an unresolved cited work.

LegalAgentBench: Evaluating LLM Agents in Legal Domain Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T05:43:31.517129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:43:31.517129Z digest=sha256:9a4084c3321a5e446d8e3c119d1eee4c4b747e2b3997111853c12e5f5e59a253

Observation ae7bd030-b13d-4cff-a7df-ad3f6b82e1df · outbound

This paper cites an unresolved cited work.

LegalAgentBench: Evaluating LLM Agents in Legal Domain Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T05:43:31.520911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:43:31.520911Z digest=sha256:2d1a54f56a4d1f041166558e536529330954cbd06d0872a034c04a9f21da84ad

Observation bcac62ce-7ac9-44f9-b891-19b79a0f44df · outbound

This paper cites React: Synergizing reasoning and acting in language models.

LegalAgentBench: Evaluating LLM Agents in Legal Domain React: Synergizing reasoning and acting in language models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T05:43:31.584743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:43:31.584743Z digest=sha256:1528fa53575aa5b9122079d3fff3370c5d0b0a6ffdab35e7ee9f99bcca2d6b9c

Observation 94c24ae8-097f-4274-8bef-f936f6919e11 · outbound

This paper cites ToolEyes: Fine-Grained Evaluation for Tool Learning Capabilities of Large Language Models in Real-world Scenarios.

LegalAgentBench: Evaluating LLM Agents in Legal Domain ToolEyes: Fine-Grained Evaluation for Tool Learning Capabilities of Large Language Models in Real-world Scenarios

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T05:43:31.689503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:43:31.689503Z digest=sha256:e30e0af331c84d1b8d42019aafb7f600cf8c2dbf84744a17c0df942877938fa4

Observation 079de58c-a35c-4ee8-acdc-7c581a8abcee · outbound

This paper cites BERTScore: Evaluating Text Generation with BERT.

LegalAgentBench: Evaluating LLM Agents in Legal Domain BERTScore: Evaluating Text Generation with BERT

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T05:43:31.807563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:43:31.807563Z digest=sha256:fbe45eb1bf60a81359beba1fda9f1d42db9debad5050a27c72519340f153be45

Observation ce05e922-7020-44a4-9fba-b29db2eec050 · outbound

This paper cites A Survey of Large Language Models.

LegalAgentBench: Evaluating LLM Agents in Legal Domain A Survey of Large Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T05:43:31.839524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:43:31.839524Z digest=sha256:96b9c834dfe36f796141cf2908bea27a65c9be0bd70b7ec150116ea02208ad27

Observation fc687d77-ce4e-437e-9d75-94ca7cb0b8d7 · outbound

This paper cites an unresolved cited work.

LegalAgentBench: Evaluating LLM Agents in Legal Domain Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T05:43:31.842830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:43:31.842830Z digest=sha256:a22ac2c60ee1cf88db10b6c58583fbe4ba09d7dc184b966aa9109dbca08225aa

Observation 2205a3ba-5b01-4343-8828-c4f9c563aa93 · outbound

This paper cites URL: " 'urlintro :=.

LegalAgentBench: Evaluating LLM Agents in Legal Domain URL: " 'urlintro :=

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T05:43:31.845329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:43:31.845329Z digest=sha256:92253d08253d86f1748dea0909bd0ecf5ea784624e75925ab2625860bf664f52

Observation 1d5d1101-085a-4af1-a894-21128cb7e3df · outbound

This paper cites write newline.

LegalAgentBench: Evaluating LLM Agents in Legal Domain write newline

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T05:43:31.849186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:43:31.849186Z digest=sha256:dabeba43e21616728a3180b8e5da31cfcb41f8dd2e139b7a8084f06b125eefdf

Pith citing papers

Observation ca4500ef-cab4-4921-879b-208520ab91a9 · inbound

Evaluating LLM-based Approaches to Legal Citation Prediction: Domain-specific Pre-training, Fine-tuning, or RAG? A Benchmark and an Australian Law Case Study cites this paper.

Evaluating LLM-based Approaches to Legal Citation Prediction: Domain-specific Pre-training, Fine-tuning, or RAG? A Benchmark and an Australian Law Case Study LegalAgentBench: Evaluating LLM Agents in Legal Domain

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T19:55:14.057214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:55:14.057214Z digest=sha256:b572f6750081e292878cb752daa797b59cc3ca68d1e2e74835172f679f5d41e0

Observation 1359ade1-72b1-4b74-8a0f-88af0a8f3512 · inbound

From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review cites this paper.

From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review LegalAgentBench: Evaluating LLM Agents in Legal Domain

Reference 105

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:57:38.422719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T02:57:37.873567Z digest=sha256:eb6559650d816fed84fe801abef4affca9a66da13eb77fcde5fd33fcd526be10

Observation 01807099-6678-49ed-a9f4-09381aab52d4 · inbound

AppealCase: A Dataset and Benchmark for Civil Case Appeal Scenarios cites this paper.

AppealCase: A Dataset and Benchmark for Civil Case Appeal Scenarios LegalAgentBench: Evaluating LLM Agents in Legal Domain

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:45.264165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:45.264165Z digest=sha256:197dbffc2c59401d400815591db96ad182df92653fcd10bb4f1ad592d3e0f927

Observation 6b5fb4e9-c1fb-4fef-982b-9a5323918bb6 · inbound

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation cites this paper.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation LegalAgentBench: Evaluating LLM Agents in Legal Domain

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:12.933539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:12.933539Z digest=sha256:f6ce0409798eee894f980fa329c9ba850bf009c1e2c8139559a62f69963c90bc

Observation 2df0fa21-1272-4e55-a493-b1785495c368 · inbound

ChemAU: Harness the Reasoning of LLMs in Chemical Research with Adaptive Uncertainty Estimation cites this paper.

ChemAU: Harness the Reasoning of LLMs in Chemical Research with Adaptive Uncertainty Estimation LegalAgentBench: Evaluating LLM Agents in Legal Domain

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:44.927958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:44.927958Z digest=sha256:6ac598f1f351b76fedf008a5d72bf35793441dc67c7162014e6befed84ba7fb3

Observation c90dfaf9-7314-43e4-902b-b16d313fe0be · inbound

Collaborative Editable Model cites this paper.

Collaborative Editable Model LegalAgentBench: Evaluating LLM Agents in Legal Domain

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T00:24:12.928087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:24:12.928087Z digest=sha256:4f125764cfa8b69c49a0965eec057c8acfa63320ef5d45163e59268c0f921849

Observation 2796e88b-1042-4170-a438-d07bc86c1fa6 · inbound

Do LLMs Overthink Basic Math Reasoning? Benchmarking the Accuracy-Efficiency Tradeoff in Language Models cites this paper.

Do LLMs Overthink Basic Math Reasoning? Benchmarking the Accuracy-Efficiency Tradeoff in Language Models LegalAgentBench: Evaluating LLM Agents in Legal Domain

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-19T06:27:07.350996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-19T06:25:12.799097Z digest=sha256:0ce2387a1b2c2dc27d161950c27ff6d241733ea3b86954a44c7303d7d7feb0a8

Observation 45614148-3f14-4d28-ac2e-48f17049ec52 · inbound

LaQual: An Automated Framework for LLM App Quality Evaluation cites this paper.

LaQual: An Automated Framework for LLM App Quality Evaluation LegalAgentBench: Evaluating LLM Agents in Legal Domain

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T16:25:19.244835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:25:19.244835Z digest=sha256:d43ec3238f20da1123be44a1a3c2c099d16ff47640cd9b69069020218cd6a88b

Observation 54ab8962-1782-4e6a-abc5-c8400e158b9a · inbound

KoBLEX: Open Legal Question Answering with Multi-hop Reasoning cites this paper.

KoBLEX: Open Legal Question Answering with Multi-hop Reasoning LegalAgentBench: Evaluating LLM Agents in Legal Domain

Reference 9474

Resolution
unresolved
no resolver link, observed 2026-08-05T12:46:10.566030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:46:10.566030Z digest=sha256:214241d11dbe78dbb2a19de7ca16df44d8594cf0f82c456ee2e94be73746cec1

Observation 9530e74a-7c18-45c6-abf6-8bb4359c6420 · inbound

ADAM: A Systematic Data Extraction Attack on Agent Memory via Adaptive Querying cites this paper.

ADAM: A Systematic Data Extraction Attack on Agent Memory via Adaptive Querying LegalAgentBench: Evaluating LLM Agents in Legal Domain

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:45:57.687924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T17:57:59.918483Z digest=sha256:b51bccd8487f649228a536f8e27f5bf0064f566b3340a7abc1640c5d4a8f9f54