REVIEW 3 major objections 6 minor 1 cited by
A Reasoning-Focused Legal Retrieval Benchmark
T0 review · 3 major / 6 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read This paper claims that answering legal questions often requires retrieving passages with almost no lexical overlap, that standard retrievers fail on two new benchmarks built to capture this, and that query expansion with generated legal…
desk verdict Useful new legal IR benchmarks with real effort, but the single-gold-label scoring makes the headline difficulty numbers hard to interpret; the paper's own text admits the problem. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the pair of benchmarks themselves, constructed through annotation processes modeled on legal research: for Bar Exam QA, a law student searched case law with Boolean search queries and selected a passage stating the rule that justifies the answer; for Housing Statute QA, question-answer pairs from the Eviction Laws Database were reformatted into yes/no questions and mapped to supporting statutes in a roughly two-million-passage corpus. The property that carries the argument is low lexical overlap between queries and gold passages, measured by TF-IDF cosine similarity and confirmed by the failure of lexical retrieval. The mechanism that recovers performance is structured legal reasoning query expansion: a prompted generative model produces a rollout that names the legal issue and states the applicable rule, and that rollout is appended to the query before retrieval. The expansion supplies the vocabulary of legal doctrine that the original facts omit, effectively translating the query into the language of the passage corpus.
What would settle it
Take a random sample of, say, 100 queries from Housing Statute QA and 100 from Bar Exam QA, ask two independent lawyers to judge whether each gold passage supports the stated answer (for Housing Statute QA, whether the cited statute actually answers the derived yes/no question), and measure agreement and the rate of unsupported labels; if a substantial fraction of gold labels fail this check, the benchmarks' reported difficulty and retriever rankings are not trustworthy as stated.
Extended reading notes
Core claim
Bar Exam QA and Housing Statute QA are retrieval tasks whose query-gold passage pairs have very low lexical similarity (mean TF-IDF cosine similarity of 0.07 and 0.09, compared with 0.25–0.27 for Natural Questions, HotpotQA, COLIEE, and CLERC). On these tasks, standard retrievers perform poorly: BM25 and E5-large-v2 reach only 5.03 and 7.00 Recall@10 on Bar Exam QA, versus 40.4 and 68.7 on Natural Questions, and Housing Statute QA shows the same pattern under both upper- and lower-bound recall definitions. The paper further claims that a law-inspired query expansion method, which prompts a generative model to identify the legal issue in the query and generate the governing legal rule before retrieval, produces statistically significant Recall@10 gains on both tasks (for example, +6.28 for BM25 and +8.86 for E5-large-v2 on Bar Exam QA), with the largest gains on lexical and smaller dense models. Downstream question answering with retrieved passages improves less dramatically, and the paper reports that the gains are capped by how well the reader model can apply even the gold passage.
Load-bearing premise
The load-bearing premise is that the hand-assigned gold passages really do support the answer to each query; if many labels are wrong or only partially relevant, the low lexical-overlap measurements and the retriever rankings built on them would shift.
Editorial extensions
If this is right
- On low lexical-overlap legal tasks, BM25 and smaller dense retrievers are not a reliable substitute for reasoning; their Recall@10 sits in the single digits on Bar Exam QA.
- Generative query expansion that encodes legal issue spotting and rule statements improves retrieval for lexical and small dense models, with gains concentrated on exactly the hard, low-overlap examples.
- On high-similarity benchmarks such as Natural Questions and HotpotQA, the same reasoning rollouts add little or nothing, so the method's value is tied to the reasoning intensity of the task.
- Evaluating end-to-end retrieval-augmented generation on these benchmarks requires a strong reader: a weak reader gains only about 20 percent from having the gold passage, so retrieval gains translate weakly into answer accuracy.
- The released benchmark splits provide a reusable testbed of roughly ten thousand query-passage-answer triples spanning bar-exam doctrine and multi-jurisdiction housing statutes.
Reading between the lines
- A testable extension the authors do not run: evaluating learned sparse and late-interaction retrievers on these benchmarks would show whether the low-overlap gap is specific to BM25 and bi-encoder dense models or generalizes across the architecture space.
- Because the paper ties retrieval difficulty to query-gold lexical similarity, it implies that legal retrieval datasets assembled from citing-context and cited-case pairs will understate the reasoning required in real lawyer queries; future benchmarks could measure reasoning load directly at the level of legal concepts rather than surface words.
- The finding that reasoning rollouts help most when lexical overlap is low suggests a cheap recipe for production legal retrieval-augmented systems: before retrieval, have a language model name the legal issue and candidate rules when the query is fact-heavy and doctrine-poor, but skip the expansion on simple lookups where it adds noise.
- Since the Housing Statute QA transformation can make some gold citations irrelevant to the derived yes/no question, the reported upper and lower recall bounds bracket the true difficulty; re-judging gold passages against the derived questions would likely widen the gap between simple retrievers and reasoning-augmented ones.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces two new legal retrieval and retrieval-augmented QA benchmarks: Bar Exam QA, built from multistate bar exam hypotheticals paired with single gold passage annotations, and Housing Statute QA, derived from the LSC Eviction Laws Database with statutory citations as gold passages. The authors report that these datasets have substantially lower lexical (TF-IDF) similarity between queries and gold passages than Natural Questions, HotpotQA, COLIEE, and CLERC (Table 2). They evaluate BM25 and E5-family dense retrievers and find low Recall@10 on Bar Exam QA (5.03 for BM25, 7.00 for E5-large-v2, Table 3) compared with those baselines on general-domain tasks. They also propose a 'structured legal reasoning' query expansion method using GPT-3.5, report that it significantly improves Recall@10 on the new benchmarks (e.g., +8.86 +/- 1.16 for E5-large-v2 on Bar Exam QA, Figure 3), and include downstream QA experiments with retrieved and gold passages.
Significance. If the benchmarks behave as claimed, they fill a real gap: existing legal IR datasets are largely built from citation graphs with high query-document lexical overlap, and few provide paired downstream QA. The paper contributes two publicly released datasets with ~10K labeled examples, reasonably large corpora (856K and 1.8M passages), confidence intervals and statistical tests for the lexical-similarity differences, and an honest upper/lower-bound treatment of the Housing Statute QA labels. The structured query expansion results are also interesting as a practical finding. The main risk to significance is that the headline 'reasoning-intensive' interpretation rests on incomplete relevance judgments; if the gold labels are sparse, the low recall numbers could reflect annotation sparsity rather than retrieval difficulty.
major comments (3)
- [§3.1, Appendix A, Table 3] The Bar Exam QA evaluation uses a single gold passage per query, selected by one law student as a 'succinct, generalizable statement of the rule' (Appendix A). Given the paper's own observation in §2.2 that 'common principles or rules are restated many times across the corpus,' this single-label scoring conflates incomplete relevance judgments with retrieval failure: a retriever that returns a different passage stating the same rule is scored as wrong. The headline Recall@10 values (5.03 and 7.00 in Table 3) and the low lexical-similarity figures in Table 2 may therefore reflect annotation sparsity rather than the reasoning difficulty of the task. I recommend reporting inter-annotator agreement, expanding the gold set for at least a random sample of queries (e.g., by pooling additional annotator judgments or accepting multiple rule restatements), and showing whether the observed retriever rankings are stable under an expanded relevance set.
- [§3.2, §5.2, Tables 24-25 and Appendix H] For Housing Statute QA, the Y/N transformation of the original LSC questions can make some inherited statutes irrelevant to the derived question, while other relevant statutes may be uncited. The paper acknowledges this in §5.2 and reports upper-bound recall (at least one gold) and lower-bound recall (all golds), but neither quantity measures which gold statutes actually support the Y/N answer. Additionally, the retrieval corpus is built from Justia's 2021 statutes, and §3.2 notes that 'Justia's coverage of state law is incomplete'; the paper does not quantify how many gold citations are absent from the corpus, which would cap achievable recall below 1. These measurement gaps mean the difficulty estimates for Housing Statute QA and the relative ordering of retrievers could change with corrected or completed labels. I suggest reporting, for a validation subset, the fraction of gold statutes that are present in the corpus and that genuinely answer the derived Y/N question, and re-estimating recall under that subset.
- [§4, §6] The conclusion that these tasks 'require a greater degree of reasoning' (§4) and that 'retrievers must themselves be reasoners too' (§7) is inferred indirectly from low lexical similarity and from the gains of reasoning-oriented query expansion. Low lexical overlap can arise from domain-specific vocabulary, legal drafting style, or incomplete labels, not only from inferential reasoning; and the query expansion results are consistent with a vocabulary-mismatch explanation as much as with a reasoning explanation. A concrete test would be to include a control condition with similarly low lexical similarity but minimal reasoning demand (e.g., synonym-substituted or paraphrased queries), or to measure human retrieval performance and oracle retrieval performance on the same queries, and to show that the retrieval gains exceed what vocabulary expansion alone would yield. Without such a check, the 'reasoning-focused' characterization is not directly validated.
minor comments (6)
- [§1] The phrase 'where where' appears in the first numbered item of the introduction; please fix the duplication.
- [Figure 1 caption] The caption reads 'Statue QA' but should be 'Statute QA'.
- [Table 33] The header 'Reasining rollout' is misspelled; it should be 'Reasoning rollout'.
- [§6] In the paragraph comparing CoT and structured reasoning, 'the difference beween CoT' should be 'the difference between CoT'.
- [Appendix H, Table 26] The E5-mistral-7b-instruct baseline on Natural Questions (Recall@10 = 11.15) is far below E5-large-v2 (68.68) and also far below that model's own performance on the new benchmarks; this likely reflects a prompting/instruction issue for the instruct model. Please clarify the experimental setup for E5-mistral-7b-instruct or discuss this anomaly, since it affects the interpretability of the comparison across datasets.
- [§5.2] The term 'lower bound' for Housing Statute QA is defined as retrieval of all gold passages, which is not a lower bound on recall in the standard relevance-judgment sense; it is a stricter, all-relevant-retrieved metric. The terminology should be clarified to avoid confusion.
Circularity Check
No significant circularity: benchmark difficulty is an external measurement, with only a mild self-referential validity confound.
full rationale
This paper makes no first-principles derivation: it constructs two new labeled benchmarks, measures lexical overlap and retriever recall against external baselines and comparison datasets (NQ, HotpotQA, COLIEE, CLERC), and reports an experimental query-expansion method. The central difficulty claim is a measurement on hand-annotated gold passages, not a quantity fitted to or defined by the method. No parameter is fitted to the benchmark labels and then re-reported as a prediction. The structured-reasoning expansion is tested on the same benchmarks used to motivate it, and Appendix F shows that the expansion is designed to produce rule text that resembles gold passages, so part of the gain is a genre-match effect; this is a validity or confound concern, not a definitional equivalence, because the benchmark difficulty and the method's performance are independently observable. Self-citations (LegalBench, CaseHOLD) are contextual and not load-bearing. The acknowledged noise in Housing Statute QA labels is bracketed with upper/lower-bound recall and is a data-quality caveat, not a reduction of the conclusion to its inputs.
Assumptions & free parameters
assumptions (2)
- domain assumption TF-IDF cosine similarity is an appropriate measure of lexical similarity and correlates with the need for reasoning in retrieval tasks.
- domain assumption The gold passages annotated by law students and the LSC database citations are correct and sufficient to answer the queries.
Cite this review
Pith. "Pith review of A Reasoning-Focused Legal Retrieval Benchmark." pith.science (2026). https://pith.science/paper/FBRG3OXN
@misc{pith2026250503970,
author = {Pith},
title = {Pith review of: A Reasoning-Focused Legal Retrieval Benchmark},
year = {2026},
howpublished = {\url{https://pith.science/paper/FBRG3OXN}},
note = {Machine review of arXiv:2505.03970}
}
read the original abstract
As the legal community increasingly examines the use of large language models (LLMs) for various legal applications, legal AI developers have turned to retrieval-augmented LLMs ("RAG" systems) to improve system performance and robustness. An obstacle to the development of specialized RAG systems is the lack of realistic legal RAG benchmarks which capture the complexity of both legal retrieval and downstream legal question-answering. To address this, we introduce two novel legal RAG benchmarks: Bar Exam QA and Housing Statute QA. Our tasks correspond to real-world legal research tasks, and were produced through annotation processes which resemble legal research. We describe the construction of these benchmarks and the performance of existing retriever pipelines. Our results suggest that legal RAG remains a challenging application, thus motivating future research.
Figures
Forward citations
Cited by 1 Pith paper
-
Temporal Misgrounding in Legal RAG: A Versioned-Corpus Benchmark for French Tax Law
Version-conditioned retrieval over a 32,436-version French tax code corpus reaches 98.3% strict accuracy on 209 temporal-reasoning questions where LLM-only and static RAG score about 3%.
Reference graph
Works this paper leans on
-
[1]
Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, et al. 2021. On the opportunities and risks of foundation models. arXiv preprint arXiv:2108.07258 (2021)
arXiv 2021
-
[2]
Ilias Chalkidis, Manos Fergadiotis, Nikolaos Manginas, Eva Katakalou, and Prodro- mos Malakasiotis. 2021. Regulatory Compliance through Doc2Doc Information Retrieval: A case study in EU/UK legislation where text similarity has limitations. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main ...
2021
-
[3]
Andrea Anne Curcio, Carol L Chomsky, and Eileen R Kaufman. 2018. How to Build a Better Bar Exam. New York State Bar Association Journal (2018), 37–41
work page 2018
-
[4]
Matthew Dahl, Varun Magesh, Mirac Suzgun, and Daniel E Ho. 2024. Large legal fictions: Profiling legal hallucinations in large language models. arXiv preprint arXiv:2401.01301 (2024)
arXiv 2024
-
[5]
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), Jill Burstein, Christy...
doi:10.18653/v1/n 2019
-
[6]
Randy Goebel, Yoshinobu Kano, Mi-Young Kim, Juliano Rabelo, Ken Satoh, and Masaharu Yoshioka. 2024. Overview of Benchmark Datasets and Methods for the Legal Information Extraction/Entailment Competition (COLIEE) 2024. In New Frontiers in Artificial Intelligence, Toyotaro Suzumura and Mayumi Bono (Eds.). Springer Nature Singapore, Singapore, 109–124
work page 2024
-
[7]
Ashley, Ran Chen, Preethi Sureshkumar, Chen Wang, Eric Nyberg, and Vern R
Matthias Grabmair, Kevin D. Ashley, Ran Chen, Preethi Sureshkumar, Chen Wang, Eric Nyberg, and Vern R. Walker. 2015. Introducing LUIMA: an experiment in legal conceptual retrieval of vaccine injury decisions using a UIMA type system and tools. In Proceedings of the 15th International Conference on Artificial Intelli- gence and Law (San Diego, California) ...
-
[8]
Neel Guha, Julian Nyarko, Daniel Ho, Christopher Ré, Adam Chilton, Alex Chohlas-Wood, Austin Peters, Brandon Waldon, Daniel Rockmore, Diego Zam- brano, et al. 2024. Legalbench: A collaboratively built benchmark for measuring legal reasoning in large language models. Advances in Neural Information Pro- cessing Systems 36 (2024)
work page 2024
Show all 55 references
-
[9]
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2021. Measuring Massive Multitask Language Under- standing. Proceedings of the International Conference on Learning Representations (ICLR) (2021)
2021
-
[10]
Nils Holzenberger, Andrew Blair-Stanek, and Benjamin Van Durme. 2020. A dataset for statutory reasoning in tax law entailment and question answering. arXiv preprint arXiv:2005.05257 (2020)
2020 arXiv
-
[11]
Abe Bohan Hou, Orion Weller, Guanghui Qin, Eugene Yang, Dawn Lawrie, Nils Holzenberger, Andrew Blair-Stanek, and Benjamin Van Durme. 2024. CLERC: A Dataset for Legal Case Retrieval and Retrieval-Augmented Analysis Generation. arXiv preprint arXiv:2406.17186 (2024)
2024 arXiv
-
[12]
Rolf Jagerman, Honglei Zhuang, Zhen Qin, Xuanhui Wang, and Michael Bender- sky. 2023. Query expansion by prompting large language models. arXiv preprint arXiv:2305.03653 (2023)
2023 arXiv
-
[13]
Pengyue Jia, Yiding Liu, Xiangyu Zhao, Xiaopeng Li, Changying Hao, Shuaiqiang Wang, and Dawei Yin. 2023. Mill: Mutual verification with large language models for zero-shot query expansion. arXiv preprint arXiv:2310.19056 (2023)
2023 arXiv
-
[14]
Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, De- vendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al . 2023. Mistral 7B. arXiv preprint arXiv:2310.06825 (2023)
2023 arXiv
-
[15]
Weld, and Luke Zettlemoyer
Mandar Joshi, Eunsol Choi, Daniel S. Weld, and Luke Zettlemoyer. 2017. TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehen- sion. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics. Association for Comput...
2017
-
[16]
Daniel Martin Katz, Michael James Bommarito, Shang Gao, and Pablo Arredondo
-
[17]
Toutanova, Llion Jones, Ming-Wei Chang, Andrew Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov
Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Matthew Kelcey, Jacob Devlin, Kenton Lee, Kristina N. Toutanova, Llion Jones, Ming-Wei Chang, Andrew Dai, Jakob Uszkoreit, Quoc Le, and Slav...
2019
-
[18]
Angeliki Lazaridou, Adhi Kuncoro, Elena Gribovskaya, Devang Agrawal, Adam Liska, Tayfun Terzi, Mai Gimenez, Cyprien de Masson d 'Autume, Tomas Ko- cisky, Sebastian Ruder, Dani Yogatama, Kris Cao, Susannah Young, and Phil Blunsom. 2021. Mind the Gap: Assessing Temporal Generali...
2021
-
[19]
Yibin Lei, Yu Cao, Tianyi Zhou, Tao Shen, and Andrew Yates. 2024. Corpus- Steered Query Expansion with Large Language Models. arXiv preprint arXiv:2402.18031 (2024)
2024 arXiv
-
[20]
Jimmy Lin. 2019. The Neural Hype and Comparisons Against Weak Baselines. SIGIR Forum 52, 2 (jan 2019), 40–51. https://doi.org/10.1145/3308774.3308781
2019
-
[21]
Antoine Louis and Gerasimos Spanakis. 2022. A Statutory Article Retrieval Dataset in French. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics. Association for Computational Linguistics, Dublin, Ireland, 6789–6803. https://aclanthology....
2022
-
[22]
Antoine Louis, Gijs van Dijck, and Gerasimos Spanakis. 2024. Interpretable long-form legal question answering with retrieval-augmented large language models. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 38. 22266–22275
2024
-
[23]
LSC. 2021. Eviction Laws Database: Local Dataset. Prepared by the Center for Public Health Law Research at Temple University’s Beasley School of Law for Legal Services Corporation. https://www.lsc.gov/initiatives/effect-state-local- laws-evictions/lsc-eviction-laws-database
2021
-
[24]
Megan Ma, Aparna Sinha, Ankit Tandon, and Jennifer Richards. 2024. Generative AI Legal Landscape 2024 . Technical Report. Technical report
2024
-
[25]
Man- ning, and Daniel E
Varun Magesh, Faiz Surani, Matthew Dahl, Mirac Suzgun, Christopher D. Man- ning, and Daniel E. Ho. 2024. Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools. arXiv:2405.20362
2024 arXiv
-
[26]
Robert Mahari, Dominik Stammbach, Elliott Ash, and AlexSandy’ Pentland. 2023. LePaRD: A Large-Scale Dataset of Judges Citing Precedents. arXiv preprint arXiv:2311.09356 (2023)
2023 arXiv
-
[27]
Macedo Maia, Siegfried Handschuh, André Freitas, Brian Davis, Ross McDer- mott, Manel Zarrouk, and Alexandra Balahur. 2018. WWW’18 Open Challenge: Financial Opinion Mining and Question Answering. In Companion Proceedings of the The Web Conference 2018 (, Lyon, France,) (WWW ’1...
2018
-
[28]
Timothy McFarlin. 2023. A More Realistic Bar Exam Will Benefit Legal Education. The Bar Examiner 92, 2 (2023)
2023
-
[29]
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language Models are Unsupervised Multitask Learners. (2019)
2019
-
[30]
Stephen E Robertson and K Sparck Jones. 1976. Relevance weighting of search terms. Journal of the American Society for Information science 27, 3 (1976), 129– 146
1976
-
[31]
Guilherme Moraes Rosa, Ruan Chaves Rodrigues, Roberto Lotufo, and Rodrigo Nogueira. 2021. Yes, bm25 is a strong baseline for legal case retrieval. arXiv preprint arXiv:2105.05686 (2021)
2021 arXiv
-
[32]
Jon Saad-Falcon, Daniel Y Fu, Simran Arora, Neel Guha, and Christopher Ré
-
[33]
Jaromír Šavelka and Kevin D Ashley. 2022. Legal information retrieval for understanding statutory terms. Artificial Intelligence and Law (2022), 1–45
2022
-
[34]
arXiv preprint arXiv:2402.07440 (2024)
Benchmarking and building long-context retrieval models with loco and m2-bert. arXiv preprint arXiv:2402.07440 (2024)
2024 arXiv
-
[35]
Nandan Thakur, Nils Reimers, Andreas Rücklé, Abhishek Srivastava, and Iryna Gurevych. 2021. BEIR: A Heterogeneous Benchmark for Zero-shot Evaluation of Information Retrieval Models. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Tra...
2021
-
[36]
Hongjin Su, Howard Yen, Mengzhou Xia, Weijia Shi, Niklas Muennighoff, Han-yu Wang, Haisu Liu, Quan Shi, Zachary S Siegel, Michael Tang, et al. 2024. BRIGHT: A Realistic and Challenging Benchmark for Reasoning-Intensive Retrieval. arXiv preprint arXiv:2407.12883 (2024)
2024 arXiv
-
[37]
Liang Wang, Nan Yang, Xiaolong Huang, Binxing Jiao, Linjun Yang, Daxin Jiang, Rangan Majumder, and Furu Wei. 2022. Text embeddings by weakly-supervised contrastive pre-training. arXiv preprint arXiv:2212.03533 (2022)
2022 arXiv
-
[38]
Santosh T.y.s.s., Rashid Haddad, and Matthias Grabmair. 2024. ECtHR-PCR: A Dataset for Precedent Understanding and Prior Case Retrieval in the European Court of Human Rights. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resou...
2024
-
[39]
Wenhui Wang, Furu Wei, Li Dong, Hangbo Bao, Nan Yang, and Ming Zhou. 2020. MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre- Trained Transformers. In Advances in Neural Information Processing Systems , H. Larochelle, M. Ranzato, R. Hadsell, M.F. Ba...
2020
-
[40]
Xiaoyue Wang, Jianyou Wang, Weili Cao, Kaicheng Wang, Ramamohan Paturi, and Leon Bergen. 2024. BIRCO: A Benchmark of Information Retrieval Tasks with Complex Objectives. arXiv preprint arXiv:2402.14151 (2024)
2024 arXiv
-
[41]
Liang Wang, Nan Yang, Xiaolong Huang, Linjun Yang, Rangan Majumder, and Furu Wei. 2023. Improving text embeddings with large language models. arXiv preprint arXiv:2401.00368 (2023)
2023 arXiv
-
[42]
Lucia Zheng, Neel Guha, Brandon R Anderson, Peter Henderson, and Daniel E Ho. 2021. When does pretraining help? assessing self-supervised learning for law and the casehold dataset of 53,000+ legal holdings. InProceedings of the eighteenth international conference on artificial...
2021
-
[43]
due process
Haoxi Zhong, Chaojun Xiao, Cunchao Tu, Tianyang Zhang, Zhiyuan Liu, and Maosong Sun. 2020. JEC-QA: a legal-domain question answering dataset. In Proceedings of the AAAI conference on artificial intelligence , Vol. 34. 9701–9708. A Bar Exam QA Dataset Construction The gold pass...
1925
-
[44]
Cohen, Ruslan Salakhutdinov, and Christopher D
Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William W. Cohen, Ruslan Salakhutdinov, and Christopher D. Manning. 2018. HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering. In Conference on Empirical Methods in Natural Language Processing (EMNLP)
2018
-
[47]
This could have been different at eleven years of age, his age at the time of the hearing
The male applicant was three years old at the time of the initial departure to Mexico and was therefore not able not form an intention to reavail himself of the protection of Mexico. This could have been different at eleven years of age, his age at the time of the hearing. At ...
-
[48]
However, nothing in the evidence or in the submissions made by the parties makes it possible to determine whether the intention of the child could have been different from that of his mother. IX. Conclusion [33] In the circumstances of this case and in light of the foregoing, ...
-
[49]
existence of justification, transparency and intelligibility
However, the Court finds that the Board erred in its consideration of the applicant’s explanation relating to his business activities in Thailand. As outlined in<FRAGMENT_SUPPRESSED> , a review on the standard of reasonableness is concerned with the "existence of justification...
-
[50]
For the foregoing reasons the Court finds the Board’s decision to be unreasonable
-
[51]
[A] guilty plea is an admission of all the elements of a formal criminal charge
The Court agrees with the parties that there is no question of general interest to certify... Table 11: Example from COLIEE (Task 1.1) [6] A Reasoning-Focused Legal Retrieval Benchmark CS&Law ’25, March 25-27, 2025, Munich, Germany Query to constitute clear error. See United S...
1969
-
[52]
Intrusion upon seclusion: This refers to the unauthorized invasion into a person’s private affairs or physical space in a way that would be highly offensive to a reasonable person
-
[53]
Public disclosure of private facts: This involves the public dissemination of private and confidential information about an individual that would be highly offensive to a reasonable person and is not of legitimate public concern
-
[54]
False light: This occurs when false or misleading information is publicly attributed to an individual, portraying them in a highly offensive and false manner
-
[55]
"public disclosure of private facts
Appropriation of name or likeness: This refers to the unauthorized use of a person’s name, likeness, or identity for commercial purposes, without their consent. Based on the facts provided, it seems that the issue relevant to Pauline’s claim against the Journal would fall unde...
2025
-
[2024]
Philosophical Transactions of the Royal Society A 382, 2270 (2024), 20230254
Gpt-4 passes the bar exam. Philosophical Transactions of the Royal Society A 382, 2270 (2024), 20230254
2024
-
[5483]
https://aclanthology.org/2024.lrec-main.486
2024
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.