REVIEW 4 major objections 6 minor 52 references
LegalViz: Legal Text Visualization by Text To Diagram Generation
T0 review · 4 major / 6 minor · reviewed 2026-08-08 · deepseek-v4-flash
Pith's one-line read The paper introduces LegalViz, a 23-language dataset of 7,010 legal text–diagram pairs, and claims that fine-tuned open models beat few-shot GPT models at generating legal diagrams.
desk verdict New multilingual legal-diagram dataset with a clear task definition, but the headline fine-tuning-vs-GPT result is undermined by a case-level train/test split that the paper never describes. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Carrying the argument is LegalViz itself: gold-standard graphs written in Graphviz DOT code, using shape conventions (double octagon for legal entities, trapezium for legal sources, square for legal statements, ellipse for deceased persons) and edge conventions (directed edges for transactions, dashed edges for non-exercisable rights, dotted edges for succession, undirected edges for equivalent or explanatory relations). The evaluation pipeline is the second half of the machinery: generated DOT code is parsed, nodes are aligned to reference nodes by bipartite matching on BERTScore similarities, and then F1 metrics are computed at the graph-structure level, node-text level, edge-label level, and legal-category level. This combination lets the paper measure whether a model captured the legal content, not just whether it emitted syntactically valid code.
What would settle it
One concrete check: redraw a random subset of the test gold graphs by independent legal experts and measure structural agreement among annotators; if expert graphs are no closer to the original annotations than the best fine-tuned model is, the gold standard is not stable and the reported rankings lose their force.
Extended reading notes
Core claim
The authors claim that legal visualization can be treated as a text-to-diagram generation task and that a relatively small, professionally annotated dataset is enough to make open models better at it than closed few-shot models. They report that fine-tuned models beat GPT models on all three graph-based metrics (Graph, Graph&Node, and Graph&Node&Edge) and on all four legal-content metrics (Entity, Relations & Transactions, Source, and Statement), with fine-tuned Gemma2-9B reaching the top scores in most columns. They also report that fine-tuning sharply raises the rate of valid DOT code generation and roughly triples the Statement score, which they interpret as evidence that the dataset teaches models to summarize legally significant facts rather than merely copy graph syntax.
Load-bearing premise
The whole evaluation assumes the gold-standard diagrams are essentially correct; they were drawn by a single legal expert and then machine-translated into 22 languages, so any inconsistency in that expert's choices, or any semantic drift introduced by translation, directly weakens the claim that the fine-tuned models are genuinely better.
Editorial extensions
If this is right
- Legal judgments, including European court and national court decisions, could be rendered as diagrams automatically, lowering the entry barrier for non-experts reading case law.
- The graph-based and legal-content evaluation metrics give the field a way to compare diagram-generation systems beyond exact-code matching.
- Fine-tuning smaller open-weight models on domain-structured datasets can close the gap with much larger closed models on structured output tasks.
- The 23-language version of the dataset allows the task to be studied across low- and high-resource European languages, with the authors reporting that low-resource languages score lower.
- Statement generation remains the hardest aspect, so summarizing the legally significant facts is the current bottleneck for automatic legal visualization.
Reading between the lines
- The gold standard rests on a single legal annotator and machine-assisted translation, so the reported margins are probably optimistic relative to a multi-annotator gold standard; an inter-annotator agreement study would show how much of the ranking is annotation style.
- Since the 7,010 instances come from more than 300 source judgments translated into 23 languages, the dataset's coverage of distinct legal situations is closer to hundreds than to thousands; extending it to new jurisdictions and document types is an untested generalization.
- If diagram generation works well, the DOT graph could serve as an interpretable intermediate representation for downstream legal tasks such as question answering, summarization, or argument retrieval, but the paper does not test this.
- The evaluation depends on BERTScore-based alignment, so graphs with fused or split nodes, paraphrases, or near-synonymous legal terms could be scored more generously or harshly than exact semantic equivalence would warrant; a human rating study on a sample would calibrate the metrics.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces LegalViz, a multilingual dataset of 7,010 legal-text/Graphviz-diagram pairs covering 23 EUR-Lex languages. The diagrams are manually annotated with legal entities, legal relationships/transactions, legal sources, and legal statements, and the non-English diagrams are produced by GPT-4 translation with manual checking. The paper also proposes new evaluation metrics for graph structure (Graph, Graph&Node, Graph&Node&Edge) and for legal-content categories (Entity, Relations & Transactions, Source, Statement). Experiments fine-tune CodeLlama, Llama, and Gemma models on the dataset and compare them with few-shot GPT-3.5/4/4o; the authors report that fine-tuned models outperform the GPT baselines on most metrics and conclude that the dataset is effective for legal text visualization.
Significance. If the central claim holds, LegalViz is a useful new resource: it is, to my knowledge, the first legal text-to-diagram dataset, it spans 23 languages, and it pairs legal content categories with a concrete graph formalism. The paper ships a public dataset and a detailed fine-tuning evaluation across several model families, which are strengths. However, the headline empirical claim depends on two things that are not yet established: (i) the train/test split must prevent the same source judgment from appearing in both training and test through its multiple language versions, and (ii) the gold annotations and the proposed metrics must be reliable enough to support claims of superiority over GPT models. Both points need to be addressed before the conclusion is supported.
major comments (4)
- [§3.2, Tables 1 and 2] The paper does not describe how the 4,710/1,150/1,150 split was made. Because each source judgment is translated into 23 languages and Table 2 shows roughly 300 instances per language, a random instance-level split can place the English version of a case in training and the German version of the same case in the test set. Since all language versions of the gold diagram are derived from the same English annotation, fine-tuned models could then score higher on test instances by recognizing a judgment already seen in another language, rather than by learning to visualize unseen legal documents. This directly affects the headline comparison in Table 3. Please report the split criterion, provide case-level identifiers, and re-run the comparison with a case-disjoint split, or otherwise demonstrate that no source judgment appears in both training and test sets.
- [§3.2] The gold diagrams were produced by a single legal annotator, and the 22 non-English versions were generated by GPT-4 translation with manual checking. No inter-annotator agreement, annotation consistency statistics, or translation-quality metrics are reported. Because the dataset is the main contribution and the fine-tuning results in Table 3 are measured against these gold graphs, the reliability of the gold data is load-bearing. Please include an annotation guideline summary, report inter-annotator agreement on a sample independently annotated by a second legal expert, and provide a manual or automatic evaluation of translation fidelity, especially for graph structure and legal terminology.
- [§4.1, Eq. for TP/FP/FN] The definition of TP for the Graph, Graph&Node, and Graph&Node&Edge metrics sums over all pairs (eh, er) for which fGraph(eh, er) = 1. If multiple reference edges share the same aligned endpoints, a single hypothesis edge can contribute to TP more than once, making TP larger than |Eh| and FP negative. Parallel labeled edges are plausible in legal diagrams, for example multiple transactions between the same two entities, so the F1 computation is not well-defined in general. Please replace the pairwise summation with a proper edge matching, for instance a maximum-weight bipartite matching on edges after node alignment, and update the reported numbers if the change affects them.
- [§4, §5.2] The proposed graph and legal-content metrics are introduced by the authors, but no evidence is given that they correlate with human judgments of diagram quality or legal correctness. The conclusion that fine-tuned models 'confirm the effectiveness' of LegalViz therefore rests entirely on these unvalidated metrics. Please add a human evaluation on a sample, for example expert and non-expert ratings of generated diagrams, or report correlation with human preference; if such validation is out of scope, the empirical claims should be softened accordingly.
minor comments (6)
- [Title page and affiliations] The first affiliation contains a typo: 'Nara Institute of Scient and Technology' should be 'Nara Institute of Science and Technology'.
- [§2, Related Work] The related work paragraph contains a spelling error: 'Chinna' should be 'China' when citing Ye et al. (2018).
- [Figure 3 caption] The caption reads 'finetuned modesl of Gemma2-9B'; 'modesl' should be 'models'.
- [§5.1, Experimental settings] The few-shot GPT experiments do not report decoding hyperparameters such as temperature, top-p, or maximum new tokens; these should be stated for reproducibility.
- [Throughout] The dataset name 'EUR-LEX' is written inconsistently as 'EUR-LEX', 'EUR-Lex', and 'EUR-lex'; please standardize the spelling.
- [Appendix J] The dataset example shows 'The Comission of the European Comminities' with spelling errors; if this is part of the released annotation, the gold data should be corrected.
Circularity Check
No significant circularity: evaluation is a held-out gold-standard comparison; the under-specified split is a reporting gap, not a circular derivation.
full rationale
Derivation chain: manually annotated English gold graphs are translated to 23 languages, open LLMs are fine-tuned on LegalViz, and generated DOT graphs are compared against held-out gold annotations. The evaluation in Sec. 4 is a reference-based comparison; nothing is fitted on the test split, and the new evaluation metrics are applied uniformly to all models. No fitted constant is later renamed as a prediction, and no self-citation is load-bearing (the only self-reference is the dataset URL). The central claim that fine-tuned models outperform GPTs is an empirical comparison against the same gold standard, so it does not reduce by definition to the dataset's construction. The one substantive weakness is under-reporting: Table 1 gives only instance counts and Sec. 3.2 states 'more than 300 unique legal texts' with '23 language variations', but the paper never explicitly states that train/val/test are disjoint by unique judgment. If the split were per instance rather than per case, cross-lingual near-duplicates could leak. The counts (val = test = 1,150, exactly 50 × 23) suggest a case-level split, so this is a reporting gap and a validity risk, not a demonstrated circular reduction. The unvalidated, author-defined metrics and single-annotator gold are correctness/scope limitations, not circularity. Score 0.
Assumptions & free parameters
assumptions (4)
- domain assumption The annotation schema (doubleoctagon for legal entities, trapezium for legal sources, square for statements, and the edge styles for transactions, family relations, and so on) is accepted as the correct representation of legal content for visualization.
- domain assumption The automatic evaluation metrics (graph F1 with BERTScore-based node/edge similarity) are assumed to measure diagram quality and legal-content fidelity.
- domain assumption GPT-4-based translation of the manually created English diagrams preserves the legal meaning and graph structure across the 22 other EU languages.
- domain assumption EUR-LEX judgments from 2006 to 2019, selected from factual sections, are representative legal documents for the visualization task.
Cite this review
Pith. "Pith review of LegalViz: Legal Text Visualization by Text To Diagram Generation." pith.science (2026). https://pith.science/paper/4LDUXAFT
@misc{pith2026250206147,
author = {Pith},
title = {Pith review of: LegalViz: Legal Text Visualization by Text To Diagram Generation},
year = {2026},
howpublished = {\url{https://pith.science/paper/4LDUXAFT}},
note = {Machine review of arXiv:2502.06147}
}
read the original abstract
Legal documents including judgments and court orders require highly sophisticated legal knowledge for understanding. To disclose expert knowledge for non-experts, we explore the problem of visualizing legal texts with easy-to-understand diagrams and propose a novel dataset of LegalViz with 23 languages and 7,010 cases of legal document and visualization pairs, using the DOT graph description language of Graphviz. LegalViz provides a simple diagram from a complicated legal corpus identifying legal entities, transactions, legal sources, and statements at a glance, that are essential in each judgment. In addition, we provide new evaluation metrics for the legal diagram visualization by considering graph structures, textual similarities, and legal contents. We conducted empirical studies on few-shot and finetuning large language models for generating legal diagrams and evaluated them with these metrics, including legal content-based evaluation within 23 languages. Models trained with LegalViz outperform existing models including GPTs, confirming the effectiveness of our dataset.
Figures
Figures from the paper (2 more)
Reference graph
Works this paper leans on
-
[1]
Iosif Angelidis, Ilias Chalkidis, and Manolis Koubarakis. 2018. https://api.semanticscholar.org/CorpusID:55699546 Named entity recognition, linking and generation for greek legislation . In International Conference on Legal Knowledge and Information Systems
work page 2018
-
[3]
Dennis Aumiller, Ashish Chouhan, and Michael Gertz. 2022. https://doi.org/10.18653/v1/2022.emnlp-main.519 EUR -lex-sum: A multi- and cross-lingual dataset for long-form summarization in the legal domain . In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 7626--7639, Abu Dhabi, United Arab Emirates. Associatio...
-
[4]
Claire Barale, Mark Klaisoongnoen, Pasquale Minervini, Michael Rovatsos, and Nehal Bhuta. 2023. https://doi.org/10.18653/v1/2023.nllp-1.24 A sy L ex: A dataset for legal language processing of refugee claims . In Proceedings of the Natural Legal Language Processing Workshop 2023, pages 244--257, Singapore. Association for Computational Linguistics
-
[5]
Jonas Belouadi, Anne Lauscher, and Steffen Eger. 2024. Automatikz: Text-guided synthesis of scientific vector graphics with tikz. In International Conference on Learning Representations (ICLR)
work page 2024
-
[6]
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gr...
2020
-
[7]
Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuan-Fang Li, Scott M
S \'e bastien Bubeck, Varun Chandrasekaran, Ronen Eldan, John A. Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuan-Fang Li, Scott M. Lundberg, Harsha Nori, Hamid Palangi, Marco Tulio Ribeiro, and Yi Zhang. 2023. https://api.semanticscholar.org/CorpusID:257663729 Sparks of artificial general intelligence: Early experiments with gpt-4 . ArXiv, a...
arXiv 2023
-
[8]
Ilias Chalkidis, Emmanouil Fergadiotis, Prodromos Malakasiotis, Nikolaos Aletras, and Ion Androutsopoulos. 2019. https://doi.org/10.18653/v1/W19-2209 Extreme multi-label legal text classification: A case study in EU legislation . In Proceedings of the Natural Legal Language Processing Workshop 2019, pages 78--87, Minneapolis, Minnesota. Association for Co...
-
[9]
Ilias Chalkidis, Manos Fergadiotis, and Ion Androutsopoulos. 2021 a . https://doi.org/10.18653/v1/2021.emnlp-main.559 M ulti EURLEX - a multi-lingual and multi-label legal document classification dataset for zero-shot cross-lingual transfer . In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 6974--6996, Onlin...
Show all 52 references
-
[10]
Ilias Chalkidis, Manos Fergadiotis, Dimitrios Tsarapatsanis, Nikolaos Aletras, Ion Androutsopoulos, and Prodromos Malakasiotis. 2021 b . https://doi.org/10.18653/v1/2021.naacl-main.22 Paragraph-level rationale extraction through regularization: A case study on E uropean court ...
2021 doi
-
[11]
Ilias Chalkidis, Nicolas Garneau, Catalina Goanta, Daniel Katz, and Anders S gaard. 2023. https://doi.org/10.18653/v1/2023.acl-long.865 L e XF iles and L egal LAMA : Facilitating E nglish multinational legal language model development . In Proceedings of the 61st Annual Meetin...
2023 doi
-
[12]
Ilias Chalkidis, Abhik Jana, Dirk Hartung, Michael Bommarito, Ion Androutsopoulos, Daniel Katz, and Nikolaos Aletras. 2022 a . https://doi.org/10.18653/v1/2022.acl-long.297 L ex GLUE : A benchmark dataset for legal language understanding in E nglish . In Proceedings of the 60t...
2022 doi
-
[13]
Ilias Chalkidis, Tommaso Pasini, Sheng Zhang, Letizia Tomada, Sebastian Schwemer, and Anders S gaard. 2022 b . https://doi.org/10.18653/v1/2022.acl-long.301 F air L ex: A multilingual benchmark for evaluating fairness in legal text processing . In Proceedings of the 60th Annua...
2022 doi
-
[15]
Huajie Chen, Deng Cai, Wei Dai, Zehui Dai, and Yadong Ding. 2019. https://doi.org/10.18653/v1/D19-1667 Charge-based prison term prediction with deep gating network . In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th Internati...
2019 doi
-
[16]
Jonathan H Choi, Kristin E Hickman, Amy B Monahan, and Daniel Schwarcz. 2021. Chatgpt goes to law school. J. Legal Educ., 71:387
2021
-
[17]
Fenia Christopoulou, Guchun Zhang, and Gerasimos Lampouras. 2024. https://aclanthology.org/2024.eacl-long.72 Text-to-code generation with modality-relative pre-training . In Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguis...
2024
-
[18]
Ona de Gibert Bonet, Aitor Garc \' a Pablos, Montse Cuadros, and Maite Melero. 2022. https://aclanthology.org/2022.lrec-1.400 S panish datasets for sensitive entity detection in the legal domain . In Proceedings of the Thirteenth Language Resources and Evaluation Conference, p...
2022
-
[19]
Kasper Drawzeski, Andrea Galassi, Agnieszka Jablonowska, Francesca Lagioia, Marco Lippi, Hans Wolfgang Micklitz, Giovanni Sartor, Giacomo Tagiuri, and Paolo Torroni. 2021. https://doi.org/10.18653/v1/2021.nllp-1.1 A corpus for multilingual analysis of online terms of service ....
2021 doi
-
[20]
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, et al. 2024. https://arxiv.org/abs/2407.21783 The llama 3 herd of models . Preprint, arXiv:2407.21783
2024 arXiv
-
[21]
Matt Dunn, Levent Sagun, Hale S irin, and Daniel Chen. 2017. https://doi.org/10.1145/3086512.3086537 Early predictability of asylum court decisions . In Proceedings of the 16th Edition of the International Conference on Articial Intelligence and Law, ICAIL '17, page 233–236, N...
2017
-
[22]
Mohamed Elaraby and Diane Litman. 2022. https://aclanthology.org/2022.coling-1.540 A rg L egal S umm: Improving abstractive summarization of legal documents with argument mining . In Proceedings of the 29th International Conference on Computational Linguistics, pages 6187--619...
2022
-
[23]
Jens Frankenreiter and Julian Nyarko. 2022. Natural language processing in legal tech. Legal Tech and the Future of Civil Justice
2022
-
[24]
Ho, Christopher Ré, Adam Chilton, Aditya Narayana, Alex Chohlas-Wood, Austin Peters, Brandon Waldon, Daniel N
Neel Guha, Julian Nyarko, Daniel E. Ho, Christopher Ré, Adam Chilton, Aditya Narayana, Alex Chohlas-Wood, Austin Peters, Brandon Waldon, Daniel N. Rockmore, Diego Zambrano, Dmitry Talisman, Enam Hoque, Faiz Surani, Frank Fagan, Galit Sarfaty, Gregory M. Dickinson, Haggai Porat...
2023
-
[25]
Dan Hendrycks, Collin Burns, Anya Chen, and Spencer Ball. 2021. Cuad: An expert-annotated nlp dataset for legal contract review. NeurIPS
2021
-
[26]
Nils Holzenberger, Andrew Blair-Stanek, and Benjamin Van Durme. 2020. https://api.semanticscholar.org/CorpusID:218581117 A dataset for statutory reasoning in tax law entailment and question answering . In NLLP@KDD
2020
-
[27]
Zikun Hu, Xiang Li, Cunchao Tu, Zhiyuan Liu, and Maosong Sun. 2018. https://aclanthology.org/C18-1041 Few-shot charge prediction with discriminative legal attributes . In Proceedings of the 27th International Conference on Computational Linguistics, pages 487--498, Santa Fe, N...
2018
-
[28]
Wonseok Hwang, Dongjun Lee, Kyoungyeon Cho, Hanuhl Lee, and Minjoon Seo. 2024. A multi-task benchmark for korean legal language understanding and judgement prediction. In Proceedings of the 36th International Conference on Neural Information Processing Systems, NIPS '22, Red H...
2024
-
[29]
Bommarito II, Daniel Martin Katz, and Eric M
Michael J. Bommarito II, Daniel Martin Katz, and Eric M. Detterman. 2021. https://doi.org/10.4337/9781788972826.00017 Chapter 11: LexNLP: Natural language processing and information extraction for legal and regulatory texts . Edward Elgar Publishing, Cheltenham, UK
2021
-
[30]
Zhijing Jin, Qipeng Guo, Xipeng Qiu, and Zheng Zhang. 2020. https://doi.org/10.18653/v1/2020.coling-main.217 G en W iki: A dataset of 1.3 million content-sharing text and graphs for unsupervised graph-to-text generation . In Proceedings of the 28th International Conference on ...
2020 doi
-
[31]
Bommarito, II, and Josh Blackman
Daniel Martin Katz, Michael J. Bommarito, II, and Josh Blackman. 2017. https://doi.org/10.1371/journal.pone.0174698 A general approach for predicting the behavior of the supreme court of the united states . PLOS ONE, 12(4):1--18
2017 doi
-
[32]
Daniel Martin Katz, Dirk Hartung, Lauritz Gerlach, Abhik Jana, and Michael James Bommarito. 2023. https://api.semanticscholar.org/CorpusID:256440319 Natural language processing in the legal domain . ArXiv, abs/2302.12039
2023 arXiv
-
[33]
Arshdeep Kaur and Bojan Bozic. 2019. https://api.semanticscholar.org/CorpusID:207824257 Convolutional neural network-based automatic prediction of judgments of the european court of human rights . In Irish Conference on Artificial Intelligence and Cognitive Science
2019
-
[34]
Rik Koncel-Kedziorski, Dhanush Bekal, Yi Luan, Mirella Lapata, and Hannaneh Hajishirzi. 2019. https://doi.org/10.18653/v1/N19-1238 T ext G eneration from K nowledge G raphs with G raph T ransformers . In Proceedings of the 2019 Conference of the North A merican Chapter of the ...
2019 doi
-
[35]
Tiffany H Kung, Morgan Cheatham, Arielle Medenilla, Czarina Sillos, Lorie De Leon, Camille Elepa \ n o, Maria Madriaga, Rimel Aggabao, Giezel Diaz-Candido, James Maningo, et al. 2023. Performance of chatgpt on usmle: potential for ai-assisted medical education using large lang...
2023
-
[36]
Pan Lu, Swaroop Mishra, Tony Xia, Liang Qiu, Kai-Wei Chang, Song-Chun Zhu, Oyvind Tafjord, Peter Clark, and Ashwin Kalyan. 2022. Learn to explain: Multimodal reasoning via thought chains for science question answering. In The 36th Conference on Neural Information Processing Sy...
2022
-
[37]
de Campos, Renato R
Pedro Henrique Luz de Araujo, Te \'o filo E. de Campos, Renato R. R. de Oliveira, Matheus Stauffer, Samuel Couto, and Paulo Bermejo. 2018. Lener-br: A dataset for named entity recognition in brazilian legal text. In Computational Processing of the Portuguese Language, pages 31...
2018
-
[38]
Masha Medvedeva, Michel Vols, and Martijn Wieling. 2020. https://doi.org/10.1007/s10506-019-09255-y Using machine learning to predict decisions of the european court of human rights . Artificial Intelligence and Law, 28(2):237--266
2020 doi
-
[39]
Joel Niklaus, Ilias Chalkidis, and Matthias St \"u rmer. 2021. https://doi.org/10.18653/v1/2021.nllp-1.3 S wiss-judgment-prediction: A multilingual legal judgment prediction benchmark . In Proceedings of the Natural Legal Language Processing Workshop 2021, pages 19--35, Punta ...
2021 doi
-
[40]
Joel Niklaus, Veton Matoshi, Pooja Rani, Andrea Galassi, Matthias St \"u rmer, and Ilias Chalkidis. 2023. https://doi.org/10.18653/v1/2023.findings-emnlp.200 LEXTREME : A multi-lingual and multi-task benchmark for the legal domain . In Findings of the Association for Computati...
2023 doi
-
[41]
OpenAI. 2023. GPT -4 technical report. Technical report
2023
-
[42]
Vasile Pais, Maria Mitrofan, Carol Luca Gasan, Vlad Coneschi, and Alexandru Ianov. 2021. https://doi.org/10.18653/v1/2021.nllp-1.2 Named entity recognition in the R omanian legal domain . In Proceedings of the Natural Legal Language Processing Workshop 2021, pages 9--18, Punta...
2021 doi
-
[43]
Welty, Christopher A
Gemma Team Morgane Riviere, Shreya Pathak, Pier Giuseppe Sessa, Cassidy Hardin, Surya Bhupatiraju, L'eonard Hussenot, Thomas Mesnard, Bobak Shahriari, Alexandre Ram'e, Johan Ferret, Peter Liu, Pouya Dehghani Tafti, Abe Friesen, Michelle Casbon, Sabela Ramos, Ravin Kumar, Charl...
2024 arXiv
-
[44]
Baptiste Rozi \`e re, Jonas Gehring, Fabian Gloeckle, Sten Sootla, Itai Gat, Xiaoqing Tan, Yossi Adi, Jingyu Liu, Tal Remez, J \'e r \'e my Rapin, Artyom Kozhevnikov, I. Evtimov, Joanna Bitton, Manish P Bhatt, Cristian Cant \'o n Ferrer, Aaron Grattafiori, Wenhan Xiong, Alexan...
2023 arXiv
-
[45]
Freda Shi, Daniel Fried, Marjan Ghazvininejad, Luke Zettlemoyer, and Sida I. Wang. 2022. https://doi.org/10.18653/v1/2022.emnlp-main.231 Natural language to code translation with execution . In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Process...
2022 doi
-
[46]
Don Tuggener, Pius von D \"a niken, Thomas Peetz, and Mark Cieliebak. 2020. https://aclanthology.org/2020.lrec-1.155 LEDGAR : A large-scale multi-label corpus for text classification of legal provisions in contracts . In Proceedings of the Twelfth Language Resources and Evalua...
2020
-
[47]
Stefanie Urchs., Jelena Mitrović., and Michael Granitzer. 2021. https://doi.org/10.5220/0010187305150521 Design and implementation of german legal decision corpora . In Proceedings of the 13th International Conference on Agents and Artificial Intelligence - Volume 2: ICAART, p...
2021 doi
-
[48]
Chaojun Xiao, Haoxiang Zhong, Zhipeng Guo, Cunchao Tu, Zhiyuan Liu, Maosong Sun, Yansong Feng, Xianpei Han, Zhen Hu, Heng Wang, and Jianfeng Xu. 2018. https://api.semanticscholar.org/CorpusID:49652844 Cail2018: A large-scale legal dataset for judgment prediction . ArXiv, abs/1...
2018 arXiv
-
[49]
Hai Ye, Xin Jiang, Zhunchen Luo, and Wenhan Chao. 2018. https://doi.org/10.18653/v1/N18-1168 Interpretable charge predictions for criminal cases: Learning to generate court views from fact descriptions . In Proceedings of the 2018 Conference of the North A merican Chapter of t...
2018 doi
-
[50]
Abhay Zala, Han Lin, Jaemin Cho, and Mohit Bansal. 2023. Diagrammergpt: Generating open-domain, open-platform diagrams via llm planning
2023
-
[51]
Weinberger, and Yoav Artzi
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020. https://openreview.net/forum?id=SkeHuCVFDr Bertscore: Evaluating text generation with BERT . In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April...
2020
-
[52]
Anderson, Peter Henderson, and Daniel E
Lucia Zheng, Neel Guha, Brandon R. Anderson, Peter Henderson, and Daniel E. Ho. 2021. https://doi.org/10.1145/3462757.3466088 When does pretraining help? assessing self-supervised learning for law and the casehold dataset of 53,000+ legal holdings . In Proceedings of the Eight...
2021
-
[53]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list...
-
[54]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...
Reviewed August 8, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.