UA-Legal-Bench is a new five-task benchmark for Ukrainian legal reasoning that demonstrates task-dependent few-shot prompting effects and the need for macro-F1 over accuracy on imbalanced classes.
LEGAL - BERT : The Muppets straight out of Law School
7 Pith papers cite this work, alongside 18 external citations. Polarity classification is still indexing.
years
2026 7verdicts
UNVERDICTED 7representative citing papers
RISE is an inference-time semantic reranking framework that refines low-confidence predictions in rhetorical role labeling using contrastively learned label representations, delivering an average +9.15 macro-F1 gain on hard examples across eight datasets and seven models.
A citation graph built from the complete Ukrainian court registry recovers legal domain boundaries via community detection and predicts legislative importance with AUC 0.9984.
Amortized optimization with policy gradients and graph knowledge selects informative word subsets to explain black-box DLM outputs.
N2I-RAG is an agentic RAG pipeline that automates binary legal indicator computation from complex normative texts with explicit traceability to provisions.
Domain-trained small language model Olava Extract outperforms frontier LLMs on structured contract extraction with macro F1 0.812, micro F1 0.842, highest precision, and 78-97% lower inference cost.
Further pre-training ModernBERT on US court opinions improves results on legal datasets compared to the base model, with gains similar to early BERT domain adaptation work.
citing papers explorer
-
UA-Legal-Bench: A Benchmark for Evaluating Large Language Models on Ukrainian Legal Reasoning
UA-Legal-Bench is a new five-task benchmark for Ukrainian legal reasoning that demonstrates task-dependent few-shot prompting effects and the need for macro-F1 over accuracy on imbalanced classes.
-
Semantic Reranking at Inference Time for Hard Examples in Rhetorical Role Labeling
RISE is an inference-time semantic reranking framework that refines low-confidence predictions in rhetorical role labeling using contrastively learned label representations, delivering an average +9.15 macro-F1 gain on hard examples across eight datasets and seven models.
-
Automatic Construction of a Legal Citation Graph from 100 Million Ukrainian Court Decisions: Large-Scale Extraction, Topological Analysis, and Ontology-Driven Clustering
A citation graph built from the complete Ukrainian court registry recovers legal domain boundaries via community detection and predicts legislative importance with AUC 0.9984.
-
Explaining Black-Box Language Models: Learning to Optimize Linguistically-Structured Word Subsets
Amortized optimization with policy gradients and graph knowledge selects informative word subsets to explain black-box DLM outputs.
-
From Norms to Indicators (N2I-RAG): An Agentic Retrieval-Augmented Generation Framework for Legal Indicator Computation
N2I-RAG is an agentic RAG pipeline that automates binary legal indicator computation from complex normative texts with explicit traceability to provisions.
-
A Few Good Clauses: Comparing LLMs vs Domain-Trained Small Language Models on Structured Contract Extraction
Domain-trained small language model Olava Extract outperforms frontier LLMs on structured contract extraction with macro F1 0.812, micro F1 0.842, highest precision, and 78-97% lower inference cost.
-
Legal Domain Adaptation of Modern BERT Models
Further pre-training ModernBERT on US court opinions improves results on legal datasets compared to the base model, with gains similar to early BERT domain adaptation work.