Pith. sign in

hub

Lexam: Benchmarking legal reasoning on 340 law exams

18 Pith papers cite this work. Polarity classification is still indexing.

18 Pith papers citing it

hub tools

citation-role summary

background 1

citation-polarity summary

years

2026 18

roles

background 1

polarities

background 1

representative citing papers

Magis-Bench: Evaluating LLMs on Magistrate-Level Legal Tasks

cs.CL · 2026-05-08 · accept · novelty 7.0

Magis-Bench is a new benchmark of 74 magistrate-level legal writing tasks from Brazilian exams where the strongest LLMs reach only 6.97/10, showing judicial reasoning remains difficult for current models.

HARVE: Hacking-Aware Reward-Head Vector Editing for Robust Reward Models

cs.LG · 2026-06-02 · unverdicted · novelty 6.0

HARVE removes the component of the reward-head vector aligned with a multi-directional hacking subspace from residual streams using a small set of contrastive examples, improving robustness on RewardHackBench across eight models without fine-tuning while preserving general capability.

Investigating Multi-Agent Deliberation in Law

cs.AI · 2026-06-29 · unverdicted · novelty 5.0

Multi-agent deliberation frameworks for legal reasoning with LLMs match baseline performance but yield distinct answers that cover cases single models miss.

GradeLegal: Automated Grading for German Legal Cases

cs.CL · 2026-05-20 · unverdicted · novelty 5.0

Reasoning-oriented LLMs reach up to 0.91 quadratic weighted kappa agreement with experts on public law cases when given sample solutions and grading rubrics, but only 0.60 on criminal law cases.

citing papers explorer

Showing 18 of 18 citing papers.