Pith. sign in

An empirical study of LLM-as-a-judge for LLM evaluation: Fine-tuned judge model is not a general substitute for GPT-4,

10 Pith papers cite this work, alongside 12 external citations. Polarity classification is still indexing.

10 Pith papers citing it
12 external citations · external index

citation-role summary

background 3 method 1

citation-polarity summary

years

2026 4 2024 6

representative citing papers

Section-Weighted Hybrid Approach for Legal Case Retrieval

cs.IR · 2026-06-02 · unverdicted · novelty 6.0

A section-aware hybrid retrieval system segments legal cases with an LLM, fuses BM25 and dense search via RRF, then applies Z-score normalized section-weighted comparisons to outperform baselines on a large benchmark.

A Survey on LLM-as-a-Judge

cs.CL · 2024-11-23 · unverdicted · novelty 4.0

A survey on LLM-as-a-Judge that reviews reliability strategies, proposes evaluation methods, and introduces a novel benchmark for assessing such systems.

ShieldGemma: Generative AI Content Moderation Based on Gemma

cs.CL · 2024-07-31 · unverdicted · novelty 4.0

ShieldGemma delivers a family of Gemma2-based classifiers that outperform Llama Guard and WildCard on public safety benchmarks while introducing a synthetic-data curation pipeline for safety tasks.

citing papers explorer

Showing 10 of 10 citing papers.