Pith. sign in

REVIEW 15 cited by

LLM-Eval: Unified Multi-Dimensional Automatic Evaluation for Open-Domain Conversations with Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.13711 v1 pith:EACVKCNY submitted 2023-05-23 cs.CL cs.AI

classification cs.CLcs.AI
keywords evaluationllm-evalopen-domainunifiedautomaticconversationconversationslanguage
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We propose LLM-Eval, a unified multi-dimensional automatic evaluation method for open-domain conversations with large language models (LLMs). Existing evaluation methods often rely on human annotations, ground-truth responses, or multiple LLM prompts, which can be expensive and time-consuming. To address these issues, we design a single prompt-based evaluation method that leverages a unified evaluation schema to cover multiple dimensions of conversation quality in a single model call. We extensively evaluate the performance of LLM-Eval on various benchmark datasets, demonstrating its effectiveness, efficiency, and adaptability compared to state-of-the-art evaluation methods. Our analysis also highlights the importance of choosing suitable LLMs and decoding strategies for accurate evaluation results. LLM-Eval offers a versatile and robust solution for evaluating open-domain conversation systems, streamlining the evaluation process and providing consistent performance across diverse scenarios.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 15 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. OmniPresent: Generating Coherent Presentation Suites from Scientific Papers

    cs.SE 2026-07 conditional novelty 6.5 of 10

    A multi-agent HTML pipeline with shared knowledge and cross-artifact verify-and-repair generates coherent poster/slides/video/page suites from papers and beats specialized baselines on OmniPreBench.

  2. Agreement Metrics for LLM-as-Judge Evaluation: What to Report and Why

    cs.CL 2026-05 conditional novelty 6.0 of 10

    On binary verdicts, Pearson, Spearman, Kendall's tau-b, phi, and the Matthews correlation are a single statistic, so most multi-metric agreement reports repeat one number under different names.

  3. AraTable: Benchmarking LLMs' Reasoning and Understanding of Arabic Tabular Data

    cs.CL 2025-07 conditional novelty 6.0 of 10

    AraTable is the first Arabic tabular QA benchmark; its experiments show LLMs are much weaker at reasoning over Arabic tables than at direct lookup.

  4. How Stylistic Similarity Shapes Preferences in Dialogue Dataset with User and Third Party Evaluations

    cs.CL 2025-07 conditional novelty 6.0 of 10

    A new open-domain dialogue dataset shows that users' own judgments of stylistic similarity correlate with their preference (Spearman r=0.67-0.75), but third-party stylistic similarity judgments do not, indicating a ga...

  5. How Significant Are the Real Performance Gains? An Unbiased Evaluation Framework for GraphRAG

    cs.CL 2025-05 conditional novelty 6.0 of 10

    A new GraphRAG evaluation framework using graph-grounded questions and bias-correction yields much smaller win rates than earlier reports, casting doubt on reported GraphRAG gains.

  6. LLM-based Evaluation Policy Extraction for Ecological Modeling

    cs.AI 2025-05 conditional novelty 6.0 of 10

    APEF learns interpretable evaluation policies for ecological time-series models by combining an LLM-driven weight optimizer with human or predefined pairwise preference annotations.

  7. Overconfidence in LLM-as-a-Judge: Diagnosis and Confidence-Driven Solution

    cs.AI 2025-08 unverdicted novelty 5.0 of 10

    LLM-as-a-Judge systems report confidence that overstates their accuracy, and the paper's TH-Score plus LLM-as-a-Fuser improves calibration.

  8. An Empirical Study of Evaluating Long-form Question Answering

    cs.IR 2025-04 conditional novelty 5.0 of 10

    In long-form QA, LLM-based evaluators correlate with human judgments better than ROUGE or BERTScore, but they are biased by answer length, question type, self-reinforcement, and rare-word usage, and fine-grained promp...

  9. LLMs to Support a Domain Specific Knowledge Assistant

    cs.CL 2025-02 conditional novelty 5.0 of 10

    A synthetic QA dataset for IFRS sustainability reporting is created with LLMs and used to build and evaluate two QA pipelines, with the fully LLM-based pipeline scoring highest.

  10. Software Engineering for Large Language Models: Research Status, Challenges and the Road Ahead

    cs.SE 2025-06 conditional novelty 4.0 of 10

    A literature review organizes LLM development into a six-phase software engineering lifecycle and identifies challenges and research directions for each phase.

  11. Investigating the Vulnerability of LLM-as-a-Judge Architectures to Prompt-Injection Attacks

    cs.CL 2025-05 conditional novelty 4.0 of 10

    Greedy Coordinate Gradient-optimized suffixes appended to one candidate answer flip Qwen2.5-3B and Falcon3-3B judge verdicts in over 30% of MT-Bench pairwise comparisons.

  12. Reusing Embeddings: Reproducible Reward Model Research in Large Language Model Alignment without GPUs

    cs.CL 2025-02 conditional novelty 4.0 of 10

    Embedding-based reward models reproduce key alignment research findings on CPU-only hardware, lowering cost and improving reproducibility.

  13. RAG Playground: A Framework for Systematic Evaluation of Retrieval Strategies and Prompt Engineering in RAG Systems

    cs.LG 2024-12 conditional novelty 4.0 of 10

    A new RAG evaluation framework reports that hybrid vector-keyword retrieval and structured self-evaluation prompting improve answer quality, reaching a 72.7% pass rate on its own unvalidated metrics.

  14. PPLqa: An Unsupervised Information-Theoretic Quality Metric for Comparing Generative Large Language Models

    cs.CL 2024-11 conditional novelty 4.0 of 10

    PPLqa, the absolute difference between the perplexity of a question plus answer and the answer alone, ranks LLM responses about as well as GPTScore and G-EVAL on one long-form benchmark, with weak overall rank correlation.

  15. Strategic Prompting for Conversational Tasks: A Comparative Analysis of Large Language Models Across Diverse Conversational Tasks

    cs.CL 2024-11 reject novelty 3.0 of 10

    No single open-source LLM among Llama, OPT, Falcon, Alpaca, and MPT performs best across reservation, empathy, counseling, persuasion, and negotiation tasks.

Pith tools