Pith. sign in

REVIEW 6 cited by

CompassJudger-1: All-in-one Judge Model Helps Model Evaluation and Evolution

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.16256 v1 pith:65HXWHSU submitted 2024-10-21 cs.CL cs.AI

classification cs.CLcs.AI
keywords evaluationcompassjudger-1judgemodeltaskstextbfvariousall-in-one
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Efficient and accurate evaluation is crucial for the continuous improvement of large language models (LLMs). Among various assessment methods, subjective evaluation has garnered significant attention due to its superior alignment with real-world usage scenarios and human preferences. However, human-based evaluations are costly and lack reproducibility, making precise automated evaluators (judgers) vital in this process. In this report, we introduce \textbf{CompassJudger-1}, the first open-source \textbf{all-in-one} judge LLM. CompassJudger-1 is a general-purpose LLM that demonstrates remarkable versatility. It is capable of: 1. Performing unitary scoring and two-model comparisons as a reward model; 2. Conducting evaluations according to specified formats; 3. Generating critiques; 4. Executing diverse tasks like a general LLM. To assess the evaluation capabilities of different judge models under a unified setting, we have also established \textbf{JudgerBench}, a new benchmark that encompasses various subjective evaluation tasks and covers a wide range of topics. CompassJudger-1 offers a comprehensive solution for various evaluation tasks while maintaining the flexibility to adapt to diverse requirements. Both CompassJudger and JudgerBench are released and available to the research community athttps://github.com/open-compass/CompassJudger. We believe that by open-sourcing these tools, we can foster collaboration and accelerate progress in LLM evaluation methodologies.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MIG: Automatic Data Selection for Instruction Tuning by Maximizing Information Gain in Semantic Space

    cs.CL 2025-04 conditional novelty 7.0 of 10

    MIG greedily selects instruction-tuning data by maximizing a concave information measure over a label graph, and with 5% of Tulu3 data it matches or exceeds full-data SFT performance.

  2. CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards

    cs.CL 2025-07 conditional novelty 6.0 of 10

    A 7B judge model trained with verifiable reward signals and a margin contrastive loss matches the judgment accuracy of models tens of times larger, and a new benchmark JudgerBenchV2 standardizes judge evaluation.

  3. OphthBench: A Comprehensive Benchmark for Evaluating Large Language Models in Chinese Ophthalmology

    cs.CL 2025-02 conditional novelty 6.0 of 10

    OphthBench is a new 591-question Chinese ophthalmology benchmark on which 39 LLMs score around 70% (after normalization), showing a clear gap between current models and clinical readiness.

  4. TrustRoboReward: Preference-Ordered Isotonic Score Editing for Multi-Paradigm Robot Reward Models

    cs.AI 2026-08 conditional novelty 5.0 of 10

    An isotonic label-cleaning step removes score-versus-preference conflicts in robot reward training data, letting a 4B VLM reward model approach GPT-5-mini.

  5. A Survey of Automatic Evaluation Methods on Text, Visual and Speech Generations

    cs.CL 2025-06 conditional novelty 5.0 of 10

    A unified taxonomy and comparative meta-evaluation of automatic evaluation methods across text, vision, and speech generation, concluding that LLM-based evaluators dominate current practice.

  6. Reward Reasoning Model

    cs.CL 2025-05 conditional novelty 4.0 of 10

    Reward models that think before judging, trained via reinforcement learning without human-written reasoning traces, outperform standard reward models and improve with more test-time compute.

Pith tools