A systematic analysis of 284 manually reviewed papers plus 1.8k+ others from 2023-2025 reveals under-reporting of human evaluation study design details, creating ambiguity in what was measured and how.
arXiv preprint arXiv:2402.02420 , year=
3 Pith papers cite this work. Polarity classification is still indexing.
representative citing papers
An Explicit Logic Channel of LLM, VFM and probabilistic inference validates and improves zero-shot MLLMs via Consistency Rate without ground-truth labels.
The SLSO framework uses iterative structured output generation, consistency checks, and regeneration to improve GPT-VLM accuracy on jaw cyst findings in panoramic radiographs compared to standard chain-of-thought prompting.
citing papers explorer
-
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation
A systematic analysis of 284 manually reviewed papers plus 1.8k+ others from 2023-2025 reveals under-reporting of human evaluation study design details, creating ambiguity in what was measured and how.
-
Explicit Logic Channel for Validation and Enhancement of MLLMs on Zero-Shot Tasks
An Explicit Logic Channel of LLM, VFM and probabilistic inference validates and improves zero-shot MLLMs via Consistency Rate without ground-truth labels.
-
Generating Findings for Jaw Cysts in Dental Panoramic Radiographs Using a GPT-Based VLM: A Preliminary Study on Building a Two-Stage Self-Correction Loop with Structured Output (SLSO) Framework
The SLSO framework uses iterative structured output generation, consistency checks, and regeneration to improve GPT-VLM accuracy on jaw cyst findings in panoramic radiographs compared to standard chain-of-thought prompting.