REVIEW 3 major objections 11 references
KnowledgeGain metric shows that LLM-selected science news improves reader learning outcomes over baselines.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.3
2026-06-28 22:55 UTC pith:JWSQUBTT
load-bearing objection KnowledgeGain points in a useful direction by targeting actual reader learning instead of similarity, but the human studies lack the details needed to trust the central claims. the 3 major comments →
KnowledgeGain: Evaluating and Optimizing Science News Generation for Reader Learning
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
KnowledgeGain, calculated as the normalized improvement in accuracy on knowledge assessment questions after reading, distinguishes learning from different science media in human readers. Calibrating an LLM simulator on human data allows efficient selection of generated articles that, when evaluated in a second human study, produce superior post-reading accuracy and KnowledgeGain scores relative to baseline generation methods.
What carries the argument
KnowledgeGain is the normalized differential accuracy on pre- and post-reading multiple-choice tests; the prompt-only LLM reader simulator is calibrated to predict this gain and used to filter articles.
Load-bearing premise
That changes in accuracy on pre- and post-reading tests reflect lasting knowledge acquisition instead of short-term memory effects or test familiarity.
What would settle it
A replication study finding no difference in delayed recall (e.g., one week later) between articles selected by the simulator and the baseline.
If this is right
- Science news generation can be optimized directly for reader learning rather than semantic similarity alone.
- LLM simulators provide a scalable way to proxy human evaluation for content selection.
- Generated science news can be tuned to align with knowledge and comprehension objectives in Bloom's Taxonomy.
- Human studies can be reduced by using the simulator for initial ranking before final validation.
Where Pith is reading between the lines
- This method could be applied to evaluate news in other fields such as medicine or technology.
- Future systems might generate explanations tailored to maximize predicted KnowledgeGain for specific readers.
- Combining KnowledgeGain with factual consistency checks could create more effective educational content pipelines.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces KnowledgeGain, a metric for science news quality that quantifies reader knowledge gain via pre- and post-reading accuracy on tests. It reports a first controlled human study showing the metric captures differential knowledge from varied science media, uses those data to calibrate a prompt-only LLM reader simulator, applies the simulator to rank/filter candidate articles, and presents a second human study in which simulator-selected articles yield higher post-reading accuracy and normalized KnowledgeGain than a strong generation baseline. The work positions this as progress toward science news aligned with Bloom's Taxonomy knowledge and comprehension objectives.
Significance. If the pre/post tests validly isolate durable learning and the studies are adequately powered and controlled, KnowledgeGain could offer a useful complement to existing NLG metrics (semantic similarity, factuality) by directly targeting reader outcomes in science communication. The two-study pipeline with LLM-assisted selection demonstrates a practical optimization loop, though its dependence on shared human data limits claims of independent validation.
major comments (3)
- [Abstract / Human Study 1] Abstract and methods description of the first human study: no details are supplied on participant sample size, exclusion criteria, MCQ construction (item novelty, difficulty calibration, parallel forms vs. repeated items, controls for guessing or test familiarity), time delay between pre- and post-tests, or the statistical tests and error bars supporting the differential-knowledge claim. These elements are load-bearing for the assertion that KnowledgeGain 'successfully captures the differential knowledge gained.'
- [LLM Reader Simulator Calibration] LLM simulator section: the simulator is calibrated exclusively on the first study's human data and then used to select articles evaluated in the second human study. This shared data source creates dependence between metric validation and article selection, weakening the independence of the reported improvements in post-reading accuracy and normalized KnowledgeGain.
- [Second Human Study / KnowledgeGain Metric] Evaluation of KnowledgeGain and second study: the central claim that pre/post accuracy scores reflect durable knowledge gain (rather than short-term recall, priming, or repeated-testing effects) rests on unstated assumptions about test design and validity. No evidence is given on question alignment with Bloom's levels, external validation of the tests, or safeguards against artifacts, which directly affects both the metric and the downstream selection benefit.
Simulated Author's Rebuttal
We thank the referee for the thorough and constructive review. The comments highlight important areas for improving methodological transparency and addressing potential limitations in our evaluation pipeline. We respond to each major comment below.
read point-by-point responses
-
Referee: [Abstract / Human Study 1] Abstract and methods description of the first human study: no details are supplied on participant sample size, exclusion criteria, MCQ construction (item novelty, difficulty calibration, parallel forms vs. repeated items, controls for guessing or test familiarity), time delay between pre- and post-tests, or the statistical tests and error bars supporting the differential-knowledge claim. These elements are load-bearing for the assertion that KnowledgeGain 'successfully captures the differential knowledge gained.'
Authors: We agree these details are essential and were omitted from the submitted manuscript. In the revised version we will expand the Human Study 1 section to report the exact participant sample size and demographics, exclusion criteria, MCQ construction process (including item novelty checks, difficulty calibration, parallel forms, and controls for guessing/test familiarity), the time interval between pre- and post-tests, and the statistical tests with error bars used to support the differential-knowledge results. revision: yes
-
Referee: [LLM Reader Simulator Calibration] LLM simulator section: the simulator is calibrated exclusively on the first study's human data and then used to select articles evaluated in the second human study. This shared data source creates dependence between metric validation and article selection, weakening the independence of the reported improvements in post-reading accuracy and normalized KnowledgeGain.
Authors: The dependence is a fair concern. The first study both validates KnowledgeGain and supplies calibration data for the simulator; the second study then applies the simulator to new candidate articles and measures outcomes with fresh participants. While the human evaluation in Study 2 remains independent, we will add an explicit limitations paragraph discussing the shared calibration data and its implications for claims of fully independent validation. revision: partial
-
Referee: [Second Human Study / KnowledgeGain Metric] Evaluation of KnowledgeGain and second study: the central claim that pre/post accuracy scores reflect durable knowledge gain (rather than short-term recall, priming, or repeated-testing effects) rests on unstated assumptions about test design and validity. No evidence is given on question alignment with Bloom's levels, external validation of the tests, or safeguards against artifacts, which directly affects both the metric and the downstream selection benefit.
Authors: We acknowledge that the manuscript does not explicitly document these test-design elements. The questions were written to target Bloom's knowledge and comprehension levels and incorporated safeguards such as varied item formats to reduce priming and repeated-testing effects. In revision we will add a dedicated subsection describing question alignment with Bloom's taxonomy, internal controls against artifacts, and any validation steps performed. External validation against established instruments was not conducted; we will note this as a limitation. revision: partial
Circularity Check
No significant circularity detected
full rationale
The paper's central derivation relies on two distinct human studies. The first study validates the KnowledgeGain metric by showing it captures differential knowledge gain across media types and supplies data to calibrate the LLM simulator. The simulator is then used only for upstream selection of candidate articles; the second human study performs an independent evaluation measuring post-reading accuracy and normalized KnowledgeGain on those selected articles versus a baseline. This chain does not reduce any claimed result to a self-definition, a parameter fitted directly to the evaluation data, or a self-citation that bears the load of the uniqueness or correctness claim. No equations, ansatzes, or renamings of known results are present that would create equivalence by construction. The empirical human evaluations remain external to the fitting step.
Axiom & Free-Parameter Ledger
axioms (2)
- domain assumption Pre- and post-reading accuracy tests validly measure knowledge gain from science news.
- domain assumption An LLM can be prompt-calibrated to reliably simulate human reader knowledge gain.
read the original abstract
Science news is an important medium to communicate discoveries between the research communities and the public. Yet, most metrics for generated or summarized text evaluate semantic similarity and factual consistency, but do not measure how much knowledge readers learn from the news. We introduce KnowledgeGain, a metric that evaluates the quality of science news by measuring how much knowledge readers gained after reading it. To evaluate the metric, we first performed a controlled human study and showed that the metric successfully captures the differential knowledge gained by human readers reading different types of science media. The data allowed us to calibrate a prompt-only LLM reader simulator. We use it to rank and filter candidate articles before human evaluation. A second human study shows that articles selected with this simulator improve post-reading accuracy and normalized KnowledgeGain over a strong generation baseline. Our work is a step toward generating science news that better meets the knowledge and comprehension goals of Bloom's Taxonomy.
Figures
Reference graph
Works this paper leans on
-
[1]
In International conference on machine learning, pages 337–371
Using large language models to simulate mul- tiple humans and replicate human subject studies. In International conference on machine learning, pages 337–371. PMLR. Aliya Amirova, Theodora Fteropoulli, Nafiso Ahmed, Martin R Cowie, and Joel Z Leibo. 2024. Framework- based qualitative analysis of free responses of large language models: Algorithmic fidelit...
2024
-
[2]
Daniel Deutsch and Dan Roth
Towards question-answering as an automatic metric for evaluating the content quality of a sum- mary.TACL. Daniel Deutsch and Dan Roth. 2020. Sacrerouge: An open-source library for using and developing summa- rization evaluation metrics. InNLP-OSS. Esin Durmus, He He, and Mona Diab. 2020. Feqa: A qa evaluation framework for faithfulness assessment in abstr...
2020
-
[3]
Nishanth Madhusudhan, Sathwik Tejaswi Madhusud- han, Vikas Yadav, and Masoud Hashemi
Factual consistency evaluation of summariza- tion in the era of large language models.Expert Systems with Applications, 254:124456. Nishanth Madhusudhan, Sathwik Tejaswi Madhusud- han, Vikas Yadav, and Masoud Hashemi. 2025. Do llms know when to not answer? investigating absten- tion abilities of large language models. InProceed- ings of the 31st Internati...
2025
-
[4]
Bleurt: Learning robust metrics for text gener- ation. InACL. Hillary C Shulman, David M Markowitz, and Todd Rogers. 2024. Reading dies in complexity: Online news consumers prefer simple writing.Science ad- vances, 10(23):eadn2555. Marlene Stoll, Martin Kerwer, Klaus Lieb, and Anita Chasiotis. 2022. Plain language summaries: A sys- tematic review of theor...
2024
-
[5]
Verbalized sampling: How to mitigate mode collapse and unlock llm diversity
Readability of the 100 most-cited neuroimag- ing papers.Frontiers in Human Neuroscience. Ran Yu, Rui Tang, Markus Rokicki, Ujwal Gadiraju, and Stefan Dietze. 2021. Topic-independent modeling of user knowledge in search-as-learning.Information Retrieval Journal. Jiayi Zhang, Simon Yu, Derek Chong, Anthony Si- cilia, Michael R Tomz, Christopher D Manning, a...
work page internal anchor Pith review arXiv 2021
-
[6]
A Human Validation Details This section provides additional details about the controlled human validation study
Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in neural information pro- cessing systems, 36:46595–46623. A Human Validation Details This section provides additional details about the controlled human validation study. Participants an- swered the same comprehension questions before and after reading one of three science communi- cation f...
-
[7]
Please complete the experiment in one continuous sitting
-
[8]
The survey contains three main parts:
Do not use external sources while answering the questions. The survey contains three main parts:
-
[9]
I do not know the answer
Background Knowledge Test: - You will be asked to answer 6 questions for each scientific topic without reading news article. - If you are unsure of an answer, simply select "I do not know the answer."
-
[10]
- You can read the news article as often as needed, but you cannot go back
Reading Task - You will read the news article about each scientific topic. - You can read the news article as often as needed, but you cannot go back
-
[11]
I do not know the answer
Knowledge Gain Test - You will be asked the same 6 questions for each sample to measure how much knowledge you have gained - If you are unsure of an answer, simply select "I do not know the answer." Dimension Score Rubric anchor Accuracy 1 Major factual errors, fabricated details, or topic drift away from the abstract. 2 Contains multiple unsupported addi...
2013
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.