{"id":"a2f7e6f1-950c-4218-aa65-6894f1114c1f","arxiv_id":"2501.00664","paper_version":3,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A new contour-based 'Eden score' compares 2D distributions by density level instead of datapoints and avoids the 'grade inflation' that inflates scores of poor generative model fits.","lead":"This paper shows that common scores for judging whether synthetic data matches real data, like correlation and earth-mover's distance, can give misleadingly high marks to poor fits. It introduces the Eden score, which compares density contours instead of individual points and agrees more closely with human judgment in a small test.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central empirical claim depends on a validation protocol that may manufacture Eden's advantage: fits in Table I are author-labeled, and the human study in Fig.","rationale":"The reader identified annulus comparability as the weakest assumption. That is a genuine technical fragility: Eq. (8) pairs the ith density annulus of p with the ith of q even though the two KDEs may have very different density scales, and the choice of five seaborn contour levels and the 5% tail cutoff are arbitrary. However, I think the more load-bearing issue is empirical. The headline claim is not 'Eden has axioms that guarantee no grade inflation' but 'Eden avoids grade inflation and agrees better with human perception.' The evidence for that claim is: (1) five hand-picked fits in Table I, with quality labels assigned by the authors, and (2) a 20-rater study in which the test pairs were deliberately selected on a score-disagreement criterion and raters were instructed to attend to all contours. Both protocols can inflate Eden's apparent performance. Section II-G explicitly states plots were chosen so that the higher-Eden plot received a lower score on at least one other method; this is case-control enrichment. If Eden is systematically different from equipoint scores, such enrichment makes Eden look better than average-case use would. The rater instruction 'considering all contours' is close to a verbal statement of Eden's equal-weighting premise, so the human gold standard is partially constructed from the hypothesis under test. The paper would need a random or representative study, neutral instructions, and a formal pre-specified definition of grade inflation to support the abstract's broad claims. These are fixable, hence the verdict should remain conditional rather than reject. I also want to credit the paper: the Anscombe/Datasaurus illustrations of correlation-score failure are convincing, the term 'equipoint vs equidensity' is a useful taxonomy, and the negative results for correlation, EMD, Jaccard, and KL on the selected fits are consistent with the literature. The concern here is about external validity, not about internal mathematics; no code/data are provided, which compounds the issue but is not the central attack.","tokens_in":15734,"tokens_out":7597,"duration_ms":79904,"concrete_test":"Run a preregistered human study on a random sample of fit pairs drawn from the same models and datasets, without conditioning on score disagreement, and ask raters a neutral question ('which plot looks like a better fit?') without the 'consider all contours' instruction. Pre-register the primary outcome: median Cohen's kappa for Eden vs. KL and vs. each equipoint score, with bootstrap confidence intervals. If Eden's advantage over KL on these unselected pairs does not exceed a pre-specified margin (e.g., Delta kappa > 0.1) or the score ordering changes, the central claim of superior agreement with human perception fails; also report Eden's ability to separate independently elicited low/high-quality labels on the same pairs.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Grade inflation is never defined formally. The paper's direct evidence that Eden avoids it is Table I: five fits classified as very low/low/high quality by the authors' visual judgment. Eden separates these classes (0.162-0.261 vs 0.853), but the classification is post hoc and the score was designed to reward exactly the contour-overlap features used for the labels. The independent validation (Sec. II-G, Fig. 4) does not resolve this. The 39 pairs were selected so that the higher-Eden plot scored lower on at least one other score (Sec. II-G), enriching for cases where Eden and equipoint scores disagree; raters were explicitly asked to choose the plot 'in which the contours matched better, considering all contours.' Both choices align the gold standard with Eden's construction. Thus the median Cohen's kappa of 0.722 for Eden and the conclusion that equipoint scores suffer grade inflation may not generalize to the population of model comparisons. The paper's broader claim that 'any reasonable equidensity score will avoid grade inflation' is also asserted without a formal argument; the only equidensity exemplar is Eden, so the claim is not tested.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper argues that commonly used 'equipoint' scores (correlation, earth-mover's, Jaccard, and KL) suffer from grade inflation when evaluating two-dimensional synthetic-versus-real distributions, and proposes a new 'equidensity' score, Eden, which averages per-annulus intersection-over-union between matched KDE contour rings (Eqs. 8-9). It illustrates grade inflation with Anscombe's quartet and the Datasaurus Dozen, reports scores for five author-labeled fits in Table I, and validates against 20 human raters on 39 plot pairs in Section II-G. The authors conclude that Eden avoids grade inflation, agrees better with human perception of goodness-of-fit than the equipoint scores, and conjecture that any reasonable equidensity score will avoid grade inflation; they also draw a connection to negative-order Rényi entropy.","tokens_in":15900,"tokens_out":5429,"duration_ms":54626,"significance":"If the central claim held, Eden would be a useful tool for generative-model evaluation, particularly where tail fidelity matters; the paper also usefully documents how correlation-type scores can give near-perfect marks to badly fit models. The validation has genuine strengths: raters were blind to scores, each pair was repeated with rotations, color swaps, and order randomization, agreement was measured per-rater with Cohen's kappa, and the confidence-interval analysis in Section III-G is a thoughtful addition. However, the evidence is not yet strong enough for the abstract's claims: the quality labels in Table I are author-assigned, the validation set is selected and the rater instruction is contour-focused, and the generalization to 'any reasonable equidensity score' is unsupported. Reproducibility would also be improved by a code/data availability statement.","major_comments":[{"comment":"Grade inflation is never defined formally, and the only ground truth for 'deserved' scores in Table I is the authors' visual classification of Fig. 2a-e as very low, low, or high quality. This classification appears to be based on the same contour-overlap features that Eden is explicitly constructed to reward (Eqs. 8-9), so the demonstration that Eden is the only score with a consistent low/high gap lacks independence from the score's construction. Please provide a formal definition of grade inflation and an independent, pre-specified set of labeled fits, or otherwise separate the construction of Eden from the evaluation labels; as written, the central claim rests on five post hoc examples.","section":"Section III, Table I"},{"comment":"The human validation does not support the population-level claim that Eden 'agrees better with human perception' because the 39 pairs were chosen so that the higher-Eden plot received a lower score on at least one other score (Section II-G), and raters were explicitly asked to choose the plot in which 'the contours matched better, considering all contours.' Both design choices favor Eden: the first enriches for Eden-versus-other disagreements, and the second directs attention to the contour-matching quantity Eden is built from. The reported 80% agreement and median kappa should be tested on randomly selected pairs and with a neutral task instruction (for example, 'which plot looks more like the real data?') before the abstract's claim can stand.","section":"Section II-G, Fig. 4"},{"comment":"The statement that 'any reasonable equidensity score will avoid grade inflation' is a conjecture, not a demonstrated result: Eden is the only equidensity score tested, and 'unreasonable' scores are excluded by definition rather than by a formal criterion. A formal argument, or at least a second, differently constructed equidensity score, is required to support this generalization; as written, the Discussion states it as a conclusion despite the authors' own limitation note in Section IV-F that only a small number of scoring methods were investigated.","section":"Abstract and Section IV-A"},{"comment":"The Eden score has several free parameters that are not tested for sensitivity: nannuli = 5 (inherited from seaborn's default), the exclusion of the lowest 5% of probability mass, the 0.1 likelihood threshold in the Jaccard definition, the EMD scaling k = 1, and the Monte Carlo sample size. Section IV-C discusses how to choose nannuli in general terms, but no experiments show that the Table I separation or the human-agreement result is stable across these choices; given that Eqs. 8-9 define the score, at least a coarse sensitivity analysis is needed.","section":"Section II-F and Section IV-C"},{"comment":"The index-paired annulus construction assumes that the ith contour of p corresponds meaningfully to the ith contour of q. When the two distributions have very different concentrations or supports, the same index can pair a high-density region in one distribution with a low-density region in the other, so the averaged IoU no longer measures local density match in the intended way. The manuscript does not test this assumption or identify when it breaks; a synthetic experiment with distributions of different variances or supports would clarify the score's behavior and its limitations.","section":"Section II-F, Eq. (8)"}],"minor_comments":[{"comment":"There are several typographical errors that should be corrected: 'the the two likelihoods' in Section II-D, 'Rényi diverence' in Section II-E, and 'perenially' in the Discussion.","section":"Section II-D, II-E, Discussion"},{"comment":"Reference [24] is missing a title and journal information; the reference list should be completed.","section":"References"},{"comment":"Table I reports standard deviations only for stochastic scores, but the caption's wording could be misread as sampling variability; consider renaming the column 'Monte Carlo SD' or otherwise clarifying that these are not repeat-sampling intervals, which are only addressed in Section III-G and Fig. 6.","section":"Table I"},{"comment":"The manuscript does not state whether code and data are available; given the number of implementation details (seaborn defaults, Monte Carlo sizes, pyemd settings, contour choices), a code and data availability statement would materially improve reproducibility.","section":"Reproducibility"},{"comment":"The connection to negative-order Rényi entropy is speculative and loosely defined; it would be clearer to label it explicitly as a hypothesis for future work rather than a definite finding.","section":"Section IV-D"}],"recommendation":"major_revision","confidential_remarks":"This is a borderline major-revision. The core idea is useful and the human-study design has real strengths, but the abstract overstates the evidence: the Table I labels are post hoc, the validation protocol is selection-biased and instruction-biased toward Eden's construction, and the 'any reasonable equidensity score' claim is unsupported. If the authors add a formal definition of grade inflation, an independent validation on unselected pairs with neutral instructions, and a sensitivity analysis for nannuli and other parameters, the paper could become acceptable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The Eden score is a genuinely new idea, and the paper is worth your time. The equipoint/equidensity distinction is a useful way to frame a real problem: many common evaluation scores reward a generative model for matching only the high-density regions of the data. The authors demonstrate this convincingly with Anscombe, Datasaurus, and real datasets, and their blind human-validation study is real effort, not a toy. The resampling analysis in Fig. 6 is also a practical contribution that deserves more attention.\n\nThat said, the validation has a soft spot that is exactly where the stress-test note lands. The Table I fits are labeled by the authors' visual judgment, and the Eden score was designed to reward the same contour-overlap features used to make those labels. The independent human study is better, but the 39 pairs were selected to enrich for cases where Eden and the equipoint scores disagree, and the rater instruction to compare contours \"considering all contours\" aligns the gold standard with Eden's construction. So the claim that Eden agrees better with human perception is suggestive, not conclusive.\n\nThe broader claim that \"any reasonable equidensity score will avoid grade inflation\" is also asserted, not argued. Only one equidensity exemplar is tested, and the formal connection to Rényi entropy is mentioned but not developed. The sensitivity of the score to nannuli=5, the contour levels, and the Jaccard threshold is not examined, and no code or data are released. These are all addressable, which is why I would not reject the paper. The core idea is plausible and important: a score that weights density regions equally should be less prone to tail neglect than one that weights datapoints equally. The paper just needs to prove that more carefully.\n\nI would send this to peer review, but with the expectation of major revision. The authors should define grade inflation formally, test sensitivity to the arbitrary parameters, include a second equidensity score such as Anderson-Darling, temper the abstract's generalizations, and release code and data. The paper is for anyone who evaluates generative models in low dimensions; it deserves a serious referee, and with revisions it could be a solid contribution.","headline":"The Eden score is a genuinely new idea and the paper is worth engaging, but the validation is weaker than the abstract implies.","tokens_in":16502,"tokens_out":1253,"would_cite":true,"duration_ms":14689,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that standard quality scores for generative models overrate fits that match only high-density regions, and introduces an 'equidensity' score, Eden, that avoids this grade inflation and better matches human judgment.","keywords":["grade inflation","generative models","quality scores","equipoint scores","equidensity scores","Eden score","Kullback-Leibler divergence","Rényi entropy"],"falsifier":"Compute the Eden score with 2, 5, 10, and 20 annuli on the fits in Fig. 2; if the ranking between the low-quality and high-quality fits reverses at any setting, then the claimed immunity to grade inflation depends on the arbitrary contour count rather than on the equidensity principle itself.","tokens_in":15454,"feed_emoji":"📊","tokens_out":8816,"duration_ms":74218,"temperature":0.7,"pith_summary":"The paper argues that several widely used quality scores for comparing two-dimensional distributions—correlation, Jaccard, earth-mover's, and Kullback-Leibler—suffer 'grade inflation': they give high marks to generative models that only match the dense regions of the training data while missing the tails. The paper introduces a new 'equidensity' score, Eden, which weights every density contour equally and, in the paper's tests, avoids grade inflation and agrees with human raters better than the other scores. If correct, the finding would change how synthetic-data quality is measured: evaluations would reward models that capture the whole distribution, not just its peaks. The paper goes further to propose that any reasonable equidensity score will avoid grade inflation, connecting such scores to Rényi entropy of negative order.","feed_headline":"Common quality scores inflate generative-model grades; Eden does not","feed_subtitle":"A new equidensity score, Eden, matches human judgment and avoids rewarding models that fit only the peaks.","key_machinery":"The load-bearing object is the equipoint/equidensity distinction, instantiated in the Eden score. Eden compares two KDEs by slicing each into concentric density annuli from the outermost contour inward (five annuli in this study, with the lowest 5% of probability mass excluded) and computing, for each index $i$, the Jaccard-style intersection-over-union of the $i$-th annulus of one distribution with the $i$-th annulus of the other: $s_i = \\operatorname{Area}(A_i^p \\cap A_i^q) / \\operatorname{Area}(A_i^p \\cup A_i^q)$. The Eden score is the mean of these $s_i$ over all annuli. Because each annulus carries equal weight, a model must match peaks and foothills alike to score well; this equal-density weighting is what distinguishes Eden from equipoint scores and is the mechanism the paper credits for avoiding grade inflation.","core_discovery":"The paper's central claim is that the grade inflation seen in correlation, Jaccard, earth-mover's, and KL scores is not accidental but structural: any score that treats every datapoint equally ('equipoint') will systematically overrate fits that align high-density regions while mismatching low-density regions. As a remedy, the paper defines equidensity scores, in which each density contour contributes equally to the final score, and presents the Eden score as the first example. Eden computes, for each of five density annuli of two kernel density estimates, the intersection-over-union of the corresponding annuli and averages these per-annulus scores. In the paper's experiments, Eden was the only score among the five tested that consistently separated low-quality from high-quality fits, and in a blind study with 20 human raters it agreed with human judgment substantially better than the equipoint scores. The paper also proposes that any reasonable equidensity score will share Eden's immunity to grade inflation, and identifies a formal connection between equidensity weighting and Rényi entropy at negative order.","pith_inferences":["One testable extension the paper leaves implicit: if the equal-weighting principle is correct, then a family of equidensity scores defined with different numbers of annuli should all rank fits identically in the limit of fine contouring; comparing the 5-annulus Eden with a 10-annulus version on the same fits would tell whether the ranking is robust.","The paper's human-validation result suggests but does not prove that human similarity judgments themselves are approximately equidensity; a direct psychophysics experiment varying whether peaks or tails are mismatched could test whether raters truly weight all contours equally.","The connection to negative-order Rényi entropy could be pushed further: if equidensity scores correspond to a negative viewpoint parameter, then the optimal number of annuli might be derivable from a desired tail-sensitivity, rather than chosen ad hoc."],"forward_implications":["If the paper is right, quality scores that count every datapoint equally will systematically overrate generative models that reproduce only the high-density core of the data; users of correlation, Jaccard, earth-mover's, or KL scores should expect such inflation and prefer equidensity scores when tail fidelity matters.","For medical and other applications where rare cases matter, equidensity scores like Eden would be a safer default for model selection, since they penalize mismatches in low-density regions as heavily as mismatches in peaks.","The proposed universality of the equipoint grade-inflation problem implies that other datapoint-weighted scores, such as multidimensional Kolmogorov-Smirnov statistics or FID (being based on earth-mover's distance), are also likely to inflate, so caution is warranted in those settings.","The connection to negative-order Rényi entropy suggests that equidensity scores may be understood as a diversity measure with a negative viewpoint parameter, opening a possible formal framework for designing and comparing further equidensity scores."],"supporting_citations":[{"why":"Provides Anscombe's quartet, the example of different distributions with identical correlation, which motivates the grade-inflation problem for the correlation score.","marker":"[13]"},{"why":"Supplies the Datasaurus Dozen, including the 'dino' dataset, showing how a low-correlation non-random distribution can yield a near-perfect correlation score when compared with random data.","marker":"[14]"},{"why":"Defines the common correlation-score formula (1 - |Rp - Rq|/2) that the paper evaluates as the first example of grade inflation.","marker":"[12]"},{"why":"Provides the earth-mover's distance implementation (pyemd) used to compute the earth-mover's score.","marker":"[25]"},{"why":"Supplies the seaborn KDE construction with levels=5 and Scott's bandwidth, which underlies the Jaccard, KL, and Eden score calculations.","marker":"[22]"},{"why":"Defines the Jaccard index (intersection-over-union) that is the basis of both the Jaccard score and the per-annulus similarity used by Eden.","marker":"[26]"},{"why":"Defines the Kullback-Leibler divergence (as a limit of Rényi divergence) from which the KL score is derived.","marker":"[24]"},{"why":"Provides the entropy/diversity framework connecting Rényi entropies and Hill diversities, supporting the paper's proposed link between equidensity scores and negative-order Rényi entropy.","marker":"[33]"}],"fun_headline_variants":["Equipoint scores overrate; Eden doesn't","Why your generative model score may be too generous","Eden: the first score that resists grade inflation","Stop rewarding generative models that only match peaks","Equidensity scores: the cure for grade inflation"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The score assumes that matching the same-numbered density rings of two distributions is a fair way to compare local fit, which may not hold when the two distributions concentrate in very different ways.","fun_headline_variants_meta":{"raw":{"variants":["Equipoint scores overrate; Eden doesn't","Why your generative model score may be too generous","Eden: the first score that resists grade inflation","Stop rewarding generative models that only match peaks","Equidensity scores: the cure for grade inflation"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000549,"raw_usage":{"total_tokens":2629,"prompt_tokens":960,"completion_tokens":1669,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":576,"completion_tokens_details":{"reasoning_tokens":1595}},"tokens_in":576,"tokens_out":1669,"duration_ms":16744,"temperature":1.0,"reasoning_tokens":1595,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T22:46:01.034089+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compute the Eden score with 2, 5, 10, and 20 annuli on the fits in Fig. 2; if the ranking between the low-quality and high-quality fits reverses at any setting, then the claimed immunity to grade inflation depends on the arbitrary contour count rather than on the equidensity principle itself.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides Anscombe's quartet, the example of different distributions with identical correlation, which motivates the grade-inflation problem for the correlation score."},{"cited_title":"Same Stats, Different Graphs: Generating Datasets with Varied Appearance and Identical Statistics through Simulated Annealing","cited_arxiv_id":null,"evidence_quote":"Supplies the Datasaurus Dozen, including the 'dino' dataset, showing how a low-correlation non-random distribution can yield a near-perfect correlation score when compared with random data."},{"cited_title":"Synthetic Data Metrics , 12 2024","cited_arxiv_id":null,"evidence_quote":"Defines the common correlation-score formula (1 - |Rp - Rq|/2) that the paper evaluates as the first example of grade inflation."},{"cited_title":"Fast and robust earth mover’s distances","cited_arxiv_id":null,"evidence_quote":"Provides the earth-mover's distance implementation (pyemd) used to compute the earth-mover's score."},{"cited_title":"seaborn: statistical data visualization","cited_arxiv_id":null,"evidence_quote":"Supplies the seaborn KDE construction with levels=5 and Scott's bandwidth, which underlies the Jaccard, KL, and Eden score calculations."},{"cited_title":"THE DISTRIBUTION OF THE FLORA IN THE ALPINE ZONE.1","cited_arxiv_id":null,"evidence_quote":"Defines the Jaccard index (intersection-over-union) that is the basis of both the Jaccard score and the per-annulus similarity used by Eden."},{"cited_title":"Kullback and R","cited_arxiv_id":null,"evidence_quote":"Defines the Kullback-Leibler divergence (as a limit of Rényi divergence) from which the KL score is derived."}],"review_version":1}