Pith. sign in

REVIEW 3 major objections 6 minor 47 references

Do Current Retrievers Cover All the Evidence? A Controlled Study of Conjunctive Cross-Page Retrieval

T0 review · 3 major / 6 minor · reviewed 2026-07-31 · grok-4.5

Pith's one-line read Finding a relevant long document is not the same as proving it covers every requested piece of evidence.

desk verdict Careful diagnostic that cleanly separates gold discovery from complete-before-subset ranking; the gap is real on their instrument, but inventory false-negatives on blocking subsets are the load-bearing soft spot. read the letter →

arxiv 2607.24165 v2 pith:H4USDCMI submitted 2026-07-27 cs.CV cs.CLcs.IR

classification cs.CVcs.CLcs.IR
keywords conjunctiveretrievalcross-pageevidencediscovery–completiongapconditioncoveragedenselexical–visualfusionrerankinglong-documentIR
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Many real searches ask for a conjunction of evidence spread across different pages of one long document. A system can retrieve a topically relevant document and still fail if a natural partial match—supporting only some of the conditions—ranks higher. This paper builds a controlled test, n-Clue on CrossPage, where each query has all-condition golds and real subset competitors, and success means a top-10 gold beats every released subset. Across seventy configurations, condition-wise decomposition and lexical–visual fusion help, generic rerankers hurt gold placement, and scaling one dense family from 0.6B to 8B does not move complete-first success. The strongest hybrid finds a gold on 81.1% of queries but succeeds complete-first on only 35.8%, and page-aware systems surface stored support for every condition on roughly 5% of queries. The bottleneck is condition coverage and complete-before-subset ordering, not gold discovery alone.

What carries the argument

n-Clue Score@k: a controlled complete-first event that succeeds only when a top-k all-condition gold precedes every released subset qrel, paired with Gold Hit and page-support diagnostics so discovery, ordering, and evidence delivery can be separated.

What would settle it

A retriever that, on the same CrossPage queries and qrels, lifts complete-first success near its Gold Hit rate—especially closing the hybrid’s 81% hit versus 36% complete-first gap—or that surfaces stored support for every condition on far more than about 5% of queries.

Watch

Extended reading notes

Core claim

On explicit conjunctive cross-page requests, representative retrievers often discover all-condition golds without ranking them ahead of natural subset matches. Condition coverage—not gold discovery or within-family scale—is the central bottleneck: the strongest displayed hybrid reaches 81.1% Gold Hit@10 but only 35.8% complete-first success, and page-aware systems deliver all stored support on only 5.1–5.3% of queries.

Load-bearing premise

That ranking a gold before known inventory-labeled subsets on this fixed, mostly World-Bank long-document test is a fair stand-in for real multi-evidence readiness.

Editorial extensions

If this is right

  • Retrieval stacks for multi-part requests need an explicit coverage layer, not only a single relevance score.
  • Condition-wise candidate lists and cross-modal fusion are higher-leverage than scaling one dense encoder alone.
  • Generic rerankers trained for ordinary relevance can worsen complete-before-subset ordering.
  • Training should treat natural grade-(n−1) documents as hard negatives that must lose to full-coverage golds.
  • Page-level systems must match conditions to distinct pages and verify the weakest condition, not only max-pool pages.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • RAG pipelines that assume “hit a relevant doc” equals “evidence is available” will systematically under-serve multi-part legal, policy, and scientific requests.
  • Leaderboards that reward graded partial relevance may rank systems that flood the top with incomplete matches above systems that protect complete-first ordering.
  • A practical next system is a condition–page matrix plus a small verifier that re-queries only the missing condition rather than another end-to-end scalar reranker.
  • The same discovery–completion split likely appears in any corpus where partial matches are common and evidence is page-scattered, not only in this instrument.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The manuscript studies conjunctive cross-page document retrieval: queries with two or three explicit conditions that must be supported on different pages of a single long document. It introduces CrossPage (1,000 queries over 2,021 real documents with a page-grounded condition inventory) and the n-Clue Score@10, which requires a top-10 all-condition gold to precede every released subset qrel. Across 70 configurations, the paper reports a large discovery–completion gap (strongest hybrid: 81.1% Gold Hit@10 vs. 35.8% n-Clue Score@10), gains from condition-wise decomposition (+6.8–7.3 points on two dense backbones) and lexical–visual RRF fusion (+8.7), consistent harm from four generic rerankers, and a null effect of scaling Qwen3-Embedding 0.6B→8B. Robustness is probed via deterministic query renderings, LLM paraphrases, task-factor slices, a four-source SourceShift stress set, blind quote-backed audits with repair, adversarial promotion bounds, and multiplicity-corrected paired inference; code, qrels, rankings, and audit records are released with one-command reproduction. The intended claim is deliberately scoped to explicit, conjunctive, cross-page requests.

Significance. If the results hold, the paper makes a useful diagnostic contribution: it separates gold discovery from complete-before-subset ordering and shows that current retrievers largely fail the latter, with condition coverage — not recall or model size — as the bottleneck. Strengths that raise confidence: a controlled instrument with distinct-page gates and privileged-oracle/floor validity checks; operation-level paired contrasts rather than a leaderboard; blockwise max-|T| plus Holm correction; family bootstrap and crossed dependence analyses; quote-backed blind audits with released records; adversarial promotion bounds; and standard-library one-command reproduction over frozen rankings. These make the claims unusually checkable and give the community a falsifiable instrument plus concrete design targets (coverage tracking, weakest-condition verification).

major comments (3)
  1. [§10.5 vs. Tables 3–4] The headline 81.1/35.8 gap and Table 4's 45.3% 'subset-first' mass depend on subset labels that certify absence by inventory silence (§7; §11: construction reads only fact-matched pages; §10.2.2 states the asymmetry explicitly). The paper's own SourceShift targeted challenge (§10.5) promoted 90/125 (72%) of high-ranked judged subsets to gold, and the CrossPage stratified audit revised 20/100 sampled pairs (19 promotions). No equivalent retrieval-conditioned audit exists for CrossPage's blocking subsets. Because subset-to-gold false negatives deflate n-Clue Score while leaving Gold Hit untouched, label error inflates the measured gap by an unquantified amount. Please run the §10.5-style challenge on (a sample of) subsets in the clean top-10 union of the 13 scorecard systems and report corrected score ranges.
  2. [Table S1h / Table 5] Table S1h shows a single adversarial judged-subset promotion per query flips every positive contrast (e.g., complementarity +8.7 has bound −24.6 at k=1). As a worst-case bound this is expected, but combined with the empirical promotion rates above it leaves the signs of the Table 5 controlled contrasts unquantified under plausible label noise. Since the contrasts are paired on shared qrels, random promotions need not flip signs; please add a plausible-rate sensitivity analysis — recompute contrasts under subset-to-gold promotions sampled at audit-estimated rates, or report per-contrast break-even promotion rates. This is executable with the released machinery.
  3. [§4.3 / Abstract vs. Tables S14, S1i] The abstract and §4.3 state that operation directions 'replicate on a four-source stress set.' The n-Clue directions do replicate on final qrels (Table 7), but the frozen SourceShift fusion−BM25 gold-NDCG contrast that motivated the stress set collapsed from +.110 to +.016 (p=.381) after the targeted challenge, with both crossed intervals including zero (Tables S14, S1i). The main text should state this outcome: it bears directly on how much weight the stress-set replication can carry, and on the label-noise question in the first comment.
minor comments (6)
  1. [Eq. (1)] The indicator symbol renders as ⊮ in the PDF; check the glyph or define the notation explicitly.
  2. [§4.4 / Table 8] Gold+Support is reported only for Qwen3-VL 2B/8B. Clarify whether ColQwen2.5 page scores were unusable for this metric, and soften the abstract's 'page-aware visual systems' wording, which generalizes from one model family.
  3. [Table 1] PMC contributes 1 of 2,176 gold qrels; the two-source framing slightly overstates diversity (§7's 'World-Bank-centered' is the accurate emphasis). The 'Single-page-solvable golds 0 / 2,176' row formatting is also ambiguous on first read.
  4. [Table 3, footnote] The dagger notes post-hoc selection of the hybrid on the raw matrix. It would defuse this concern to state that the same system is also the top raw-form fair row (Table S1a, 31.1), so the selection is among near-equivalent variants rather than cherry-picked.
  5. [§3.1] Gold+Support credits only the three highest-scoring surfaced pages, which is exactly tight for n=3 queries; a sweep over the number of surfaced pages would clarify how much of the 5.1–5.3% is metric strictness versus genuine delivery failure.
  6. [Table 6] Some cells are small (n=146, 185). The text notes the slices are descriptive, but adding interval estimates or explicit cell-size cautions would help readers avoid over-reading single cells.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: empirical measurement study; scores are observed indicator events on frozen rankings, not quantities forced by definition or self-citation.

full rationale

The paper’s load-bearing claims are comparative measurements (Gold Hit vs n-Clue Score@10; paired base→variant deltas for scale, decomposition, fusion, reranking) on a fixed instrument with frozen qrels and released rankings. n-Clue Score is an explicit indicator event—top-10 gold precedes every released subset—not a fitted parameter renamed as a prediction, and different systems obtain different scores (e.g., BM25-AND 26.8 vs hybrid 35.8), so the discovery–completion gap is not true by construction. Controls separate instrument from systems: privileged construction-fact oracles reach 100 while single-page floors score 0; systems never receive construction facts or evidence-page annotations. Post-hoc selection of the displayed hybrid is acknowledged as descriptive and does not define the primary estimand. There is no self-citation uniqueness theorem, ansatz smuggled from author prior work, or renaming of a known closed-form result. Inventory false-negative risk (SourceShift promotions) is a construct-validity concern, not circular derivation. Conclusion: self-contained empirical contrasts; circularity score 0.

Assumptions & free parameters 4 free parameters · 5 assumptions · 3 invented entities

Load-bearing structure is experimental design, not fitted theory. The claim rests on task formalization (conjunctive distinct-page golds vs subset competitors), inventory-based labeling, the complete-first@10 rule, and frozen off-the-shelf retrievers. Free choices are metric/depth/fusion hyperparameters and corpus construction gates rather than constants fit to prove the gap.

free parameters (4)
  • n-Clue success depth k=10 = k=10
    Primary estimand requires a gold in top-10 before every released subset; changing k would rescale absolute scores though the gap narrative could persist.
  • RRF fusion hyperparameters = equal weights, RRF k=60, depth 100
    Equal-weight RRF with k=60 and depth 100 is fixed without tuning on labels; still a design choice that affects the hybrid endpoint.
  • SourceShift per-source OCR quality gate = 98% per-source page extraction
    Pre-specified 98% extraction gate drops four sources and shapes the stress-set distribution.
  • Displayed decomposition/fusion endpoints = max raw-form n-Clue within block
    For multi-variant blocks, the paper reports the highest raw-form n-Clue endpoint post hoc while keeping all variants in the supplement; inference uses blockwise selection correction.
assumptions (5)
  • domain assumption A document is gold iff the inventory supports every condition with a distinct-page assignment; subsets are natural corpus documents with 1≤g<n.
    Defines the complete-vs-partial competitors central to the estimand (Study Design §3.1–3.2).
  • domain assumption Inventory-zero query–document pairs are grade 0; complete-first success is relative to the versioned inventory, not exhaustive human labels of all pairs.
    Stated in Scope/Limitations; false negatives could in principle alter ordering diagnostics.
  • ad hoc to paper Graded partial-credit NDCG is the wrong primary estimand for this task because it rewards ranking many subsets.
    Motivated with BM25-OR vs AND contrast; reasonable but a normative evaluation choice, not a theorem.
  • domain assumption Frozen inference-only public retrievers and equal-weight fusion without label-tuned weights are fair probes of current practice.
    Underpins the ‘current retrievers’ claim; excludes trained coverage-aware agents and many proprietary systems.
  • standard math Bootstrap over exact-gold-set families and paired randomization adequately handle query dependence for the reported intervals/tests.
    Standard resampling practice; supplement also reports crossed document-weight stresses.
invented entities (3)
  • n-Clue Score@k (complete-first success)
    purpose: Primary metric separating ranking a full-coverage gold before known subset matches from mere gold discovery.
    Defined in Eq. (1); operationalizes the discovery–completion gap.
  • CrossPage / n-Clue measurement instrument independent evidence
    purpose: Controlled corpus+qrel instrument with page-grounded condition inventories and natural subset competitors.
    Constructed artifact; validated via audits and controls rather than external prior existence.
  • Gold+Support@k page-evidence delivery metric
    purpose: Conservative check whether top-k golds surface stored support pages for every condition.
    Secondary diagnostic showing ~5% full support delivery for page-scored visual systems.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Do Current Retrievers Cover All the Evidence? A Controlled Study of Conjunctive Cross-Page Retrieval." pith.science (2026). https://pith.science/paper/H4USDCMI

@misc{pith2026260724165,
  author       = {Pith},
  title        = {Pith review of: Do Current Retrievers Cover All the Evidence? A Controlled Study of Conjunctive Cross-Page Retrieval},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/H4USDCMI}},
  note         = {Machine review of arXiv:2607.24165}
}
read the original abstract

Finding a long document relevant to a multi-part request is not the same as establishing that it contains every requested piece of evidence. We study this gap for conjunctive document retrieval, where two or three explicit conditions must be supported on different pages of one document. We use n-Clue as a controlled measurement instrument: 1{,}000 queries over 2{,}021 documents pair all-condition golds with naturally occurring documents that satisfy only a subset, and a complete-first success requires a top-10 gold to precede every released subset qrel. Across 70 configurations, condition-wise decomposition improves two dense backbones by 6.8--7.3 points and lexical--visual fusion adds 8.7, while four generic rerankers all reduce Gold-NDCG; these directions replicate on a four-source stress set. Scaling one dense family from 0.6B to 8B changes complete-first success by 0.0 points. The strongest displayed hybrid illustrates the resulting gap: it finds a gold for 81.1\% of queries but succeeds complete-first on only 35.8\%, and the gap persists across condition count, target length, candidate density, query rendering, and the four-source stress set. Finally, page-aware visual systems surface stored support for every condition on only 5.1--5.3\% of queries. These results identify condition coverage, rather than gold discovery alone, as the central bottleneck.

Figures

Figures reproduced from arXiv: 2607.24165 by the authors.

Figure 1
Figure 1. Discovery and completion are different retrieval [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Raw-form diagnostic matrix. (a) All 54 fair methods under Gold Hit@10 and n-Clue Score@10. (b) Controlled [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

47 extracted references · 4 canonical work pages

  1. [1]

    NeurIPS Datasets and Benchmarks Track , year=

    Thakur, Nandan and Reimers, Nils and R. NeurIPS Datasets and Benchmarks Track , year=

  2. [2]

    EACL , year=

    Muennighoff, Niklas and Tazi, Nouamane and Magne, Lo. EACL , year=

  3. [3]

    doi:10.18653/v1/2025.findings-emnlp.726 , year=

    Lu, Xuan and Liu, Sifan and Yin, Bochao and Li, Yongqi and Chen, Xinghao and Su, Hui and Jin, Yaohui and Zeng, Wenjun and Shen, Xiaoyu , booktitle=. doi:10.18653/v1/2025.findings-emnlp.726 , year=

  4. [4]

    doi:10.1609/aaai.v40i40.40706 , year=

    Xu, Ganlin and Yin, Zhitao and Zhang, Linghao and Liang, Jiaqing and Lu, Weijia and Zhang, Xiaodong and Yang, Zhifei and Jiang, Sihang and Yang, Deqing , booktitle=. doi:10.1609/aaai.v40i40.40706 , year=

  5. [5]

    Proceedings of NAACL-HLT (Volume 1: Long Papers) , pages=

    Multi-Conditional Ranking with Large Language Models , author=. Proceedings of NAACL-HLT (Volume 1: Long Papers) , pages=. doi:10.18653/v1/2025.naacl-long.146 , year=

  6. [6]

    doi:10.18653/v1/2025.emnlp-main.1576 , year=

    Dong, Kuicai and Chang, Yujing and Goh Xin Deik, Derrick and Li, Dexun and Tang, Ruiming and Liu, Yong , booktitle=. doi:10.18653/v1/2025.emnlp-main.1576 , year=

  7. [7]

    Proceedings of the 31st International Conference on Computational Linguistics , pages=

    Counting-Stars: A Multi-evidence, Position-aware, and Scalable Benchmark for Evaluating Long-Context Large Language Models , author=. Proceedings of the 31st International Conference on Computational Linguistics , pages=

  8. [8]

    doi:10.18653/v1/2026.findings-acl.187 , year=

    Yang, Shouqing and Zhang, Qi and Yang, Yuhang and Xu, Ruikang and Hou, Yuwei and Jia, Zhulin and Gao, Lirong and Wang, Haobo and Chen, Jinglei and Wang, Jiexiang and Guo, Sheng and Zheng, Bo and Chen, Gang , booktitle=. doi:10.18653/v1/2026.findings-acl.187 , year=

Show all 47 references
  1. [9]

    doi:10.18653/v1/2026.acl-long.2129 , year=

    Zheng, Yuanlei and Fu, Pei and Li, Hang and Wang, Ziyang and Zhang, Yuyi and Ruan, Wenyu and Zhang, Xiaojin and Wei, Zhongyu and Luo, Zhenbo and Luan, Jian and Chen, Wei and Bai, Xiang , booktitle=. doi:10.18653/v1/2026.acl-long.2129 , year=

  2. [10]

    Nguyen, Tri and Rosenberg, Mir and Song, Xia and Gao, Jianfeng and Tiwary, Saurabh and Majumder, Rangan and Deng, Li , journal=

  3. [11]

    Yang, Zhilin and Qi, Peng and Zhang, Saizheng and Bengio, Yoshua and Cohen, William W and Salakhutdinov, Ruslan and Manning, Christopher D , booktitle=

  4. [12]

    Trivedi, Harsh and Balasubramanian, Niranjan and Khot, Tushar and Sabharwal, Ashish , journal=

  5. [13]

    Constructing A Multi-hop

    Ho, Xanh and Nguyen, Anh-Khoa Duong and Sugawara, Saku and Aizawa, Akiko , booktitle=. Constructing A Multi-hop

  6. [14]

    ICLR , year=

    Faysse, Manuel and Sibille, Hugues and Wu, Tony and Omrani, Bilel and Viaud, Gautier and Hudelot, C. ICLR , year=

  7. [15]

    Khattab, Omar and Zaharia, Matei , booktitle=

  8. [16]

    Santhanam, Keshav and Khattab, Omar and Saad-Falcon, Jon and Potts, Christopher and Zaharia, Matei , booktitle=

  9. [17]

    SIGIR , year=

    Formal, Thibault and Piwowarski, Benjamin and Clinchant, St. SIGIR , year=

  10. [18]

    Chen, Jianlv and Xiao, Shitao and Zhang, Peitian and Luo, Kun and Lian, Defu and Liu, Zheng , booktitle=

  11. [19]

    arXiv preprint arXiv:2212.03533 , year=

    Text Embeddings by Weakly-Supervised Contrastive Pre-training , author=. arXiv preprint arXiv:2212.03533 , year=

  12. [20]

    arXiv preprint arXiv:2308.03281 , year=

    Towards General Text Embeddings with Multi-stage Contrastive Learning , author=. arXiv preprint arXiv:2308.03281 , year=

  13. [21]

    Zhang, Yanzhao and Li, Mingxin and Long, Dingkun and Zhang, Xin and Lin, Huan and Yang, Baosong and Xie, Pengjun and Yang, An and Liu, Dayiheng and Lin, Junyang and Huang, Fei and Zhou, Jingren , journal=

  14. [22]

    Lee, Chankyu and Roy, Rajarshi and Xu, Mengyao and Raiman, Jonathan and Shoeybi, Mohammad and Catanzaro, Bryan and Ping, Wei , booktitle=

  15. [23]

    arXiv preprint arXiv:2506.18902 , year=

    G. arXiv preprint arXiv:2506.18902 , year=

  16. [24]

    Transactions on Machine Learning Research , year=

    Unsupervised Dense Information Retrieval with Contrastive Learning , author=. Transactions on Machine Learning Research , year=

  17. [25]

    EMNLP , year=

    Unifying Multimodal Retrieval via Document Screenshot Embedding , author=. EMNLP , year=

  18. [26]

    SIGIR , year=

    Reciprocal Rank Fusion Outperforms Condorcet and Individual Rank Learning Methods , author=. SIGIR , year=

  19. [27]

    Findings of EMNLP , year=

    Document Ranking with a Pretrained Sequence-to-Sequence Model , author=. Findings of EMNLP , year=

  20. [28]

    Sun, Weiwei and Yan, Lingyong and Ma, Xinyu and Wang, Shuaiqiang and Ren, Pengjie and Chen, Zhumin and Yin, Dawei and Ren, Zhaochun , booktitle=. Is

  21. [29]

    The Probabilistic Relevance Framework:

    Robertson, Stephen and Zaragoza, Hugo , journal=. The Probabilistic Relevance Framework:

  22. [30]

    EMNLP , year=

    Dense Passage Retrieval for Open-Domain Question Answering , author=. EMNLP , year=

  23. [31]

    ICLR , year=

    Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text Retrieval , author=. ICLR , year=

  24. [32]

    Retrieval-Augmented Generation for Knowledge-Intensive

    Lewis, Patrick and Perez, Ethan and Piktus, Aleksandra and Petroni, Fabio and Karpukhin, Vladimir and Goyal, Naman and others , booktitle=. Retrieval-Augmented Generation for Knowledge-Intensive

  25. [33]

    The Annals of Statistics , volume=

    Bootstrap Methods: Another Look at the Jackknife , author=. The Annals of Statistics , volume=

  26. [34]

    and Tang, Michael and Sun, Ruoxi and Yoon, Jinsung and Arik, Sercan O

    Su, Hongjin and Yen, Howard and Xia, Mengzhou and Shi, Weijia and Muennighoff, Niklas and Wang, Han-yu and Liu, Haisu and Shi, Quan and Siegel, Zachary S. and Tang, Michael and Sun, Ruoxi and Yoon, Jinsung and Arik, Sercan O. and Chen, Danqi and Yu, Tao , booktitle=

  27. [35]

    Weller, Orion and Lawrie, Dawn and Van Durme, Benjamin , booktitle=

  28. [36]

    Weller, Orion and Chang, Benjamin and MacAvaney, Sean and Lo, Kyle and Cohan, Arman and Van Durme, Benjamin and Lawrie, Dawn and Soldaini, Luca , booktitle=

  29. [37]

    Song, Tingyu and Gan, Guo and Shang, Mingsheng and Zhao, Yilun , booktitle=

  30. [38]

    EMNLP , pages=

    Fact or Fiction: Verifying Scientific Claims , author=. EMNLP , pages=

  31. [39]

    Wadden, David and Lo, Kyle and Kuehl, Bailey and Cohan, Arman and Beltagy, Iz and Wang, Lucy Lu and Hajishirzi, Hannaneh , booktitle=

  32. [40]

    Zhu, Dawei and Wang, Liang and Yang, Nan and Song, Yifan and Wu, Wenhao and Wei, Furu and Li, Sujian , booktitle=

  33. [41]

    Cumulated Gain-Based Evaluation of

    J. Cumulated Gain-Based Evaluation of. ACM Transactions on Information Systems , volume=

  34. [42]

    NeurIPS Datasets and Benchmarks Track , year=

    Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks , author=. NeurIPS Datasets and Benchmarks Track , year=

  35. [43]

    Journal of the London Mathematical Society , volume=

    On Representatives of Subsets , author=. Journal of the London Mathematical Society , volume=

  36. [44]

    Journal of Computer and System Sciences , volume=

    Optimal Aggregation Algorithms for Middleware , author=. Journal of Computer and System Sciences , volume=

  37. [45]

    Educational and Psychological Measurement , volume=

    A Coefficient of Agreement for Nominal Scales , author=. Educational and Psychological Measurement , volume=

  38. [46]

    Psychological Bulletin , volume=

    Measuring Nominal Scale Agreement among Many Raters , author=. Psychological Bulletin , volume=

  39. [47]

    Communications of the ACM , volume=

    Datasheets for Datasets , author=. Communications of the ACM , volume=

Pith tools

Reviewed July 31, 2026 · model on record in the stance chip above.