REVIEW 4 major objections 6 minor 147 references
"All that Glitters": Approaches to Evaluations with Unreliable Model and Human Annotations
T0 review · 4 major / 6 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read Standard agreement metrics, applied to unreliable human labels, can make encoder models look super-human at rating classroom teaching; stricter psychometric measures reveal spurious correlations and nonrandom bias in model and human raters.
desk verdict Brings psychometric tools NLP should know, but its central validity measure likely overcorrects on its own assumption; deserves peer review with major revisions. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing device is Eq. (4), the disattenuated convergent correlation: the observed correlation between human and model ratings of the same teacher on different lessons, divided by the square root of the product of the two rater families' generalizability coefficients $E\rho^2$. Because low reliability mechanically shrinks observed correlations, this correction separates "the two raters track the same underlying construct" from "the correlation is an artifact of measurement error." The generalizability coefficient $E\rho^2$ itself—the share of rating variance attributable to the teacher rather than to lesson, segment, rater, and item facets—is estimated from a nested random-effects model $R \times (S:O:I)$ and carries the confidence and decision-study analyses. Rater-level behavior is quantified by a hierarchical rater model whose item-response-theory stage estimates ideal scores and whose signal-detection stage assigns each rater a bias parameter $\phi$ and a variability parameter $\psi^2$, extended with race covariates to test fairness. These pieces recombine in the decision study of Eq. (9), which reweights variance components to predict reliability under proposed human-in-the-loop designs.
What would settle it
Apply the same six-metric protocol to a held-out second cohort of classroom transcripts rated by multiple experts, or re-run Eq. (4) using only same-week lessons of each teacher. If the near-1.0 disattenuated correlations shrink or exceed 1.0 when lessons are close in time, the stability assumption fails and the corrected correlations are artifacts of the correction; if encoder superiority on $E\rho^2$, the spurious-item pattern, and the GPT negative-bias trend fail to reproduce on new transcripts, the discovery is an artifact of this dataset rather than a property of the methods.
Extended reading notes
Core claim
The central claim is that when expert human labels are unreliable, standard inter-rater statistics—percent agreement, Cohen's $\kappa$, quadratic weighted kappa, Pearson, Spearman, and Kendall correlations, ICCs—can certify a model as "super-human" while masking what the model actually learned. On those measures, the paper's five encoder models outperform the best human raters on nearly every item of the MQI and CLASS instruments. Under Generalizability Theory (the $E\rho^2$ generalizability coefficient and the $\Phi$ dependability coefficient), which attribute rating variance to teacher, lesson, segment, and rater facets, the encoder advantage shrinks and reverses on items like teacher explanations (EXPL) and student explanations (STEXPL). Disattenuating human-model correlations by the reliabilities of each rater family shows that some apparent model-human agreement is spurious, while items with near-1.0 corrected correlations indicate genuine shared signal only if a teacher's latent ability is stable across lessons. A hierarchical rater model—an item-response-theory true-score stage feeding a signal-detection rater stage—estimates each rater's leniency/severity bias $\phi$ and consistency $\psi^2$, and conditioning those on teacher race shows a negative bias trend against Black teachers in the GPT family, smaller but nonrandom biases in encoders, and detectable racial biases in some human raters on negatively worded items. Closing decision studies estimate that a three-encoder ensemble observing whole class periods can double the reliability achievable for the rare-construct item LANGIMP relative to ten fifteen-minute human visits, at a large time saving, while GPT ensembles would reduce human rating reliability in human-in-the-loop use.
Load-bearing premise
The near-perfect disattenuated correlations assume a teacher's true instructional ability stays about the same from lesson to lesson, so that a human rating one lesson and a model rating another are measuring the same thing; if ability shifts between lessons, the corrected correlation cannot distinguish shared construct from lesson-to-lesson variability.
Editorial extensions
If this is right
- Reporting only concordance metrics can certify a model as super-human against an unreliable human baseline, so high-stakes annotation evaluations should also report generalizability and dependability.
- Encoder models trained on transcripts can, for specific MQI items such as LANGIMP, deliver rating reliability roughly double what ten short human visits achieve, at a large time saving.
- GPT-style prompt-engineered ratings of classroom instruction currently underperform expert humans on nearly every metric, and decision-study estimates predict their variance would lower human rating reliability in human-in-the-loop use.
- Nonrandom racial bias is detectable at the individual-rater level even at very low label reliabilities: GPT models show a negative bias trend against Black teachers, encoders show smaller but real biases on sparse items, and some human raters show racial differences on negatively worded items.
- The evaluation protocol transfers to other NLP tasks with unreliable expert annotations: g-studies identify label weaknesses before training, disattenuation tests whether model-human correlations reflect shared constructs, and d-studies estimate the value of model assistance before costly trials are run.
Reading between the lines
- The same six-part protocol would transfer to other benchmarks whose "gold" labels come from a small pool of expensive experts, such as essay scoring, clinical chart review, or content moderation; reliability-corrected model rankings on such benchmarks would likely differ from the raw-correlation leaderboards currently reported.
- A direct test of the spuriousness story would retrain the encoder family with speaker-identity markers added to the transcripts: the paper's design deliberately withheld speaker information, and its claim that the EXPL and STEXPL failures stem from that omission predicts that those items' generalizability would recover once speaker roles are visible.
- The human-in-the-loop decision studies offer a small, cheap falsifiable pilot: run a handful of classrooms where principals rate under their usual 15-minute visits while an encoder scores full periods in the background, then compare achieved reliability and time spent against the predicted curves; the claimed savings of roughly two to ten hours per teacher are currently extrapolations, not measure
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper argues that common concordance metrics (correlation, percent agreement, kappa) can be misleading when human annotations are unreliable, and it demonstrates a set of psychometric tools—generalizability theory, disattenuated correlations, hierarchical rater models, fairness analyses, and decision studies—for evaluating model and human annotations of classroom teaching quality. Using the NCTE dataset, the authors compare human expert ratings, GPT-family ratings from a prior study, and five newly trained transformer encoder models. They report that encoder models appear to achieve state-of-the-art or 'super-human' concordance under standard metrics, but that more rigorous generalizability and validity analyses reveal spurious correlations and racial biases, while GPT models perform poorly and would worsen human-in-the-loop reliability.
Significance. If the qualitative conclusions hold, the paper makes a valuable methodological contribution to NLP evaluation: it demonstrates how to quantify label quality, bias, and validity when human ratings are noisy, and it provides an applied case study with public code and data links. The replication of the NCTE g-study, the use of hierarchical rater models to disentangle rater bias from construct signal, and the explicit treatment of limitations (including the acknowledgment that the encoder models are trained on the same noisy human labels they are later evaluated against) are genuine strengths. The central qualitative claim—that standard metrics can mask model and label quality—is supportable, but several load-bearing numerical claims are fragile as currently presented, particularly the disattenuated-correlation evidence for spuriousness and the magnitude of the 'super-human' performance.
major comments (4)
- [§5.3 and Appendix G (Eq. 13)] The disattenuated correlation in Eq. (4) rests on the assumption stated in §5.3.1: 'If an individual teacher's latent instructional ability theta_i is about the same from lesson to lesson with the same students.' The paper's own hierarchical rater model in Eq. (13) of Appendix G explicitly models lesson-level latent abilities theta_oi varying around teacher-level Theta_i, and the g-study of Eq. (1)–(2) treats lesson-within-teacher variance (nu_o:i) as error. If theta actually varies across lessons, the numerator of Eq. (4) mixes construct overlap with lesson-to-lesson instability, while the denominator divides by an E_rho2 estimate that includes that instability as error, so the correction overcorrects. The near-1.0 disattenuated correlations in Figure 3 and Figure 4(c) (e.g., EXPL at 1.0†) may then be artifacts of the correction rather than evidence that humans and models track the same stable construct. This threatens the paper's second stated contribution (Section 1, 'methods for detection of spurious correlations via disattenuating low human-model correlations'). A concrete test would be to split items or teachers by the magnitude of lesson-within-teacher variance and show that the disattenuated correlations are not driven by high-variance items, or to estimate the model of Eq. (13) and use the teacher-level variance component in the denominator.
- [§5.1, Tables 1 and 9] The headline claim that encoder models achieve 'super-human' results across all classroom annotation tasks (abstract, §5.1.3) is not adequately supported because no majority-class baseline is reported. On highly imbalanced items such as MGEN, MAJERR, and LANGIMP, where the dominant score category accounts for a large fraction of labels (see Figure 6 and Figure 8), a trivial classifier that always predicts the modal category would achieve high percent agreement and respectable kappa values. For example, Table 9 shows encoder percent agreement of 0.95–0.96 on MGEN; without a modal-class baseline, this value is uninterpretable. The generalizability and HRM analyses partially mitigate this concern, but the quantitative 'super-human' framing is load-bearing for the paper's message. Please report majority-class baselines and chance-corrected agreement measures (e.g., prevalence-adjusted kappa) for all items.
- [Table 2 and §5.2.2] The generalizability and dependability estimates E_R^2 and Phi in Table 2 are reported as point estimates without confidence intervals or other uncertainty quantification, despite the low reliability levels (many values between 0.00 and 0.20) and small item counts. For instance, the human-vs-encoder differences on EXPL (0.15 vs 0.00) and SMQR (0.14 vs 0.09) in Table 2 may be within sampling error. Since Section 5.3.2 uses these exact values to declare correlations 'spurious' (e.g., EXPL and STEXPL), the absence of uncertainty intervals makes the numeric conclusions fragile. Please provide bootstrap or Bayesian intervals for the variance components and derived coefficients.
- [Abstract and §7 (Limitations)] The encoder models were trained on the same human ratings that are used later in the evaluation, so the 'super-human' concordance result in the abstract partly reflects the models fitting the evaluation target rather than an independent assessment of quality. The paper acknowledges this in Section 7 ('the signal is still trained on noisy human ratings'), and the psychometric analyses are motivated by this dependency, but the abstract's unqualified statement 'the encoder family of models achieve state-of-the-art, even "super-human", results across all classroom annotation tasks' overstates the finding. Recommend rewording the abstract to indicate that the models achieve state-of-the-art concordance with the training labels, with the caveat that this concordance is not evidence of construct validity.
minor comments (6)
- [§4] The text reads 'Encoder models were trained a single GPU in Google Colab'; 'a' should be 'on a single GPU'.
- [Table 1 caption] The caption describes the last metric as 'Kendall's concordance correlation'; this should be 'Kendall's tau rank correlation coefficient'.
- [§5.3.1 and footnote 12] The sentence 'Disattenuated correlations of 1.0 do not mean perfect correlation: it generally means that measurement error is not randomly distributed' appears twice (in the main text and in footnote 12) and is confusing; consider rephrasing or deleting one occurrence.
- [Appendix D.1.1] The citation 'Gao, 2022' for SimCSE refers to a reference in the bibliography that is not the SimCSE paper (the listed Shuai Gao entry is about a French-Mongolian MT system). The correct reference is the SimCSE paper by Gao et al. (2021). Similarly, in Table 7 the E5 embedding model is cited as 'Wang et al. (2022)', but the bibliography entry points to a Jiarui Wang et al. paper on text style transfer, not the E5 embeddings paper. These citation errors should be corrected.
- [Figure 4 caption] Panel (b) is labeled 'Reliabilities' but includes metrics such as percent agreement and percent agreement ±1, which are agreement indices rather than reliability coefficients; consider using a more precise label such as 'Agreement metrics'.
- [§5.1.3] The sentence 'Using nearly any standardized combination of metrics across all items from Section 5.1, Encoder models perform better than the single highest performing expert human rater' is broader than what Table 9 shows; for items like MMETH and STEXPL, human correlations are higher than the encoder family values. Suggest qualifying this as 'on average across items' or specifying the exception pattern.
Circularity Check
No significant circularity: encoder evaluation uses a disjoint held-out test set; psychometric corrections are separate estimates; remaining concerns are validity assumptions, not circular reductions.
full rationale
After walking the derivation chain, I find no circular step that reduces a claimed result to its own inputs. The encoder 'super-human' concordance result is a supervised-learning evaluation against a lesson-stratified held-out test set: Section 4 states that 'all model outputs in this study were conducted with a lesson-level-stratified held-out test set (see Figure 8) that was not used during model development,' and the Limitations passage that 'the signal is still trained on noisy human ratings' is a data-quality caveat, not an identity between the training target and the evaluation target. The g-theory generalizability estimates (Eqs. 1-3) and the disattenuated correlations (Eq. 4) are separate quantities: the numerator is a cross-family, cross-lesson correlation, while the denominator uses variance components estimated from each family's own ratings; neither is defined in terms of the other, so no self-definitional reduction is present. The HRM bias and fairness results are estimated jointly from the rating data by MCMC with stated priors, and the GPT ratings are imported from an external study (Wang and Demszky, 2023), not from a self-citation chain. The only self-citations (Hardy 2021 as an example of QWK use, and a note about a forthcoming paper) are illustrative and not load-bearing. The Eq. 4 lesson-stability assumption is a substantive identification assumption whose violation would threaten validity, but that is a correctness or robustness concern, not circularity: the formula does not reduce to its inputs by construction.
Assumptions & free parameters
free parameters (4)
- encoder hyperparameters (dropout, attention heads, epochs, learning rate, weight decay) =
dropout=0.75, heads=32, epochs=15-75, lr=2.5e-5, weight decay=0.0003
- G-theory variance components for Eρ² and Φ per item per rater family =
posterior variance estimates from lme4 (not tabulated as single numbers)
- HRM rater bias φ and variability ψ per rater per item per race =
posterior means from MCMC (not listed in main text)
- Decision study design counts (segments, observations, raters) =
n_s=6, n_o=1, n_r=1 for HIL; human 15 minutes, model 45 minutes
assumptions (5)
- domain assumption Latent teaching ability theta is multivariate normal with zero mean and identity covariance (Eq. 5), and teacher-year latent abilities are normal (Eq. 13).
- domain assumption A teacher's latent instructional ability is approximately stable across lessons with the same students (Section 5.3.1).
- domain assumption Local independence: ratings are conditionally independent given the true score, rater, and item (Eqs. 5-8).
- domain assumption G-theory random effects: raters, observations, and segments are sampled from a defined universe, and variance components generalize to that universe (Section 5.2.1).
- domain assumption For the fairness analysis, teacher race is the only covariate influencing rater bias beyond item and rater, with no unmeasured confounders (Section 5.5.1).
Cite this review
Pith. "Pith review of "All that Glitters": Approaches to Evaluations with Unreliable Model and Human Annotations." pith.science (2026). https://pith.science/paper/YAJC7B5R
@misc{pith2026241115634,
author = {Pith},
title = {Pith review of: "All that Glitters": Approaches to Evaluations with Unreliable Model and Human Annotations},
year = {2026},
howpublished = {\url{https://pith.science/paper/YAJC7B5R}},
note = {Machine review of arXiv:2411.15634}
}
read the original abstract
"Gold" and "ground truth" human-mediated labels have error. The effects of this error can escape commonly reported metrics of label quality or obscure questions of accuracy, bias, fairness, and usefulness during model evaluation. This study demonstrates methods for answering such questions even in the context of very low reliabilities from expert humans. We analyze human labels, GPT model ratings, and transformer encoder model annotations describing the quality of classroom teaching, an important, expensive, and currently only human task. We answer the question of whether such a task can be automated using two Large Language Model (LLM) architecture families--encoders and GPT decoders, using novel approaches to evaluating label quality across six dimensions: Concordance, Confidence, Validity, Bias, Fairness, and Helpfulness. First, we demonstrate that using standard metrics in the presence of poor labels can mask both label and model quality: the encoder family of models achieve state-of-the-art, even "super-human", results across all classroom annotation tasks. But not all these positive results remain after using more rigorous evaluation measures which reveal spurious correlations and nonrandom racial biases across models and humans. This study then expands these methods to estimate how model use would change to human label quality if models were used in a human-in-the-loop context, finding that the variance captured in GPT model labels would worsen reliabilities for humans influenced by these models. We identify areas where some LLMs, within the generalizability of the current data, could improve the quality of expensive human ratings of classroom instruction.
Figures
Figures from the paper (9 more)
Reference graph
Works this paper leans on
-
[1]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRING...
-
[2]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...
-
[3]
Gavin Abercrombie, Verena Rieser, and Dirk Hovy. 2023. https://doi.org/10.48550/arXiv.2301.10684 Consistency is Key : Disentangling Label Variation in Natural Language Processing with Intra - Annotator Agreement . arXiv preprint. ArXiv:2301.10684 [cs]
-
[4]
Adams, Mark Wilson, and Wen-chung Wang
Raymond J. Adams, Mark Wilson, and Wen-chung Wang. 1997. https://doi.org/10.1177/0146621697211001 The multidimensional random coefficients multinomial logit model . Applied Psychological Measurement, 21(1):1--23. Place: US Publisher: Sage Publications
-
[5]
Julius Adebayo, Justin Gilmer, Michael Muelly, Ian Goodfellow, Moritz Hardt, and Been Kim. 2020. https://doi.org/10.48550/arXiv.1810.03292 Sanity Checks for Saliency Maps . arXiv preprint. ArXiv:1810.03292 [cs, stat]
-
[6]
Nikhil Agarwal, Alex Moehring, Pranav Rajpurkar, and Tobias Salz. 2023. https://doi.org/10.3386/w31422 Combining Human Expertise with Artificial Intelligence : Experimental Evidence from Radiology
doi:10.3386/w31422 2023
-
[7]
Elena Aguilar. 2013. Developing a Work Plan : How Do I Determine What to Do ? In The art of coaching: effective strategies for school transformation, pages 119--144. Jossey-Bass, A Wiley Brand, San Francisco
2013
-
[8]
Sterling Alic, Dorottya Demszky, Zid Mancenido, Jing Liu, Heather Hill, and Dan Jurafsky. 2022. https://doi.org/10.18653/v1/2022.bea-1.27 Computationally identifying funneling and focusing questions in classroom discourse . In Proceedings of the 17th Workshop on Innovative Use of NLP for Building Educational Applications (BEA 2022), pages 224--233, Seattl...
Show all 147 references
-
[9]
Amos Azaria, Rina Azoulay, and Shulamit Reches. 2024. https://doi.org/10.1162/dint_a_00235 ChatGPT is a Remarkable Tool — For Experts . Data Intelligence, 6(1):240--296
2024 doi
- [10]
- [11]
-
[12]
Chin, Thomas J
Andrew Bacher-Hicks, Mark J. Chin, Thomas J. Kane, and Douglas O. Staiger. 2017. https://doi.org/10.3386/w23478 An Evaluation of Bias in Three Measures of Teacher Quality : Value - Added , Classroom Observations , and Student Surveys
2017 doi
-
[13]
Chin, Thomas J
Andrew Bacher-Hicks, Mark J. Chin, Thomas J. Kane, and Douglas O. Staiger. 2019. https://doi.org/10.1016/j.econedurev.2019.101919 An experimental evaluation of three teacher quality measures: Value -added, classroom observations, and student surveys . Economics of Education Re...
2019
-
[14]
Paul Bambrick-Santoyo. 2016. Get better faster: a 90-day plan for coaching new teachers. Jossey-Bass, A Wiley Brand, San Francisco, CA
2016
-
[15]
Paul Bambrick-Santoyo. 2018. Leverage leadership 2.0: a practical guide to building exceptional schools. Jossey-Bass, San Francisco, CA
2018
-
[16]
Douglas Bates, Martin Mächler, Ben Bolker, and Steve Walker. 2015. https://doi.org/10.18637/jss.v067.i01 Fitting Linear Mixed - Effects Models Using lme4 . Journal of Statistical Software, 67:1--48
2015 doi
-
[17]
Williamson, and and Robert J
Isaac 1 Bejar, David M. Williamson, and and Robert J. Mislevy. 2006. Human Scoring . In Automated Scoring of Complex Tasks in Computer - Based Testing . Routledge. Num Pages: 34
2006
-
[18]
Howcroft
Anya Belz, Simon Mille, and David M. Howcroft. 2020. https://doi.org/10.18653/v1/2020.inlg-1.24 Disentangling the Properties of Human Evaluation Methods : A Classification System to Support Comparability , Meta - Evaluation and Reproducibility Testing . In Proceedings of the 1...
2020 doi
-
[19]
Anya Belz, Craig Thomson, Ehud Reiter, and Simon Mille. 2023. https://doi.org/10.18653/v1/2023.findings-acl.226 Non- Repeatable Experiments and Non - Reproducible Results : The Reproducibility Crisis in Human Evaluation in NLP . In Findings of the Association for Computational...
2023 doi
-
[21]
Eckstein, Noémi Éltető, Thomas L
Marcel Binz, Elif Akata, Matthias Bethge, Franziska Brändle, Fred Callaway, Julian Coda-Forno, Peter Dayan, Can Demircan, Maria K. Eckstein, Noémi Éltető, Thomas L. Griffiths, Susanne Haridi, Akshay K. Jagadish, Li Ji-An, Alexander Kipnis, Sreejan Kumar, Tobias Ludwig, Marvin ...
- [22]
-
[23]
David Blazar. 2018. https://doi.org/10.1162/edfp_a_00251 Validating Teacher Effects on Students ’ Attitudes and Behaviors : Evidence from Random Assignment of Teachers to Students . Education Finance and Policy, 13(3):281--309
2018 doi
-
[24]
Charalambous, and Heather C
David Blazar, David Braslow, Charalambos Y. Charalambous, and Heather C. Hill. 2017. https://doi.org/10.1080/10627197.2017.1309274 Attending to General and Mathematics - Specific Dimensions of Teaching : Exploring Factors Across Two Observation Instruments . Educational Assess...
2017
-
[25]
David Blazar and Cynthia Pollard. 2022. https://www.edworkingpapers.com/ai22-591 Challenges and Tradeoffs of “ Good ” Teaching : The Pursuit of Multiple Educational Outcomes . Technical report, Annenberg Institute at Brown University. Publication Title: EdWorkingPapers.com
2022
-
[26]
Robert L. Brennan. 2001 a . https://doi.org/10.1007/978-1-4757-3456-0 Generalizability Theory . Springer, New York, NY
2001 doi
-
[27]
Robert L. Brennan. 2001 b . https://doi.org/10.1007/978-1-4757-3456-0_6 Variability of Statistics in Generalizability Theory . In Robert L. Brennan, editor, Generalizability Theory , Statistics for Social Sciences and Public Policy , pages 179--213. Springer, New York, NY
2001 doi
-
[28]
Robert L. Brennan. 2013. Generalizability Theory . Springer Science & Business Media. Google-Books-ID: nbHbBwAAQBAJ
2013
-
[29]
Briggs and Mark Wilson
Derek C. Briggs and Mark Wilson. 2007. https://doi.org/10.1111/j.1745-3984.2007.00031.x Generalizability in item response modeling . Journal of Educational Measurement, 44(2):131--155. Place: United Kingdom Publisher: Blackwell Publishing
2007
-
[30]
Casabianca
Jodi M. Casabianca. 2021. https://doi.org/10.1111/emip.12478 Digital Module 27: Hierarchical Rater Models . Educational Measurement: Issues and Practice, 40(4):103--104. \_eprint: https://onlinelibrary.wiley.com/doi/pdf/10.1111/emip.12478
2021 doi
-
[31]
Casabianca, Daniel F
Jodi M. Casabianca, Daniel F. McCaffrey, Drew H. Gitomer, Courtney A. Bell, Bridget K. Hamre, and Robert C. Pianta. 2013. https://doi.org/10.1177/0013164413486987 Effect of Observation Mode on Measures of Secondary Mathematics Teaching . Educational and Psychological Measureme...
2013 doi
-
[32]
Charalambous and Seán Delaney
Charalambos Y. Charalambous and Seán Delaney. 2019. https://doi.org/10.1163/9789004418875_014 13 Mathematics Teaching Practices and Practice - Based Pedagogies . Brill. Section: International Handbook of Mathematics Teacher Education: Volume 1
2019 doi
-
[33]
Eric P. Charles. 2005. https://doi.org/10.1037/1082-989X.10.2.206 The Correction for Attenuation Due to Measurement Error : Clarifying Concepts and Creating Confidence Sets . Psychological Methods, 10(2):206--226. Place: US Publisher: American Psychological Association
2005 doi
- [34]
-
[35]
Chengyu Cui, Chun Wang, and Gongjun Xu. 2024. https://doi.org/10.1007/s11336-024-09955-8 Variational Estimation for Multidimensional Generalized Partial Credit Model . Psychometrika
2024 doi
- [36]
-
[37]
Linda Darling-Hammond. 2014. https://scholarworks.umb.edu/nejpp/vol26/iss1/4 What Can PISA Tell Us about U . S . Education Policy ? New England Journal of Public Policy, 26(1)
2014
-
[38]
Linda Darling-Hammond, Lisa Flook, Channa Cook-Harvey, Brigid Barron, and David Osher. 2020. https://doi.org/10.1080/10888691.2018.1537791 Implications for educational practice of the science of learning and development . Applied Developmental Science, 24(2):97--140. Publisher...
2020
-
[39]
Lawrence T. Decarlo. 2003. https://doi.org/10.3758/BF03195496 Using the PLUM procedure of SPSS to fit unequal variance and generalized signal detection models . Behavior Research Methods, Instruments, & Computers, 35(1):49--56
2003 doi
-
[40]
Lawrence T. DeCarlo. 2008. https://doi.org/10.1002/j.2333-8504.2008.tb02149.x Studies of a Latent - Class Signal - Detection Model for Constructed - Response Scoring . ETS Research Report Series, 2008(2):i--55. \_eprint: https://onlinelibrary.wiley.com/doi/pdf/10.1002/j.2333-8...
2008
-
[41]
Lawrence T. DeCarlo. 2023. https://doi.org/10.1111/jedm.12358 Classical Item Analysis from a Signal Detection Perspective . Journal of Educational Measurement, 60(3):520--547. \_eprint: https://onlinelibrary.wiley.com/doi/pdf/10.1111/jedm.12358
2023 doi
-
[42]
DeCarlo, YoungKoung Kim, and Matthew S
Lawrence T. DeCarlo, YoungKoung Kim, and Matthew S. Johnson. 2011. https://www.jstor.org/stable/23018150 A Hierarchical Rater Model for Constructed Responses , with a Signal Detection Rater Model . Journal of Educational Measurement, 48(3):333--356. Publisher: National Council...
2011
- [43]
-
[44]
Dorottya Demszky and Heather Hill. 2023. https://doi.org/10.18653/v1/2023.bea-1.44 The NCTE Transcripts : A Dataset of Elementary Math Classroom Transcripts . In Proceedings of the 18th Workshop on Innovative Use of NLP for Building Educational Applications ( BEA 2023) , pages...
2023 doi
-
[46]
Hill, Shyamoli Sanghi, and Ariel Chung
Dorottya Demszky, Jing Liu, Heather C. Hill, Shyamoli Sanghi, and Ariel Chung. 2023. https://edworkingpapers.com/ai23-875 Improving Teachers ’ Questioning Quality through Automated Feedback : A Mixed - Methods Randomized Controlled Trial in Brick -and- Mortar Classrooms . Tech...
2023
- [47]
-
[48]
Dorottya Demszky, Rose Wang, Sean Geraghty, and Carol Yu. 2024. https://doi.org/10.1145/3636555.3636924 Does Feedback on Talk Time Increase Student Engagement ? Evidence from a Randomized Controlled Trial on a Math Tutoring Platform . In Proceedings of the 14th Learning Analyt...
2024
-
[49]
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. https://doi.org/10.18653/v1/N19-1423 BERT : Pre -training of Deep Bidirectional Transformers for Language Understanding . In Proceedings of the 2019 Conference of the North American Chapter of the Associat...
2019 doi
- [50]
-
[51]
Donnelly, Nathaniel Blanchard, Andrew M
Patrick J. Donnelly, Nathaniel Blanchard, Andrew M. Olney, Sean Kelly, Martin Nystrand, and Sidney K. D'Mello. 2017. https://doi.org/10.1145/3027385.3027417 Words matter: automatic detection of teacher questions in live classroom discourse using linguistics, acoustics, and con...
2017
-
[52]
Cynthia Dwork, Moritz Hardt, Toniann Pitassi, Omer Reingold, and Richard Zemel. 2012. https://doi.org/10.1145/2090236.2090255 Fairness through awareness . In Proceedings of the 3rd Innovations in Theoretical Computer Science Conference , ITCS '12, pages 214--226, New York, NY,...
2012
-
[53]
Detecting Illusory Halo Effects in Rater - Mediated Assessment : A Mixture Rasch Facets Modeling Approach
Thomas Eckes and Kuan-Yu Jin. Detecting Illusory Halo Effects in Rater - Mediated Assessment : A Mixture Rasch Facets Modeling Approach
-
[54]
Anjalie Field, Su Lin Blodgett, Zeerak Waseem, and Yulia Tsvetkov. 2021. https://doi.org/10.18653/v1/2021.acl-long.149 A Survey of Race , Racism , and Anti - Racism in NLP . In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th...
2021 doi
- [55]
-
[56]
Shuai Gao. 2022. https://aclanthology.org/2022.jeptalnrecital-recital.8 Syst \`e me de traduction automatique neuronale fran c ais-mongol (historique, mise en place et \'e valuations) ( F rench- M ongolian neural machine translation system (history, implementation, and evaluat...
2022
- [57]
-
[58]
Gordon, Michelle S
Mitchell L. Gordon, Michelle S. Lam, Joon Sung Park, Kayur Patel, Jeff Hancock, Tatsunori Hashimoto, and Michael S. Bernstein. 2022. https://doi.org/10.1145/3491102.3502004 Jury Learning : Integrating Dissenting Voices into Machine Learning Models . In Proceedings of the 2022 ...
2022
-
[59]
Jason Grissom, Susanna Loeb, and Benjamin Master. 2013. https://cepa.stanford.edu/content/effective-instructional-time-use-school-leaders-longitudinal-evidence-observations-principals Effective Instructional Time Use for School Leaders : Longitudinal Evidence from Observations...
2013
-
[60]
Louis Guttman. 1945. https://doi.org/10.1007/BF02288892 A basis for analyzing test-retest reliability . Psychometrika, 10(4):255--282
1945 doi
-
[61]
Zaretta Hammond. 2015. Culturally responsive teaching and the brain: promoting authentic engagement and rigor among culturally and linguistically diverse students. Corwin, a SAGE company, Thousand Oaks, California. OCLC: ocn889185083
2015
- [62]
- [63]
-
[64]
Ursula Hebert-Johnson, Michael Kim, Omer Reingold, and Guy Rothblum. 2018. https://proceedings.mlr.press/v80/hebert-johnson18a.html Multicalibration: Calibration for the ( Computationally - Identifiable ) Masses . In Proceedings of the 35th International Conference on Machine ...
2018
-
[65]
Peter Henderson, Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen, Ruoxi Jia, and Prateek Mittal. 2024. Safety Risks from Customizing Foundation Models via Fine -tuning
2024
-
[66]
Juyeon Heo, Christina Heinze-Deml, Oussama Elachqar, Shirley Ren, Udhay Nallasamy, Andy Miller, Kwan Ho Ryan Chan, and Jaya Narain. 2024. https://arxiv.org/abs/2410.14516v1 Do LLMs "know" internally when they follow instructions?
2024 arXiv
-
[67]
Hill, Merrie L
Heather C. Hill, Merrie L. Blunk, Charalambos Y. Charalambous, Jennifer M. Lewis, Geoffrey C. Phelps, Laurie Sleep, and Deborah Loewenberg Ball. 2008. https://www.jstor.org/stable/27739893 Mathematical Knowledge for Teaching and the Mathematical Quality of Instruction : An Exp...
2008
-
[68]
Hill, Charalambos Y
Heather C. Hill, Charalambos Y. Charalambous, David Blazar, Daniel McGinn, Matthew A. Kraft, Mary Beisiegel, Andrea Humez, Erica Litke, and Kathleen Lynch. 2012 a . https://doi.org/10.1080/10627197.2012.715019 Validating Arguments for Observational Instruments : Attending to M...
2012
-
[69]
Hill, Charalambos Y
Heather C. Hill, Charalambos Y. Charalambous, and Matthew A. Kraft. 2012 b . https://doi.org/10.3102/0013189X12437203 When Rater Reliability Is Not Enough : Teacher Observation Systems and a Case for the Generalizability Study . Educational Researcher, 41(2):56--64. Publisher:...
2012 doi
-
[70]
Ho and Thomas J
Andrew D. Ho and Thomas J. Kane. 2013. https://eric.ed.gov/?id=ED540957 The Reliability of Classroom Observations by School Personnel . Research Paper . MET Project . Technical report, Bill & Melinda Gates Foundation. Publication Title: Bill & Melinda Gates Foundation ERIC Num...
2013
-
[71]
Valentin Hofmann, Pratyusha Ria Kalluri, Dan Jurafsky, and Sharese King. 2024 a . https://doi.org/10.1038/s41586-024-07856-5 AI generates covertly racist decisions about people based on their dialect . Nature, pages 1--8. Publisher: Nature Publishing Group
2024 doi
- [72]
- [73]
-
[74]
Amin Hosseiny Marani, Joshua Levine, and Eric P.S. Baumer. 2022. https://doi.org/10.1145/3511808.3557410 One Rating to Rule Them All ? Evidence of Multidimensionality in Human Assessment of Topic Labeling Quality . In Proceedings of the 31st ACM International Conference on Inf...
2022
-
[75]
Qi (Helen) Huang and Daniel M. Bolt. 2023. https://doi.org/10.3758/s13428-023-02275-2 Unipolar IRT and the Author Recognition Test ( ART ) . Behavior Research Methods
2023 doi
-
[76]
Jacobs, Ryan J
Cassandra L. Jacobs, Ryan J. Hubbard, and Kara D. Federmeier. 2022. https://aclanthology.org/2022.scil-1.22 Masked language models directly encode linguistic uncertainty . In Proceedings of the Society for Computation in Linguistics 2022, pages 225--228, online. Association fo...
2022
-
[77]
Xuejun (Ryan) Ji. 2023. https://doi.org/10.14288/1.0437518 Using cross-classified mixed effects model for validation studies : a flexible and pragmatic validation method . Ph.D. thesis, University of British Columbia
2023 doi
-
[78]
Irina Jurenka, Markus Kunesch, Kevin R McKee, Daniel Gillick, Shaojian Zhu, Shubham Milind Phal, Katherine Hermann, Daniel Kasenberg, Avishkar Bhoopchand, Ankit Anand, Miruna Pîslar, Stephanie Chan, Lisa Wang, Jennifer She, Parsa Mahmoudieh, Wei-Jen Ko, Andrea Huber, Brett Wil...
2024
-
[79]
Thomas Kane, Heather Hill, and Douglas Staiger. 2015. https://doi.org/10.3886/ICPSR36095.V4 National Center for Teacher Effectiveness Main Study : Version 4
2015 doi
-
[80]
Kane, Daniel F
Thomas J. Kane, Daniel F. McCaffrey, Trey Miller, and Douglas O. Staiger. 2013. https://eric.ed.gov/?id=ED540959 Have We Identified Effective Teachers ? Validating Measures of Effective Teaching Using Random Assignment . Research Paper . MET Project . Technical report, Bill & ...
2013
-
[81]
Kane and Douglas O
Thomas J. Kane and Douglas O. Staiger. 2012. https://eric.ed.gov/?id=ED540960 Gathering Feedback for Teaching : Combining High - Quality Observations with Student Surveys and Achievement Gains . Research Paper . MET Project . Technical report, Bill & Melinda Gates Foundation. ...
2012
-
[82]
Maximilian Kasy and Rediet Abebe. 2021. https://doi.org/10.1145/3442188.3445919 Fairness, Equality , and Power in Algorithmic Decision - Making . In Proceedings of the 2021 ACM Conference on Fairness , Accountability , and Transparency , pages 576--586, Virtual Event Canada. ACM
2021
-
[83]
Gabriella Kazai, Jaap Kamps, and Natasa Milic-Frayling. 2013. https://doi.org/10.1007/s10791-012-9205-0 An analysis of human factors and label accuracy in crowdsourcing relevance judgments . Information Retrieval, 16(2):138--178
2013 doi
-
[84]
Olney, Patrick Donnelly, Martin Nystrand, and Sidney K
Sean Kelly, Andrew M. Olney, Patrick Donnelly, Martin Nystrand, and Sidney K. D’Mello. 2018. https://doi.org/10.3102/0013189X18785613 Automatically Measuring Question Authenticity in Real - World Classrooms . Educational Researcher, 47(7):451--464. Publisher: American Educatio...
2018 doi
- [85]
- [86]
- [87]
-
[88]
David Klahr. 2013. https://doi.org/10.1073/pnas.1212738110 What do we mean? On the importance of not abandoning scientific rigor when talking about science education . Proceedings of the National Academy of Sciences, 110(supplement\_3):14075--14080. Publisher: Proceedings of t...
2013 doi
-
[89]
Kromrey, Robert H
J. Kromrey, Robert H. Fay, and Aarti P. Bellara. 2008. https://www.semanticscholar.org/paper/Macro-for-Computing-Confidence-Intervals-for-Kromrey-Fay/62c7827b2d2ebb01dd9cd78a757001513f335141 Macro for Computing Confidence Intervals for Disattenuated Correlation Coefficients
2008
-
[90]
Doug Lemov. 2021. Teach like a champion 3.0: 63 techniques that put students on the path to college, third edition edition. Jossey-Bass, a Wiley imprint, Hoboken, NJ
2021
-
[91]
Doug Lemov and Norman Atkins. 2015. Teach like a champion 2.0: 62 techniques that put students on the path to college, second edition edition. Jossey-Bass, San Francisco, CA
2015
- [92]
-
[93]
Peter Liljedahl, Tracy Johnston Zager, and Laura Wheeler. 2021. Building thinking classrooms in mathematics: 14 teaching practices for enhancing learning: Grades K -12 . Corwin Mathematics . Corwin, Thousand Oaks, California London New Delhi Singapore
2021
-
[94]
Jing Liu and Julie Cohen. 2021. https://doi.org/10.3102/01623737211009267 Measuring Teaching Practices at Scale : A Novel Application of Text -as- Data Methods . Educational Evaluation and Policy Analysis, 43(4):587--614. Publisher: American Educational Research Association
2021 doi
-
[95]
Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang
Nelson F. Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang. 2023 a . https://doi.org/10.48550/arXiv.2307.03172 Lost in the Middle : How Language Models Use Long Contexts . arXiv preprint. ArXiv:2307.03172 [cs]
- [96]
- [97]
- [98]
-
[99]
French, and Helen Patrick
Panayota Mantzicopoulos, Brian F. French, and Helen Patrick. 2018. https://doi.org/10.1080/10409289.2018.1477903 The Mathematical Quality of Instruction ( MQI ) in Kindergarten : An Evaluation of the Stability of the MQI Using Generalizability Theory . Early Education and Deve...
2018
-
[100]
Mariano and Brian W
Louis T. Mariano and Brian W. Junker. 2007. https://doi.org/10.3102/1076998606298033 Covariates of the Rating Process in Hierarchical Models for Multiple Ratings of Test Items . Journal of Educational and Behavioral Statistics, 32(3):287--314
2007 doi
-
[101]
Thomas McCoy, Shunyu Yao, Dan Friedman, Matthew Hardy, and Thomas L
R. Thomas McCoy, Shunyu Yao, Dan Friedman, Matthew Hardy, and Thomas L. Griffiths. 2023. https://arxiv.org/abs/2309.13638v1 Embers of Autoregression : Understanding Large Language Models Through the Problem They are Trained to Solve
2023 arXiv
-
[102]
Samuel Messick. 1998. https://www.jstor.org/stable/27522333 Test Validity : A Matter of Consequence . Social Indicators Research, 45(1/3):35--44. Publisher: Springer
1998
-
[103]
Muchinsky
Paul M. Muchinsky. 1996. https://doi.org/10.1177/0013164496056001004 The Correction for Attenuation . Educational and Psychological Measurement, 56(1):63--75. Publisher: SAGE Publications Inc
1996 doi
-
[104]
Eiji Muraki. 1992. https://doi.org/10.1177/014662169201600206 A Generalized Partial Credit Model : Application of an EM Algorithm . Applied Psychological Measurement, 16(2):159--176. Publisher: SAGE Publications Inc
1992 doi
-
[105]
Murphy and S
Daniel L. Murphy and S. Natasha Beretvas. 2015. https://doi.org/10.1080/08957347.2015.1042158 A Comparison of Teacher Effectiveness Measures Calculated Using Three Multilevel Models for Raters Effects . Applied Measurement in Education, 28(3):219--236. Publisher: Routledge \_e...
2015
-
[106]
You Gotta be a Doctor , Lin
Huy Nghiem, John Prindle, Jieyu Zhao, and Hal Daumé III. 2024. https://doi.org/10.48550/arXiv.2406.12232 " You Gotta be a Doctor , Lin ": An Investigation of Name - Based Bias of Large Language Models in Employment Recommendations . arXiv preprint. ArXiv:2406.12232
-
[107]
Patz, Brian W
Richard J. Patz, Brian W. Junker, Matthew S. Johnson, and Louis T. Mariano. 2002. https://www.jstor.org/stable/3648122 The Hierarchical Rater Model for Rated Test Items and Its Application to Large - Scale Educational Assessment Data . Journal of Educational and Behavioral Sta...
2002
-
[108]
Pianta and Bridget K
Robert C. Pianta and Bridget K. Hamre. 2009. https://doi.org/10.3102/0013189X09332374 Conceptualization, Measurement , and Improvement of Classroom Processes : Standardized Observation Can Leverage Capacity . Educational Researcher, 38(2):109--119. Publisher: American Educatio...
2009 doi
-
[109]
Pianta, Karen M
Robert C. Pianta, Karen M. La Paro, and Bridget K. Hamre. 2008. Classroom Assessment Scoring System ( CLASS ) Manual , K -3 . Paul H. Brookes Publishing Company. Google-Books-ID: NBeaGgAACAAJ
2008
- [110]
-
[111]
Martyn Plummer. 2003. JAGS : A program for analysis of Bayesian graphical models using Gibbs sampling. Working Papers
2003
- [112]
-
[113]
Marco Tulio Ribeiro, Tongshuang Wu, Carlos Guestrin, and Sameer Singh. 2020. https://doi.org/10.18653/v1/2020.acl-main.442 Beyond Accuracy : Behavioral Testing of NLP Models with CheckList . In Proceedings of the 58th Annual Meeting of the Association for Computational Linguis...
2020 doi
-
[114]
Rickford and Sharese King
John R. Rickford and Sharese King. 2016. https://doi.org/10.1353/lan.2016.0078 Language and linguistics on trial: Hearing Rachel Jeantel (and other vernacular speakers) in the courtroom and beyond . Language, 92(4):948--988
2016
- [115]
-
[116]
Olney, Sean Kelly, Martin Nystrand, Sidney D'Mello, Nathan Blanchard, Xiaoyi Sun, Marcy Glaus, and Art Graesser
Borhan Samei, Andrew M. Olney, Sean Kelly, Martin Nystrand, Sidney D'Mello, Nathan Blanchard, Xiaoyi Sun, Marcy Glaus, and Art Graesser. 2014. https://eric.ed.gov/?id=ED566380 Domain Independent Assessment of Dialogic Properties of Classroom Discourse . Technical report. Publi...
2014
- [117]
-
[118]
Jon Saphier, Mary Ann Haley-Speca, and Robert Gower. 2008. The skillful teacher: building your teaching skills, 6th ed edition. Research for Better Teaching, Acton, Mass
2008
-
[119]
Schwartz, Jessica M
Daniel L. Schwartz, Jessica M. Tsang, and Kristen P. Blair. 2016. The ABCs of how we learn: 26 scientifically proven approaches, how they work, and when to use them , first edition edition. Norton books in education. W.W. Norton & Company, New York
2016
-
[120]
Mark D. Shermis. 2014. https://doi.org/10.1016/j.asw.2013.04.001 State-of-the-art automated essay scoring: Competition , results, and future directions from a United States demonstration . Assessing Writing, 20:53--76
2014 doi
- [121]
-
[122]
Robert E. Slavin. 2002. https://doi.org/10.3102/0013189X031007015 Evidence- Based Education Policies : Transforming Educational Practice and Research . Educational Researcher, 31(7):15--21. Publisher: American Educational Research Association
2002 doi
- [123]
-
[124]
Mukund Sundararajan, Ankur Taly, and Qiqi Yan. 2017. https://proceedings.mlr.press/v70/sundararajan17a.html Axiomatic Attribution for Deep Networks . In Proceedings of the 34th International Conference on Machine Learning , pages 3319--3328. PMLR. ISSN: 2640-3498
2017
-
[125]
Martin, and Tamara Sumner
Abhijit Suresh, Jennifer Jacobs, Charis Harty, Margaret Perkoff, James H. Martin, and Tamara Sumner. 2022. https://aclanthology.org/2022.lrec-1.497 The T alk M oves dataset: K-12 mathematics lesson transcripts annotated for teacher and student discursive moves . In Proceedings...
2022
- [126]
-
[127]
https://www.r-project.org/ R: A Language and Environment for Statistical Computing
R Core Team. https://www.r-project.org/ R: A Language and Environment for Statistical Computing
- [128]
-
[129]
UpLevel. 2024. https://resources.uplevelteam.com/gen-ai-for-coding Gen AI for Coding Research Report . Technical report, Uplevel Data Labs
2024
- [130]
-
[131]
Chris van der Lee, Albert Gatt, Emiel van Miltenburg, Sander Wubben, and Emiel Krahmer. 2019. https://doi.org/10.18653/v1/W19-8643 Best practices for the human evaluation of automatically generated text . In Proceedings of the 12th International Conference on Natural Language ...
2019 doi
-
[132]
Jiarui Wang, Richong Zhang, Junfan Chen, Jaein Kim, and Yongyi Mao. 2022. https://doi.org/10.18653/v1/2022.emnlp-main.521 Text style transferring via adversarial masking and styled filling . In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Process...
2022 doi
-
[133]
Rose Wang and Dorottya Demszky. 2023. https://doi.org/10.18653/v1/2023.bea-1.53 Is ChatGPT a Good Teacher Coach ? Measuring Zero - Shot Performance For Scoring and Providing Actionable Insights on Classroom Instruction . In Proceedings of the 18th Workshop on Innovative Use of...
2023 doi
-
[134]
Melissa Warr, Nicole Jakubczyk Oster, and Roger Isaac. 2024. https://doi.org/10.1080/15391523.2024.2395295 Implicit bias in large language models: Experimental proof and implications for education . Journal of Research on Technology in Education, 0(0):1--24. Publisher: Routled...
2024
-
[135]
Zeerak Waseem. 2016. https://doi.org/10.18653/v1/W16-5618 Are You a Racist or Am I Seeing Things ? Annotator Influence on Hate Speech Detection on Twitter . In Proceedings of the First Workshop on NLP and Computational Social Science , pages 138--142, Austin, Texas. Associatio...
2016 doi
-
[136]
Albert Webson, Alyssa Loo, Qinan Yu, and Ellie Pavlick. 2023. https://doi.org/10.18653/v1/2023.findings-emnlp.514 Are Language Models Worse than Humans at Following Prompts ? It 's Complicated . In Findings of the Association for Computational Linguistics : EMNLP 2023 , pages ...
2023 doi
-
[137]
Albert Webson and Ellie Pavlick. 2022. https://doi.org/10.18653/v1/2022.naacl-main.167 Do Prompt - Based Models Really Understand the Meaning of Their Prompts ? In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics...
2022 doi
- [138]
-
[139]
Whitehurst, Matthew M
Grover J. Whitehurst, Matthew M. Chingos, and Katharine M. Lindquist. 2014. Evaluating Teachers with Classroom Observations : Lessons Learned in Four Districts . Technical report, Brookings Institution. Publication Title: Brookings Institution ERIC Number: ED553815
2014
-
[140]
Stefanie A. Wind. 2019. https://doi.org/10.1111/jedm.12222 Nonparametric Evidence of Validity , Reliability , and Fairness for Rater - Mediated Assessments : An Illustration Using Mokken Scale Analysis . Journal of Educational Measurement, 56(3):478--504. \_eprint: https://onl...
2019 doi
-
[141]
Wind and Wenjing Guo
Stefanie A. Wind and Wenjing Guo. 2019. https://doi.org/10.1177/0013164419834613 Exploring the Combined Effects of Rater Misfit and Differential Rater Functioning in Performance Assessments . Educational and Psychological Measurement, 79(5):962--987. Publisher: SAGE Publications Inc
2019 doi
- [142]
- [143]
-
[144]
Rich Zemel, Yu Wu, Kevin Swersky, Toni Pitassi, and Cynthia Dwork. 2013. https://proceedings.mlr.press/v28/zemel13.html Learning Fair Representations . In Proceedings of the 30th International Conference on Machine Learning , pages 325--333. PMLR. ISSN: 1938-7228
2013
- [145]
- [146]
-
[147]
Quintana, Anita Delahay, and Xu Wang
Xiaofei Zhou, Christopher Kok, Rebecca M. Quintana, Anita Delahay, and Xu Wang. 2023. https://doi.org/10.1145/3573051.3593388 How Learning Experience Designers Make Design Decisions : The Role of Data , the Reliance on Subject Matter Expertise , and the Opportunities for Data ...
2023
-
[148]
Eva A. O. Zijlmans, Jesper Tijmstra, L. Andries van der Ark, and Klaas Sijtsma. 2018 a . https://doi.org/10.1177/0013164417728358 Item- Score Reliability in Empirical - Data Sets and Its Relationship With Other Item Indices . Educational and Psychological Measurement, 78(6):99...
2018 doi
-
[149]
Eva A. O. Zijlmans, L. Andries van der Ark, Jesper Tijmstra, and Klaas Sijtsma. 2018 b . https://doi.org/10.1177/0146621618758290 Methods for Estimating Item - Score Reliability . Applied Psychological Measurement, 42(7):553--570. Publisher: SAGE Publications Inc
2018 doi
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.