Pith. sign in

REVIEW 4 major objections 6 minor 147 references

"All that Glitters": Approaches to Evaluations with Unreliable Model and Human Annotations

T0 review · 4 major / 6 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read Standard agreement metrics, applied to unreliable human labels, can make encoder models look super-human at rating classroom teaching; stricter psychometric measures reveal spurious correlations and nonrandom bias in model and human raters.

desk verdict Brings psychometric tools NLP should know, but its central validity measure likely overcorrects on its own assumption; deserves peer review with major revisions. read the letter →

arxiv 2411.15634 v1 pith:YAJC7B5R submitted 2024-11-23 cs.CL cs.AIstat.AP

classification cs.CLcs.AIstat.AP
keywords LLMevaluationlabelreliabilitygeneralizabilitytheoryhierarchicalratermodelsdisattenuatedcorrelationclassroomobservationracialbiashuman-in-the-loop
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper asks whether automated systems can take over a task that is currently done only by expert humans: rating the quality of classroom teaching from transcripts. Because those expert ratings are themselves highly unreliable, the paper argues that the usual concordance metrics (agreement rates, kappa, correlations) cannot answer that question, and shows that on those metrics its encoder models outscore the best human raters—an apparent "super-human" result. The paper's central claim is that switching to psychometric measures that decompose label variance overturns part of that verdict: several encoder wins dissolve as spurious correlation, human-model agreement on some items is not evidence that the two track the same construct, and both GPT models and individual human raters show measurable nonrandom racial bias. On the useful side, decision-study estimates predict that encoder ratings could raise the reliability of observations of rare teaching behaviors at a fraction of current cost, while GPT ratings would drag human reliability down.

What carries the argument

The load-bearing device is Eq. (4), the disattenuated convergent correlation: the observed correlation between human and model ratings of the same teacher on different lessons, divided by the square root of the product of the two rater families' generalizability coefficients $E\rho^2$. Because low reliability mechanically shrinks observed correlations, this correction separates "the two raters track the same underlying construct" from "the correlation is an artifact of measurement error." The generalizability coefficient $E\rho^2$ itself—the share of rating variance attributable to the teacher rather than to lesson, segment, rater, and item facets—is estimated from a nested random-effects model $R \times (S:O:I)$ and carries the confidence and decision-study analyses. Rater-level behavior is quantified by a hierarchical rater model whose item-response-theory stage estimates ideal scores and whose signal-detection stage assigns each rater a bias parameter $\phi$ and a variability parameter $\psi^2$, extended with race covariates to test fairness. These pieces recombine in the decision study of Eq. (9), which reweights variance components to predict reliability under proposed human-in-the-loop designs.

What would settle it

Apply the same six-metric protocol to a held-out second cohort of classroom transcripts rated by multiple experts, or re-run Eq. (4) using only same-week lessons of each teacher. If the near-1.0 disattenuated correlations shrink or exceed 1.0 when lessons are close in time, the stability assumption fails and the corrected correlations are artifacts of the correction; if encoder superiority on $E\rho^2$, the spurious-item pattern, and the GPT negative-bias trend fail to reproduce on new transcripts, the discovery is an artifact of this dataset rather than a property of the methods.

Watch

Extended reading notes

Core claim

The central claim is that when expert human labels are unreliable, standard inter-rater statistics—percent agreement, Cohen's $\kappa$, quadratic weighted kappa, Pearson, Spearman, and Kendall correlations, ICCs—can certify a model as "super-human" while masking what the model actually learned. On those measures, the paper's five encoder models outperform the best human raters on nearly every item of the MQI and CLASS instruments. Under Generalizability Theory (the $E\rho^2$ generalizability coefficient and the $\Phi$ dependability coefficient), which attribute rating variance to teacher, lesson, segment, and rater facets, the encoder advantage shrinks and reverses on items like teacher explanations (EXPL) and student explanations (STEXPL). Disattenuating human-model correlations by the reliabilities of each rater family shows that some apparent model-human agreement is spurious, while items with near-1.0 corrected correlations indicate genuine shared signal only if a teacher's latent ability is stable across lessons. A hierarchical rater model—an item-response-theory true-score stage feeding a signal-detection rater stage—estimates each rater's leniency/severity bias $\phi$ and consistency $\psi^2$, and conditioning those on teacher race shows a negative bias trend against Black teachers in the GPT family, smaller but nonrandom biases in encoders, and detectable racial biases in some human raters on negatively worded items. Closing decision studies estimate that a three-encoder ensemble observing whole class periods can double the reliability achievable for the rare-construct item LANGIMP relative to ten fifteen-minute human visits, at a large time saving, while GPT ensembles would reduce human rating reliability in human-in-the-loop use.

Load-bearing premise

The near-perfect disattenuated correlations assume a teacher's true instructional ability stays about the same from lesson to lesson, so that a human rating one lesson and a model rating another are measuring the same thing; if ability shifts between lessons, the corrected correlation cannot distinguish shared construct from lesson-to-lesson variability.

Editorial extensions

If this is right

  • Reporting only concordance metrics can certify a model as super-human against an unreliable human baseline, so high-stakes annotation evaluations should also report generalizability and dependability.
  • Encoder models trained on transcripts can, for specific MQI items such as LANGIMP, deliver rating reliability roughly double what ten short human visits achieve, at a large time saving.
  • GPT-style prompt-engineered ratings of classroom instruction currently underperform expert humans on nearly every metric, and decision-study estimates predict their variance would lower human rating reliability in human-in-the-loop use.
  • Nonrandom racial bias is detectable at the individual-rater level even at very low label reliabilities: GPT models show a negative bias trend against Black teachers, encoders show smaller but real biases on sparse items, and some human raters show racial differences on negatively worded items.
  • The evaluation protocol transfers to other NLP tasks with unreliable expert annotations: g-studies identify label weaknesses before training, disattenuation tests whether model-human correlations reflect shared constructs, and d-studies estimate the value of model assistance before costly trials are run.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The same six-part protocol would transfer to other benchmarks whose "gold" labels come from a small pool of expensive experts, such as essay scoring, clinical chart review, or content moderation; reliability-corrected model rankings on such benchmarks would likely differ from the raw-correlation leaderboards currently reported.
  • A direct test of the spuriousness story would retrain the encoder family with speaker-identity markers added to the transcripts: the paper's design deliberately withheld speaker information, and its claim that the EXPL and STEXPL failures stem from that omission predicts that those items' generalizability would recover once speaker roles are visible.
  • The human-in-the-loop decision studies offer a small, cheap falsifiable pilot: run a handful of classrooms where principals rate under their usual 15-minute visits while an encoder scores full periods in the background, then compare achieved reliability and time spent against the predicted curves; the claimed savings of roughly two to ten hours per teacher are currently extrapolations, not measure
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper argues that common concordance metrics (correlation, percent agreement, kappa) can be misleading when human annotations are unreliable, and it demonstrates a set of psychometric tools—generalizability theory, disattenuated correlations, hierarchical rater models, fairness analyses, and decision studies—for evaluating model and human annotations of classroom teaching quality. Using the NCTE dataset, the authors compare human expert ratings, GPT-family ratings from a prior study, and five newly trained transformer encoder models. They report that encoder models appear to achieve state-of-the-art or 'super-human' concordance under standard metrics, but that more rigorous generalizability and validity analyses reveal spurious correlations and racial biases, while GPT models perform poorly and would worsen human-in-the-loop reliability.

Significance. If the qualitative conclusions hold, the paper makes a valuable methodological contribution to NLP evaluation: it demonstrates how to quantify label quality, bias, and validity when human ratings are noisy, and it provides an applied case study with public code and data links. The replication of the NCTE g-study, the use of hierarchical rater models to disentangle rater bias from construct signal, and the explicit treatment of limitations (including the acknowledgment that the encoder models are trained on the same noisy human labels they are later evaluated against) are genuine strengths. The central qualitative claim—that standard metrics can mask model and label quality—is supportable, but several load-bearing numerical claims are fragile as currently presented, particularly the disattenuated-correlation evidence for spuriousness and the magnitude of the 'super-human' performance.

major comments (4)
  1. [§5.3 and Appendix G (Eq. 13)] The disattenuated correlation in Eq. (4) rests on the assumption stated in §5.3.1: 'If an individual teacher's latent instructional ability theta_i is about the same from lesson to lesson with the same students.' The paper's own hierarchical rater model in Eq. (13) of Appendix G explicitly models lesson-level latent abilities theta_oi varying around teacher-level Theta_i, and the g-study of Eq. (1)–(2) treats lesson-within-teacher variance (nu_o:i) as error. If theta actually varies across lessons, the numerator of Eq. (4) mixes construct overlap with lesson-to-lesson instability, while the denominator divides by an E_rho2 estimate that includes that instability as error, so the correction overcorrects. The near-1.0 disattenuated correlations in Figure 3 and Figure 4(c) (e.g., EXPL at 1.0†) may then be artifacts of the correction rather than evidence that humans and models track the same stable construct. This threatens the paper's second stated contribution (Section 1, 'methods for detection of spurious correlations via disattenuating low human-model correlations'). A concrete test would be to split items or teachers by the magnitude of lesson-within-teacher variance and show that the disattenuated correlations are not driven by high-variance items, or to estimate the model of Eq. (13) and use the teacher-level variance component in the denominator.
  2. [§5.1, Tables 1 and 9] The headline claim that encoder models achieve 'super-human' results across all classroom annotation tasks (abstract, §5.1.3) is not adequately supported because no majority-class baseline is reported. On highly imbalanced items such as MGEN, MAJERR, and LANGIMP, where the dominant score category accounts for a large fraction of labels (see Figure 6 and Figure 8), a trivial classifier that always predicts the modal category would achieve high percent agreement and respectable kappa values. For example, Table 9 shows encoder percent agreement of 0.95–0.96 on MGEN; without a modal-class baseline, this value is uninterpretable. The generalizability and HRM analyses partially mitigate this concern, but the quantitative 'super-human' framing is load-bearing for the paper's message. Please report majority-class baselines and chance-corrected agreement measures (e.g., prevalence-adjusted kappa) for all items.
  3. [Table 2 and §5.2.2] The generalizability and dependability estimates E_R^2 and Phi in Table 2 are reported as point estimates without confidence intervals or other uncertainty quantification, despite the low reliability levels (many values between 0.00 and 0.20) and small item counts. For instance, the human-vs-encoder differences on EXPL (0.15 vs 0.00) and SMQR (0.14 vs 0.09) in Table 2 may be within sampling error. Since Section 5.3.2 uses these exact values to declare correlations 'spurious' (e.g., EXPL and STEXPL), the absence of uncertainty intervals makes the numeric conclusions fragile. Please provide bootstrap or Bayesian intervals for the variance components and derived coefficients.
  4. [Abstract and §7 (Limitations)] The encoder models were trained on the same human ratings that are used later in the evaluation, so the 'super-human' concordance result in the abstract partly reflects the models fitting the evaluation target rather than an independent assessment of quality. The paper acknowledges this in Section 7 ('the signal is still trained on noisy human ratings'), and the psychometric analyses are motivated by this dependency, but the abstract's unqualified statement 'the encoder family of models achieve state-of-the-art, even "super-human", results across all classroom annotation tasks' overstates the finding. Recommend rewording the abstract to indicate that the models achieve state-of-the-art concordance with the training labels, with the caveat that this concordance is not evidence of construct validity.
minor comments (6)
  1. [§4] The text reads 'Encoder models were trained a single GPU in Google Colab'; 'a' should be 'on a single GPU'.
  2. [Table 1 caption] The caption describes the last metric as 'Kendall's concordance correlation'; this should be 'Kendall's tau rank correlation coefficient'.
  3. [§5.3.1 and footnote 12] The sentence 'Disattenuated correlations of 1.0 do not mean perfect correlation: it generally means that measurement error is not randomly distributed' appears twice (in the main text and in footnote 12) and is confusing; consider rephrasing or deleting one occurrence.
  4. [Appendix D.1.1] The citation 'Gao, 2022' for SimCSE refers to a reference in the bibliography that is not the SimCSE paper (the listed Shuai Gao entry is about a French-Mongolian MT system). The correct reference is the SimCSE paper by Gao et al. (2021). Similarly, in Table 7 the E5 embedding model is cited as 'Wang et al. (2022)', but the bibliography entry points to a Jiarui Wang et al. paper on text style transfer, not the E5 embeddings paper. These citation errors should be corrected.
  5. [Figure 4 caption] Panel (b) is labeled 'Reliabilities' but includes metrics such as percent agreement and percent agreement ±1, which are agreement indices rather than reliability coefficients; consider using a more precise label such as 'Agreement metrics'.
  6. [§5.1.3] The sentence 'Using nearly any standardized combination of metrics across all items from Section 5.1, Encoder models perform better than the single highest performing expert human rater' is broader than what Table 9 shows; for items like MMETH and STEXPL, human correlations are higher than the encoder family values. Suggest qualifying this as 'on average across items' or specifying the exception pattern.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: encoder evaluation uses a disjoint held-out test set; psychometric corrections are separate estimates; remaining concerns are validity assumptions, not circular reductions.

full rationale

After walking the derivation chain, I find no circular step that reduces a claimed result to its own inputs. The encoder 'super-human' concordance result is a supervised-learning evaluation against a lesson-stratified held-out test set: Section 4 states that 'all model outputs in this study were conducted with a lesson-level-stratified held-out test set (see Figure 8) that was not used during model development,' and the Limitations passage that 'the signal is still trained on noisy human ratings' is a data-quality caveat, not an identity between the training target and the evaluation target. The g-theory generalizability estimates (Eqs. 1-3) and the disattenuated correlations (Eq. 4) are separate quantities: the numerator is a cross-family, cross-lesson correlation, while the denominator uses variance components estimated from each family's own ratings; neither is defined in terms of the other, so no self-definitional reduction is present. The HRM bias and fairness results are estimated jointly from the rating data by MCMC with stated priors, and the GPT ratings are imported from an external study (Wang and Demszky, 2023), not from a self-citation chain. The only self-citations (Hardy 2021 as an example of QWK use, and a note about a forthcoming paper) are illustrative and not load-bearing. The Eq. 4 lesson-stability assumption is a substantive identification assumption whose violation would threaten validity, but that is a correctness or robustness concern, not circularity: the formula does not reduce to its inputs by construction.

Assumptions & free parameters 4 free parameters · 5 assumptions · 0 invented entities

The central estimates (Eρ², disattenuated correlations, rater bias, d-study reliability) are all produced by statistical models fitted to the same noisy NCTE labels. They are not derived from first principles or external benchmarks. The free parameters listed are the ones the central claims quantitatively depend on.

free parameters (4)
  • encoder hyperparameters (dropout, attention heads, epochs, learning rate, weight decay) = dropout=0.75, heads=32, epochs=15-75, lr=2.5e-5, weight decay=0.0003
    Chosen by hand to make training work on noisy labels; listed in Table 7 and Appendix D.3. The central encoder performance claims depend on these choices.
  • G-theory variance components for Eρ² and Φ per item per rater family = posterior variance estimates from lme4 (not tabulated as single numbers)
    Estimated from the data in Eq. 1 and used directly in Equations 2-4 and 9-10. The disattenuated correlations and d-study outputs are functions of these fitted values.
  • HRM rater bias φ and variability ψ per rater per item per race = posterior means from MCMC (not listed in main text)
    Estimated from the data with weakly informative priors in Sections 5.4 and 5.5. They are the basis for the bias and fairness claims.
  • Decision study design counts (segments, observations, raters) = n_s=6, n_o=1, n_r=1 for HIL; human 15 minutes, model 45 minutes
    Chosen scenario parameters in Section 5.6.1. They determine the reliability estimates in Figure 4(f).
assumptions (5)
  • domain assumption Latent teaching ability theta is multivariate normal with zero mean and identity covariance (Eq. 5), and teacher-year latent abilities are normal (Eq. 13).
    Needed for the HRM and IRT estimation of true scores and rater bias. The paper notes in Limitations that this may not hold for models.
  • domain assumption A teacher's latent instructional ability is approximately stable across lessons with the same students (Section 5.3.1).
    Load-bearing for disattenuated correlations: without it, the numerator of Eq. 4 is contaminated by lesson-to-lesson variation.
  • domain assumption Local independence: ratings are conditionally independent given the true score, rater, and item (Eqs. 5-8).
    Standard IRT and SDT assumption. Violations, such as within-lesson autocorrelation, are acknowledged in Limitations.
  • domain assumption G-theory random effects: raters, observations, and segments are sampled from a defined universe, and variance components generalize to that universe (Section 5.2.1).
    Required for all Eρ² and d-study estimates. The paper follows the original study's design.
  • domain assumption For the fairness analysis, teacher race is the only covariate influencing rater bias beyond item and rater, with no unmeasured confounders (Section 5.5.1).
    The model estimates φ separately for Black and White teachers and interprets differences as racial bias. Confounders correlated with race would break this.

how reviews work

0 comments
Cite this review

Pith. "Pith review of "All that Glitters": Approaches to Evaluations with Unreliable Model and Human Annotations." pith.science (2026). https://pith.science/paper/YAJC7B5R

@misc{pith2026241115634,
  author       = {Pith},
  title        = {Pith review of: "All that Glitters": Approaches to Evaluations with Unreliable Model and Human Annotations},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/YAJC7B5R}},
  note         = {Machine review of arXiv:2411.15634}
}
read the original abstract

"Gold" and "ground truth" human-mediated labels have error. The effects of this error can escape commonly reported metrics of label quality or obscure questions of accuracy, bias, fairness, and usefulness during model evaluation. This study demonstrates methods for answering such questions even in the context of very low reliabilities from expert humans. We analyze human labels, GPT model ratings, and transformer encoder model annotations describing the quality of classroom teaching, an important, expensive, and currently only human task. We answer the question of whether such a task can be automated using two Large Language Model (LLM) architecture families--encoders and GPT decoders, using novel approaches to evaluating label quality across six dimensions: Concordance, Confidence, Validity, Bias, Fairness, and Helpfulness. First, we demonstrate that using standard metrics in the presence of poor labels can mask both label and model quality: the encoder family of models achieve state-of-the-art, even "super-human", results across all classroom annotation tasks. But not all these positive results remain after using more rigorous evaluation measures which reveal spurious correlations and nonrandom racial biases across models and humans. This study then expands these methods to estimate how model use would change to human label quality if models were used in a human-in-the-loop context, finding that the variance captured in GPT model labels would worsen reliabilities for humans influenced by these models. We identify areas where some LLMs, within the generalizability of the current data, could improve the quality of expensive human ratings of classroom instruction.

Figures

Figures reproduced from arXiv: 2411.15634 by the authors.

Figure 1
Figure 1. Data Processes and Sources for Studying Teaching and Annotation Quality [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 6
Figure 6. Each instrument item is intended to measure a different aspect of teaching quality. [PITH_FULL_IMAGE:figures/full_fig_p004_6.png] view at source ↗
Figure 2
Figure 2. Spearman correlation coefficients and confidence intervals by MQI Item for all rater families and studies. [PITH_FULL_IMAGE:figures/full_fig_p006_2.png] view at source ↗
Figures from the paper (9 more)
Figure 3
Figure 3. Figure 3: Correlations (fainter color hues, numerator of Eq. 4), disattenuated correlations (darker color hues, Eq. 4), and their respective 95% confidence intervals between human raters and model raters by MQI item. Item-level rater-label generalizability for both human and mod…
Figure 15
Figure 15. Figure 15: Investigating the validity of a construct would require more robust qualitative review of the content. [PITH_FULL_IMAGE:figures/full_fig_p010_15.png]
Figure 4
Figure 4. Figure 4: Section 5 Study Method Results for four focus MQI Items across Human (Kane et al., 2015), Encoder (this study), and GPT (Wang and Demszky, 2023) rater families. (a) Distributions. Score distributions by rater type. (b) Reliabilities. Inter-rater reliability metrics int…
Figure 5
Figure 5. Figure 5: Overview of technical details the two instructional frameworks used for evaluating instruction. [PITH_FULL_IMAGE:figures/full_fig_p033_5.png]
Figure 7
Figure 7. Figure 7: Model Pipeline: General sentence-encoder model architecture. [PITH_FULL_IMAGE:figures/full_fig_p037_7.png]
Figure 10
Figure 10. Figure 10: Rater biases, [PITH_FULL_IMAGE:figures/full_fig_p053_10.png]
Figure 11
Figure 11. Figure 11: Fairness across Racial Lines. Section 5.5: Standardized difference in rater bias 𝜙𝑟 (x axis) and rater combined variability/consistency, 𝜓𝑟, (y axis) across Black teachers and White teachers. Leftward values are more severe towards Black teachers, rightward are more l…
Figure 12
Figure 12. Figure 12: Variance components for Generalizability Calculations [PITH_FULL_IMAGE:figures/full_fig_p055_12.png]
Figure 13
Figure 13. Figure 13: Estimates for Family-wise Item-level Generalizability, [PITH_FULL_IMAGE:figures/full_fig_p056_13.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

147 extracted references · 36 canonical work pages

  1. [1]

    online" 'onlinestring :=

    ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRING...

  2. [2]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...

  3. [3]

    Gavin Abercrombie, Verena Rieser, and Dirk Hovy. 2023. https://doi.org/10.48550/arXiv.2301.10684 Consistency is Key : Disentangling Label Variation in Natural Language Processing with Intra - Annotator Agreement . arXiv preprint. ArXiv:2301.10684 [cs]

  4. [4]

    Adams, Mark Wilson, and Wen-chung Wang

    Raymond J. Adams, Mark Wilson, and Wen-chung Wang. 1997. https://doi.org/10.1177/0146621697211001 The multidimensional random coefficients multinomial logit model . Applied Psychological Measurement, 21(1):1--23. Place: US Publisher: Sage Publications

  5. [5]

    Julius Adebayo, Justin Gilmer, Michael Muelly, Ian Goodfellow, Moritz Hardt, and Been Kim. 2020. https://doi.org/10.48550/arXiv.1810.03292 Sanity Checks for Saliency Maps . arXiv preprint. ArXiv:1810.03292 [cs, stat]

  6. [6]

    Nikhil Agarwal, Alex Moehring, Pranav Rajpurkar, and Tobias Salz. 2023. https://doi.org/10.3386/w31422 Combining Human Expertise with Artificial Intelligence : Experimental Evidence from Radiology

  7. [7]

    Elena Aguilar. 2013. Developing a Work Plan : How Do I Determine What to Do ? In The art of coaching: effective strategies for school transformation, pages 119--144. Jossey-Bass, A Wiley Brand, San Francisco

  8. [8]

    Sterling Alic, Dorottya Demszky, Zid Mancenido, Jing Liu, Heather Hill, and Dan Jurafsky. 2022. https://doi.org/10.18653/v1/2022.bea-1.27 Computationally identifying funneling and focusing questions in classroom discourse . In Proceedings of the 17th Workshop on Innovative Use of NLP for Building Educational Applications (BEA 2022), pages 224--233, Seattl...

Show all 147 references
  1. [9]

    Amos Azaria, Rina Azoulay, and Shulamit Reches. 2024. https://doi.org/10.1162/dint_a_00235 ChatGPT is a Remarkable Tool — For Experts . Data Intelligence, 6(1):240--296

  2. [10]

    Joris Baan, Wilker Aziz, Barbara Plank, and Raquel Fernández. 2022. https://doi.org/10.48550/arXiv.2210.16133 Stop Measuring Calibration When Humans Disagree . arXiv preprint. ArXiv:2210.16133 [cs]

  3. [11]

    Joris Baan, Raquel Fernández, Barbara Plank, and Wilker Aziz. 2024. https://doi.org/10.48550/arXiv.2402.16102 Interpreting Predictive Probabilities : Model Confidence or Human Label Variation ? arXiv preprint. ArXiv:2402.16102 [cs] version: 1

  4. [12]

    Chin, Thomas J

    Andrew Bacher-Hicks, Mark J. Chin, Thomas J. Kane, and Douglas O. Staiger. 2017. https://doi.org/10.3386/w23478 An Evaluation of Bias in Three Measures of Teacher Quality : Value - Added , Classroom Observations , and Student Surveys

  5. [13]

    Chin, Thomas J

    Andrew Bacher-Hicks, Mark J. Chin, Thomas J. Kane, and Douglas O. Staiger. 2019. https://doi.org/10.1016/j.econedurev.2019.101919 An experimental evaluation of three teacher quality measures: Value -added, classroom observations, and student surveys . Economics of Education Re...

  6. [14]

    Paul Bambrick-Santoyo. 2016. Get better faster: a 90-day plan for coaching new teachers. Jossey-Bass, A Wiley Brand, San Francisco, CA

  7. [15]

    Paul Bambrick-Santoyo. 2018. Leverage leadership 2.0: a practical guide to building exceptional schools. Jossey-Bass, San Francisco, CA

  8. [16]

    Douglas Bates, Martin Mächler, Ben Bolker, and Steve Walker. 2015. https://doi.org/10.18637/jss.v067.i01 Fitting Linear Mixed - Effects Models Using lme4 . Journal of Statistical Software, 67:1--48

  9. [17]

    Williamson, and and Robert J

    Isaac 1 Bejar, David M. Williamson, and and Robert J. Mislevy. 2006. Human Scoring . In Automated Scoring of Complex Tasks in Computer - Based Testing . Routledge. Num Pages: 34

  10. [18]

    Howcroft

    Anya Belz, Simon Mille, and David M. Howcroft. 2020. https://doi.org/10.18653/v1/2020.inlg-1.24 Disentangling the Properties of Human Evaluation Methods : A Classification System to Support Comparability , Meta - Evaluation and Reproducibility Testing . In Proceedings of the 1...

  11. [19]

    Anya Belz, Craig Thomson, Ehud Reiter, and Simon Mille. 2023. https://doi.org/10.18653/v1/2023.findings-acl.226 Non- Repeatable Experiments and Non - Reproducible Results : The Reproducibility Crisis in Human Evaluation in NLP . In Findings of the Association for Computational...

  12. [21]

    Eckstein, Noémi Éltető, Thomas L

    Marcel Binz, Elif Akata, Matthias Bethge, Franziska Brändle, Fred Callaway, Julian Coda-Forno, Peter Dayan, Can Demircan, Maria K. Eckstein, Noémi Éltető, Thomas L. Griffiths, Susanne Haridi, Akshay K. Jagadish, Li Ji-An, Alexander Kipnis, Sreejan Kumar, Tobias Ludwig, Marvin ...

  13. [22]

    Abeba Birhane, Pratyusha Kalluri, Dallas Card, William Agnew, Ravit Dotan, and Michelle Bao. 2022. https://doi.org/10.48550/arXiv.2106.15590 The Values Encoded in Machine Learning Research . arXiv preprint. ArXiv:2106.15590

  14. [23]

    David Blazar. 2018. https://doi.org/10.1162/edfp_a_00251 Validating Teacher Effects on Students ’ Attitudes and Behaviors : Evidence from Random Assignment of Teachers to Students . Education Finance and Policy, 13(3):281--309

  15. [24]

    Charalambous, and Heather C

    David Blazar, David Braslow, Charalambos Y. Charalambous, and Heather C. Hill. 2017. https://doi.org/10.1080/10627197.2017.1309274 Attending to General and Mathematics - Specific Dimensions of Teaching : Exploring Factors Across Two Observation Instruments . Educational Assess...

  16. [25]

    David Blazar and Cynthia Pollard. 2022. https://www.edworkingpapers.com/ai22-591 Challenges and Tradeoffs of “ Good ” Teaching : The Pursuit of Multiple Educational Outcomes . Technical report, Annenberg Institute at Brown University. Publication Title: EdWorkingPapers.com

  17. [26]

    Robert L. Brennan. 2001 a . https://doi.org/10.1007/978-1-4757-3456-0 Generalizability Theory . Springer, New York, NY

  18. [27]

    Robert L. Brennan. 2001 b . https://doi.org/10.1007/978-1-4757-3456-0_6 Variability of Statistics in Generalizability Theory . In Robert L. Brennan, editor, Generalizability Theory , Statistics for Social Sciences and Public Policy , pages 179--213. Springer, New York, NY

  19. [28]

    Robert L. Brennan. 2013. Generalizability Theory . Springer Science & Business Media. Google-Books-ID: nbHbBwAAQBAJ

  20. [29]

    Briggs and Mark Wilson

    Derek C. Briggs and Mark Wilson. 2007. https://doi.org/10.1111/j.1745-3984.2007.00031.x Generalizability in item response modeling . Journal of Educational Measurement, 44(2):131--155. Place: United Kingdom Publisher: Blackwell Publishing

  21. [30]

    Casabianca

    Jodi M. Casabianca. 2021. https://doi.org/10.1111/emip.12478 Digital Module 27: Hierarchical Rater Models . Educational Measurement: Issues and Practice, 40(4):103--104. \_eprint: https://onlinelibrary.wiley.com/doi/pdf/10.1111/emip.12478

  22. [31]

    Casabianca, Daniel F

    Jodi M. Casabianca, Daniel F. McCaffrey, Drew H. Gitomer, Courtney A. Bell, Bridget K. Hamre, and Robert C. Pianta. 2013. https://doi.org/10.1177/0013164413486987 Effect of Observation Mode on Measures of Secondary Mathematics Teaching . Educational and Psychological Measureme...

  23. [32]

    Charalambous and Seán Delaney

    Charalambos Y. Charalambous and Seán Delaney. 2019. https://doi.org/10.1163/9789004418875_014 13 Mathematics Teaching Practices and Practice - Based Pedagogies . Brill. Section: International Handbook of Mathematics Teacher Education: Volume 1

  24. [33]

    Eric P. Charles. 2005. https://doi.org/10.1037/1082-989X.10.2.206 The Correction for Attenuation Due to Measurement Error : Clarifying Concepts and Creating Confidence Sets . Psychological Methods, 10(2):206--226. Place: US Publisher: American Psychological Association

  25. [34]

    Gaebler, Hamed Nilforoshan, Ravi Shroff, and Sharad Goel

    Sam Corbett-Davies, Johann D. Gaebler, Hamed Nilforoshan, Ravi Shroff, and Sharad Goel. 2023. https://doi.org/10.48550/arXiv.1808.00023 The Measure and Mismeasure of Fairness . arXiv preprint. ArXiv:1808.00023 [cs]

  26. [35]

    Chengyu Cui, Chun Wang, and Gongjun Xu. 2024. https://doi.org/10.1007/s11336-024-09955-8 Variational Estimation for Multidimensional Generalized Partial Credit Model . Psychometrika

  27. [36]

    Alexander D'Amour, Katherine Heller, Dan Moldovan, Ben Adlam, Babak Alipanahi, Alex Beutel, Christina Chen, Jonathan Deaton, Jacob Eisenstein, Matthew D. Hoffman, Farhad Hormozdiari, Neil Houlsby, Shaobo Hou, Ghassen Jerfel, Alan Karthikesalingam, Mario Lucic, Yian Ma, Cory Mc...

  28. [37]

    Linda Darling-Hammond. 2014. https://scholarworks.umb.edu/nejpp/vol26/iss1/4 What Can PISA Tell Us about U . S . Education Policy ? New England Journal of Public Policy, 26(1)

  29. [38]

    Linda Darling-Hammond, Lisa Flook, Channa Cook-Harvey, Brigid Barron, and David Osher. 2020. https://doi.org/10.1080/10888691.2018.1537791 Implications for educational practice of the science of learning and development . Applied Developmental Science, 24(2):97--140. Publisher...

  30. [39]

    Lawrence T. Decarlo. 2003. https://doi.org/10.3758/BF03195496 Using the PLUM procedure of SPSS to fit unequal variance and generalized signal detection models . Behavior Research Methods, Instruments, & Computers, 35(1):49--56

  31. [40]

    Lawrence T. DeCarlo. 2008. https://doi.org/10.1002/j.2333-8504.2008.tb02149.x Studies of a Latent - Class Signal - Detection Model for Constructed - Response Scoring . ETS Research Report Series, 2008(2):i--55. \_eprint: https://onlinelibrary.wiley.com/doi/pdf/10.1002/j.2333-8...

  32. [41]

    Lawrence T. DeCarlo. 2023. https://doi.org/10.1111/jedm.12358 Classical Item Analysis from a Signal Detection Perspective . Journal of Educational Measurement, 60(3):520--547. \_eprint: https://onlinelibrary.wiley.com/doi/pdf/10.1111/jedm.12358

  33. [42]

    DeCarlo, YoungKoung Kim, and Matthew S

    Lawrence T. DeCarlo, YoungKoung Kim, and Matthew S. Johnson. 2011. https://www.jstor.org/stable/23018150 A Hierarchical Rater Model for Constructed Responses , with a Signal Detection Rater Model . Journal of Educational Measurement, 48(3):333--356. Publisher: National Council...

  34. [43]

    Dorottya Demszky and Heather Hill. 2022. https://doi.org/10.48550/ARXIV.2211.11772 The NCTE Transcripts : A Dataset of Elementary Math Classroom Transcripts . Publisher: arXiv Version Number: 1

  35. [44]

    Dorottya Demszky and Heather Hill. 2023. https://doi.org/10.18653/v1/2023.bea-1.44 The NCTE Transcripts : A Dataset of Elementary Math Classroom Transcripts . In Proceedings of the 18th Workshop on Innovative Use of NLP for Building Educational Applications ( BEA 2023) , pages...

  36. [46]

    Hill, Shyamoli Sanghi, and Ariel Chung

    Dorottya Demszky, Jing Liu, Heather C. Hill, Shyamoli Sanghi, and Ariel Chung. 2023. https://edworkingpapers.com/ai23-875 Improving Teachers ’ Questioning Quality through Automated Feedback : A Mixed - Methods Randomized Controlled Trial in Brick -and- Mortar Classrooms . Tech...

  37. [47]

    Dorottya Demszky, Jing Liu, Zid Mancenido, Julie Cohen, Heather Hill, Dan Jurafsky, and Tatsunori Hashimoto. 2021. https://doi.org/10.48550/ARXIV.2106.03873 Measuring Conversational Uptake : A Case Study on Student - Teacher Interactions . Publisher: arXiv Version Number: 1

  38. [48]

    Dorottya Demszky, Rose Wang, Sean Geraghty, and Carol Yu. 2024. https://doi.org/10.1145/3636555.3636924 Does Feedback on Talk Time Increase Student Engagement ? Evidence from a Randomized Controlled Trial on a Math Tutoring Platform . In Proceedings of the 14th Learning Analyt...

  39. [49]

    Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. https://doi.org/10.18653/v1/N19-1423 BERT : Pre -training of Deep Bidirectional Transformers for Language Understanding . In Proceedings of the 2019 Conference of the North American Chapter of the Associat...

  40. [50]

    Frances Ding, Moritz Hardt, John Miller, and Ludwig Schmidt. 2022. https://doi.org/10.48550/arXiv.2108.04884 Retiring Adult : New Datasets for Fair Machine Learning . arXiv preprint. ArXiv:2108.04884 [cs, stat]

  41. [51]

    Donnelly, Nathaniel Blanchard, Andrew M

    Patrick J. Donnelly, Nathaniel Blanchard, Andrew M. Olney, Sean Kelly, Martin Nystrand, and Sidney K. D'Mello. 2017. https://doi.org/10.1145/3027385.3027417 Words matter: automatic detection of teacher questions in live classroom discourse using linguistics, acoustics, and con...

  42. [52]

    Cynthia Dwork, Moritz Hardt, Toniann Pitassi, Omer Reingold, and Richard Zemel. 2012. https://doi.org/10.1145/2090236.2090255 Fairness through awareness . In Proceedings of the 3rd Innovations in Theoretical Computer Science Conference , ITCS '12, pages 214--226, New York, NY,...

  43. [53]

    Detecting Illusory Halo Effects in Rater - Mediated Assessment : A Mixture Rasch Facets Modeling Approach

    Thomas Eckes and Kuan-Yu Jin. Detecting Illusory Halo Effects in Rater - Mediated Assessment : A Mixture Rasch Facets Modeling Approach

  44. [54]

    Anjalie Field, Su Lin Blodgett, Zeerak Waseem, and Yulia Tsvetkov. 2021. https://doi.org/10.18653/v1/2021.acl-long.149 A Survey of Race , Racism , and Anti - Racism in NLP . In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th...

  45. [55]

    Eve Fleisig, Genevieve Smith, Madeline Bossi, Ishita Rustagi, Xavier Yin, and Dan Klein. 2024. https://doi.org/10.48550/arXiv.2406.08818 Linguistic Bias in ChatGPT : Language Models Reinforce Dialect Discrimination . arXiv preprint. ArXiv:2406.08818 [cs] version: 1

  46. [56]

    Shuai Gao. 2022. https://aclanthology.org/2022.jeptalnrecital-recital.8 Syst \`e me de traduction automatique neuronale fran c ais-mongol (historique, mise en place et \'e valuations) ( F rench- M ongolian neural machine translation system (history, implementation, and evaluat...

  47. [57]

    Amelia Glaese, Nat McAleese, Maja Trębacz, John Aslanides, Vlad Firoiu, Timo Ewalds, Maribeth Rauh, Laura Weidinger, Martin Chadwick, Phoebe Thacker, Lucy Campbell-Gillingham, Jonathan Uesato, Po-Sen Huang, Ramona Comanescu, Fan Yang, Abigail See, Sumanth Dathathri, Rory Greig...

  48. [58]

    Gordon, Michelle S

    Mitchell L. Gordon, Michelle S. Lam, Joon Sung Park, Kayur Patel, Jeff Hancock, Tatsunori Hashimoto, and Michael S. Bernstein. 2022. https://doi.org/10.1145/3491102.3502004 Jury Learning : Integrating Dissenting Voices into Machine Learning Models . In Proceedings of the 2022 ...

  49. [59]

    Jason Grissom, Susanna Loeb, and Benjamin Master. 2013. https://cepa.stanford.edu/content/effective-instructional-time-use-school-leaders-longitudinal-evidence-observations-principals Effective Instructional Time Use for School Leaders : Longitudinal Evidence from Observations...

  50. [60]

    Louis Guttman. 1945. https://doi.org/10.1007/BF02288892 A basis for analyzing test-retest reliability . Psychometrika, 10(4):255--282

  51. [61]

    Zaretta Hammond. 2015. Culturally responsive teaching and the brain: promoting authentic engagement and rigor among culturally and linguistically diverse students. Corwin, a SAGE company, Thousand Oaks, California. OCLC: ocn889185083

  52. [62]

    Moritz Hardt, Eric Price, and Nathan Srebro. 2016. https://doi.org/10.48550/arXiv.1610.02413 Equality of Opportunity in Supervised Learning . arXiv preprint. ArXiv:1610.02413 [cs]

  53. [63]

    Mike Hardy. 2021. https://doi.org/10.48550/arXiv.2112.11973 Toward Educator -focused Automated Scoring Systems for Reading and Writing . arXiv preprint. ArXiv:2112.11973 [cs]

  54. [64]

    Ursula Hebert-Johnson, Michael Kim, Omer Reingold, and Guy Rothblum. 2018. https://proceedings.mlr.press/v80/hebert-johnson18a.html Multicalibration: Calibration for the ( Computationally - Identifiable ) Masses . In Proceedings of the 35th International Conference on Machine ...

  55. [65]

    Peter Henderson, Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen, Ruoxi Jia, and Prateek Mittal. 2024. Safety Risks from Customizing Foundation Models via Fine -tuning

  56. [66]

    Juyeon Heo, Christina Heinze-Deml, Oussama Elachqar, Shirley Ren, Udhay Nallasamy, Andy Miller, Kwan Ho Ryan Chan, and Jaya Narain. 2024. https://arxiv.org/abs/2410.14516v1 Do LLMs "know" internally when they follow instructions?

  57. [67]

    Hill, Merrie L

    Heather C. Hill, Merrie L. Blunk, Charalambos Y. Charalambous, Jennifer M. Lewis, Geoffrey C. Phelps, Laurie Sleep, and Deborah Loewenberg Ball. 2008. https://www.jstor.org/stable/27739893 Mathematical Knowledge for Teaching and the Mathematical Quality of Instruction : An Exp...

  58. [68]

    Hill, Charalambos Y

    Heather C. Hill, Charalambos Y. Charalambous, David Blazar, Daniel McGinn, Matthew A. Kraft, Mary Beisiegel, Andrea Humez, Erica Litke, and Kathleen Lynch. 2012 a . https://doi.org/10.1080/10627197.2012.715019 Validating Arguments for Observational Instruments : Attending to M...

  59. [69]

    Hill, Charalambos Y

    Heather C. Hill, Charalambos Y. Charalambous, and Matthew A. Kraft. 2012 b . https://doi.org/10.3102/0013189X12437203 When Rater Reliability Is Not Enough : Teacher Observation Systems and a Case for the Generalizability Study . Educational Researcher, 41(2):56--64. Publisher:...

  60. [70]

    Ho and Thomas J

    Andrew D. Ho and Thomas J. Kane. 2013. https://eric.ed.gov/?id=ED540957 The Reliability of Classroom Observations by School Personnel . Research Paper . MET Project . Technical report, Bill & Melinda Gates Foundation. Publication Title: Bill & Melinda Gates Foundation ERIC Num...

  61. [71]

    Valentin Hofmann, Pratyusha Ria Kalluri, Dan Jurafsky, and Sharese King. 2024 a . https://doi.org/10.1038/s41586-024-07856-5 AI generates covertly racist decisions about people based on their dialect . Nature, pages 1--8. Publisher: Nature Publishing Group

  62. [72]

    Valentin Hofmann, Pratyusha Ria Kalluri, Dan Jurafsky, and Sharese King. 2024 b . https://doi.org/10.48550/arXiv.2403.00742 Dialect prejudice predicts AI decisions about people's character, employability, and criminality . arXiv preprint. ArXiv:2403.00742 [cs]

  63. [73]

    Tom Hosking, Phil Blunsom, and Max Bartolo. 2024. https://doi.org/10.48550/arXiv.2309.16349 Human Feedback is not Gold Standard . arXiv preprint. ArXiv:2309.16349

  64. [74]

    Amin Hosseiny Marani, Joshua Levine, and Eric P.S. Baumer. 2022. https://doi.org/10.1145/3511808.3557410 One Rating to Rule Them All ? Evidence of Multidimensionality in Human Assessment of Topic Labeling Quality . In Proceedings of the 31st ACM International Conference on Inf...

  65. [75]

    Qi (Helen) Huang and Daniel M. Bolt. 2023. https://doi.org/10.3758/s13428-023-02275-2 Unipolar IRT and the Author Recognition Test ( ART ) . Behavior Research Methods

  66. [76]

    Jacobs, Ryan J

    Cassandra L. Jacobs, Ryan J. Hubbard, and Kara D. Federmeier. 2022. https://aclanthology.org/2022.scil-1.22 Masked language models directly encode linguistic uncertainty . In Proceedings of the Society for Computation in Linguistics 2022, pages 225--228, online. Association fo...

  67. [77]

    Xuejun (Ryan) Ji. 2023. https://doi.org/10.14288/1.0437518 Using cross-classified mixed effects model for validation studies : a flexible and pragmatic validation method . Ph.D. thesis, University of British Columbia

  68. [78]

    Irina Jurenka, Markus Kunesch, Kevin R McKee, Daniel Gillick, Shaojian Zhu, Shubham Milind Phal, Katherine Hermann, Daniel Kasenberg, Avishkar Bhoopchand, Ankit Anand, Miruna Pîslar, Stephanie Chan, Lisa Wang, Jennifer She, Parsa Mahmoudieh, Wei-Jen Ko, Andrea Huber, Brett Wil...

  69. [79]

    Thomas Kane, Heather Hill, and Douglas Staiger. 2015. https://doi.org/10.3886/ICPSR36095.V4 National Center for Teacher Effectiveness Main Study : Version 4

  70. [80]

    Kane, Daniel F

    Thomas J. Kane, Daniel F. McCaffrey, Trey Miller, and Douglas O. Staiger. 2013. https://eric.ed.gov/?id=ED540959 Have We Identified Effective Teachers ? Validating Measures of Effective Teaching Using Random Assignment . Research Paper . MET Project . Technical report, Bill & ...

  71. [81]

    Kane and Douglas O

    Thomas J. Kane and Douglas O. Staiger. 2012. https://eric.ed.gov/?id=ED540960 Gathering Feedback for Teaching : Combining High - Quality Observations with Student Surveys and Achievement Gains . Research Paper . MET Project . Technical report, Bill & Melinda Gates Foundation. ...

  72. [82]

    Maximilian Kasy and Rediet Abebe. 2021. https://doi.org/10.1145/3442188.3445919 Fairness, Equality , and Power in Algorithmic Decision - Making . In Proceedings of the 2021 ACM Conference on Fairness , Accountability , and Transparency , pages 576--586, Virtual Event Canada. ACM

  73. [83]

    Gabriella Kazai, Jaap Kamps, and Natasa Milic-Frayling. 2013. https://doi.org/10.1007/s10791-012-9205-0 An analysis of human factors and label accuracy in crowdsourcing relevance judgments . Information Retrieval, 16(2):138--178

  74. [84]

    Olney, Patrick Donnelly, Martin Nystrand, and Sidney K

    Sean Kelly, Andrew M. Olney, Patrick Donnelly, Martin Nystrand, and Sidney K. D’Mello. 2018. https://doi.org/10.3102/0013189X18785613 Automatically Measuring Question Authenticity in Real - World Classrooms . Educational Researcher, 47(7):451--464. Publisher: American Educatio...

  75. [85]

    Douwe Kiela, Max Bartolo, Yixin Nie, Divyansh Kaushik, Atticus Geiger, Zhengxuan Wu, Bertie Vidgen, Grusha Prasad, Amanpreet Singh, Pratik Ringshia, Zhiyi Ma, Tristan Thrush, Sebastian Riedel, Zeerak Waseem, Pontus Stenetorp, Robin Jia, Mohit Bansal, Christopher Potts, and Adi...

  76. [86]

    Been Kim, Martin Wattenberg, Justin Gilmer, Carrie Cai, James Wexler, Fernanda Viegas, and Rory Sayres. 2018. https://doi.org/10.48550/arXiv.1711.11279 Interpretability Beyond Feature Attribution : Quantitative Testing with Concept Activation Vectors ( TCAV ) . arXiv preprint....

  77. [87]

    Kingma and Jimmy Ba

    Diederik P. Kingma and Jimmy Ba. 2017. https://doi.org/10.48550/arXiv.1412.6980 Adam: A Method for Stochastic Optimization . arXiv preprint. ArXiv:1412.6980 [cs]

  78. [88]

    David Klahr. 2013. https://doi.org/10.1073/pnas.1212738110 What do we mean? On the importance of not abandoning scientific rigor when talking about science education . Proceedings of the National Academy of Sciences, 110(supplement\_3):14075--14080. Publisher: Proceedings of t...

  79. [89]

    Kromrey, Robert H

    J. Kromrey, Robert H. Fay, and Aarti P. Bellara. 2008. https://www.semanticscholar.org/paper/Macro-for-Computing-Confidence-Intervals-for-Kromrey-Fay/62c7827b2d2ebb01dd9cd78a757001513f335141 Macro for Computing Confidence Intervals for Disattenuated Correlation Coefficients

  80. [90]

    Doug Lemov. 2021. Teach like a champion 3.0: 63 techniques that put students on the path to college, third edition edition. Jossey-Bass, a Wiley imprint, Hoboken, NJ

  81. [91]

    Doug Lemov and Norman Atkins. 2015. Teach like a champion 2.0: 62 techniques that put students on the path to college, second edition edition. Jossey-Bass, San Francisco, CA

  82. [92]

    Zehan Li, Xin Zhang, Yanzhao Zhang, Dingkun Long, Pengjun Xie, and Meishan Zhang. 2023. https://doi.org/10.48550/arXiv.2308.03281 Towards General Text Embeddings with Multi -stage Contrastive Learning . arXiv preprint. ArXiv:2308.03281 [cs]

  83. [93]

    Peter Liljedahl, Tracy Johnston Zager, and Laura Wheeler. 2021. Building thinking classrooms in mathematics: 14 teaching practices for enhancing learning: Grades K -12 . Corwin Mathematics . Corwin, Thousand Oaks, California London New Delhi Singapore

  84. [94]

    Jing Liu and Julie Cohen. 2021. https://doi.org/10.3102/01623737211009267 Measuring Teaching Practices at Scale : A Novel Application of Text -as- Data Methods . Educational Evaluation and Policy Analysis, 43(4):587--614. Publisher: American Educational Research Association

  85. [95]

    Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang

    Nelson F. Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang. 2023 a . https://doi.org/10.48550/arXiv.2307.03172 Lost in the Middle : How Language Models Use Long Contexts . arXiv preprint. ArXiv:2307.03172 [cs]

  86. [96]

    Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. 2023 b . https://doi.org/10.48550/arXiv.2303.16634 G- Eval : NLG Evaluation using GPT -4 with Better Human Alignment . arXiv preprint. ArXiv:2303.16634 [cs]

  87. [97]

    Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. https://doi.org/10.48550/arXiv.1907.11692 RoBERTa : A Robustly Optimized BERT Pretraining Approach . arXiv preprint. ArXiv:1907.11692 [cs]

  88. [98]

    Scott Lundberg and Su-In Lee. 2017. https://doi.org/10.48550/arXiv.1705.07874 A Unified Approach to Interpreting Model Predictions . arXiv preprint. ArXiv:1705.07874 [cs, stat]

  89. [99]

    French, and Helen Patrick

    Panayota Mantzicopoulos, Brian F. French, and Helen Patrick. 2018. https://doi.org/10.1080/10409289.2018.1477903 The Mathematical Quality of Instruction ( MQI ) in Kindergarten : An Evaluation of the Stability of the MQI Using Generalizability Theory . Early Education and Deve...

  90. [100]

    Mariano and Brian W

    Louis T. Mariano and Brian W. Junker. 2007. https://doi.org/10.3102/1076998606298033 Covariates of the Rating Process in Hierarchical Models for Multiple Ratings of Test Items . Journal of Educational and Behavioral Statistics, 32(3):287--314

  91. [101]

    Thomas McCoy, Shunyu Yao, Dan Friedman, Matthew Hardy, and Thomas L

    R. Thomas McCoy, Shunyu Yao, Dan Friedman, Matthew Hardy, and Thomas L. Griffiths. 2023. https://arxiv.org/abs/2309.13638v1 Embers of Autoregression : Understanding Large Language Models Through the Problem They are Trained to Solve

  92. [102]

    Samuel Messick. 1998. https://www.jstor.org/stable/27522333 Test Validity : A Matter of Consequence . Social Indicators Research, 45(1/3):35--44. Publisher: Springer

  93. [103]

    Muchinsky

    Paul M. Muchinsky. 1996. https://doi.org/10.1177/0013164496056001004 The Correction for Attenuation . Educational and Psychological Measurement, 56(1):63--75. Publisher: SAGE Publications Inc

  94. [104]

    Eiji Muraki. 1992. https://doi.org/10.1177/014662169201600206 A Generalized Partial Credit Model : Application of an EM Algorithm . Applied Psychological Measurement, 16(2):159--176. Publisher: SAGE Publications Inc

  95. [105]

    Murphy and S

    Daniel L. Murphy and S. Natasha Beretvas. 2015. https://doi.org/10.1080/08957347.2015.1042158 A Comparison of Teacher Effectiveness Measures Calculated Using Three Multilevel Models for Raters Effects . Applied Measurement in Education, 28(3):219--236. Publisher: Routledge \_e...

  96. [106]

    You Gotta be a Doctor , Lin

    Huy Nghiem, John Prindle, Jieyu Zhao, and Hal Daumé III. 2024. https://doi.org/10.48550/arXiv.2406.12232 " You Gotta be a Doctor , Lin ": An Investigation of Name - Based Bias of Large Language Models in Employment Recommendations . arXiv preprint. ArXiv:2406.12232

  97. [107]

    Patz, Brian W

    Richard J. Patz, Brian W. Junker, Matthew S. Johnson, and Louis T. Mariano. 2002. https://www.jstor.org/stable/3648122 The Hierarchical Rater Model for Rated Test Items and Its Application to Large - Scale Educational Assessment Data . Journal of Educational and Behavioral Sta...

  98. [108]

    Pianta and Bridget K

    Robert C. Pianta and Bridget K. Hamre. 2009. https://doi.org/10.3102/0013189X09332374 Conceptualization, Measurement , and Improvement of Classroom Processes : Standardized Observation Can Leverage Capacity . Educational Researcher, 38(2):109--119. Publisher: American Educatio...

  99. [109]

    Pianta, Karen M

    Robert C. Pianta, Karen M. La Paro, and Bridget K. Hamre. 2008. Classroom Assessment Scoring System ( CLASS ) Manual , K -3 . Paul H. Brookes Publishing Company. Google-Books-ID: NBeaGgAACAAJ

  100. [110]

    Weinberger

    Geoff Pleiss, Manish Raghavan, Felix Wu, Jon Kleinberg, and Kilian Q. Weinberger. 2017. https://doi.org/10.48550/arXiv.1709.02012 On Fairness and Calibration . arXiv preprint. ArXiv:1709.02012 [cs, stat]

  101. [111]

    Martyn Plummer. 2003. JAGS : A program for analysis of Bayesian graphical models using Gibbs sampling. Working Papers

  102. [112]

    Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen, Ruoxi Jia, Prateek Mittal, and Peter Henderson. 2023. https://doi.org/10.48550/arXiv.2310.03693 Fine-tuning Aligned Language Models Compromises Safety , Even When Users Do Not Intend To ! arXiv preprint. ArXiv:2310.03693

  103. [113]

    Marco Tulio Ribeiro, Tongshuang Wu, Carlos Guestrin, and Sameer Singh. 2020. https://doi.org/10.18653/v1/2020.acl-main.442 Beyond Accuracy : Behavioral Testing of NLP Models with CheckList . In Proceedings of the 58th Annual Meeting of the Association for Computational Linguis...

  104. [114]

    Rickford and Sharese King

    John R. Rickford and Sharese King. 2016. https://doi.org/10.1353/lan.2016.0078 Language and linguistics on trial: Hearing Rachel Jeantel (and other vernacular speakers) in the courtroom and beyond . Language, 92(4):948--988

  105. [115]

    Cynthia Rudin. 2019. https://doi.org/10.48550/arXiv.1811.10154 Stop Explaining Black Box Machine Learning Models for High Stakes Decisions and Use Interpretable Models Instead . arXiv preprint. ArXiv:1811.10154 [cs, stat]

  106. [116]

    Olney, Sean Kelly, Martin Nystrand, Sidney D'Mello, Nathan Blanchard, Xiaoyi Sun, Marcy Glaus, and Art Graesser

    Borhan Samei, Andrew M. Olney, Sean Kelly, Martin Nystrand, Sidney D'Mello, Nathan Blanchard, Xiaoyi Sun, Marcy Glaus, and Art Graesser. 2014. https://eric.ed.gov/?id=ED566380 Domain Independent Assessment of Dialogic Properties of Classroom Discourse . Technical report. Publi...

  107. [117]

    Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2020. https://doi.org/10.48550/arXiv.1910.01108 DistilBERT , a distilled version of BERT : smaller, faster, cheaper and lighter . arXiv preprint. ArXiv:1910.01108

  108. [118]

    Jon Saphier, Mary Ann Haley-Speca, and Robert Gower. 2008. The skillful teacher: building your teaching skills, 6th ed edition. Research for Better Teaching, Acton, Mass

  109. [119]

    Schwartz, Jessica M

    Daniel L. Schwartz, Jessica M. Tsang, and Kristen P. Blair. 2016. The ABCs of how we learn: 26 scientifically proven approaches, how they work, and when to use them , first edition edition. Norton books in education. W.W. Norton & Company, New York

  110. [120]

    Mark D. Shermis. 2014. https://doi.org/10.1016/j.asw.2013.04.001 State-of-the-art automated essay scoring: Competition , results, and future directions from a United States demonstration . Assessing Writing, 20:53--76

  111. [121]

    Evan Shieh, Faye-Marie Vassel, Cassidy Sugimoto, and Thema Monroe-White. 2024. https://doi.org/10.48550/arXiv.2404.07475 Laissez- Faire Harms : Algorithmic Biases in Generative Language Models . arXiv preprint. ArXiv:2404.07475

  112. [122]

    Robert E. Slavin. 2002. https://doi.org/10.3102/0013189X031007015 Evidence- Based Education Policies : Transforming Educational Practice and Research . Educational Researcher, 31(7):15--21. Publisher: American Educational Research Association

  113. [123]

    Jiaming Song, Pratyusha Kalluri, Aditya Grover, Shengjia Zhao, and Stefano Ermon. 2020. https://doi.org/10.48550/arXiv.1812.04218 Learning Controllable Fair Representations . arXiv preprint. ArXiv:1812.04218 [cs, stat]

  114. [124]

    Mukund Sundararajan, Ankur Taly, and Qiqi Yan. 2017. https://proceedings.mlr.press/v70/sundararajan17a.html Axiomatic Attribution for Deep Networks . In Proceedings of the 34th International Conference on Machine Learning , pages 3319--3328. PMLR. ISSN: 2640-3498

  115. [125]

    Martin, and Tamara Sumner

    Abhijit Suresh, Jennifer Jacobs, Charis Harty, Margaret Perkoff, James H. Martin, and Tamara Sumner. 2022. https://aclanthology.org/2022.lrec-1.497 The T alk M oves dataset: K-12 mathematics lesson transcripts annotated for teacher and student discursive moves . In Proceedings...

  116. [126]

    Anaïs Tack, Ekaterina Kochmar, Zheng Yuan, Serge Bibauw, and Chris Piech. 2023. https://doi.org/10.48550/arXiv.2306.06941 The BEA 2023 Shared Task on Generating AI Teacher Responses in Educational Dialogues . arXiv preprint. ArXiv:2306.06941

  117. [127]

    https://www.r-project.org/ R: A Language and Environment for Statistical Computing

    R Core Team. https://www.r-project.org/ R: A Language and Environment for Statistical Computing

  118. [128]

    Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, W...

  119. [129]

    UpLevel. 2024. https://resources.uplevelteam.com/gen-ai-for-coding Gen AI for Coding Research Report . Technical report, Uplevel Data Labs

  120. [130]

    Michelle Vaccaro, Abdullah Almaatouq, and Thomas Malone. 2024. https://doi.org/10.48550/arXiv.2405.06087 When Are Combinations of Humans and AI Useful ? arXiv preprint. ArXiv:2405.06087 [cs]

  121. [131]

    Chris van der Lee, Albert Gatt, Emiel van Miltenburg, Sander Wubben, and Emiel Krahmer. 2019. https://doi.org/10.18653/v1/W19-8643 Best practices for the human evaluation of automatically generated text . In Proceedings of the 12th International Conference on Natural Language ...

  122. [132]

    Jiarui Wang, Richong Zhang, Junfan Chen, Jaein Kim, and Yongyi Mao. 2022. https://doi.org/10.18653/v1/2022.emnlp-main.521 Text style transferring via adversarial masking and styled filling . In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Process...

  123. [133]

    Rose Wang and Dorottya Demszky. 2023. https://doi.org/10.18653/v1/2023.bea-1.53 Is ChatGPT a Good Teacher Coach ? Measuring Zero - Shot Performance For Scoring and Providing Actionable Insights on Classroom Instruction . In Proceedings of the 18th Workshop on Innovative Use of...

  124. [134]

    Melissa Warr, Nicole Jakubczyk Oster, and Roger Isaac. 2024. https://doi.org/10.1080/15391523.2024.2395295 Implicit bias in large language models: Experimental proof and implications for education . Journal of Research on Technology in Education, 0(0):1--24. Publisher: Routled...

  125. [135]

    Zeerak Waseem. 2016. https://doi.org/10.18653/v1/W16-5618 Are You a Racist or Am I Seeing Things ? Annotator Influence on Hate Speech Detection on Twitter . In Proceedings of the First Workshop on NLP and Computational Social Science , pages 138--142, Austin, Texas. Associatio...

  126. [136]

    Albert Webson, Alyssa Loo, Qinan Yu, and Ellie Pavlick. 2023. https://doi.org/10.18653/v1/2023.findings-emnlp.514 Are Language Models Worse than Humans at Following Prompts ? It 's Complicated . In Findings of the Association for Computational Linguistics : EMNLP 2023 , pages ...

  127. [137]

    Albert Webson and Ellie Pavlick. 2022. https://doi.org/10.18653/v1/2022.naacl-main.167 Do Prompt - Based Models Really Understand the Meaning of Their Prompts ? In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics...

  128. [138]

    Jacob Whitehill and Jennifer LoCasale-Crouch. 2024. https://doi.org/10.48550/arXiv.2310.01132 Automated Evaluation of Classroom Instructional Support with LLMs and BoWs : Connecting Global Predictions to Specific Feedback . arXiv preprint. ArXiv:2310.01132 [cs]

  129. [139]

    Whitehurst, Matthew M

    Grover J. Whitehurst, Matthew M. Chingos, and Katharine M. Lindquist. 2014. Evaluating Teachers with Classroom Observations : Lessons Learned in Four Districts . Technical report, Brookings Institution. Publication Title: Brookings Institution ERIC Number: ED553815

  130. [140]

    Stefanie A. Wind. 2019. https://doi.org/10.1111/jedm.12222 Nonparametric Evidence of Validity , Reliability , and Fairness for Rater - Mediated Assessments : An Illustration Using Mokken Scale Analysis . Journal of Educational Measurement, 56(3):478--504. \_eprint: https://onl...

  131. [141]

    Wind and Wenjing Guo

    Stefanie A. Wind and Wenjing Guo. 2019. https://doi.org/10.1177/0013164419834613 Exploring the Combined Effects of Rater Misfit and Differential Rater Functioning in Performance Assessments . Educational and Psychological Measurement, 79(5):962--987. Publisher: SAGE Publications Inc

  132. [142]

    Paiheng Xu, Jing Liu, Nathan Jones, Julie Cohen, and Wei Ai. 2024. https://doi.org/10.48550/arXiv.2404.02444 The Promises and Pitfalls of Using Language Models to Measure Instruction Quality in Education . arXiv preprint. ArXiv:2404.02444

  133. [143]

    Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Ruslan Salakhutdinov, and Quoc V. Le. 2020. https://doi.org/10.48550/arXiv.1906.08237 XLNet : Generalized Autoregressive Pretraining for Language Understanding . arXiv preprint. ArXiv:1906.08237

  134. [144]

    Rich Zemel, Yu Wu, Kevin Swersky, Toni Pitassi, and Cynthia Dwork. 2013. https://proceedings.mlr.press/v28/zemel13.html Learning Fair Representations . In Proceedings of the 30th International Conference on Machine Learning , pages 325--333. PMLR. ISSN: 1938-7228

  135. [145]

    Shengjia Zhao and Stefano Ermon. 2021. https://doi.org/10.48550/arXiv.2011.07476 Right Decisions from Wrong Predictions : A Mechanism Design Alternative to Individual Calibration . arXiv preprint. ArXiv:2011.07476 [cs, math, stat]

  136. [146]

    Hwang, Xiang Ren, and Maarten Sap

    Kaitlyn Zhou, Jena D. Hwang, Xiang Ren, and Maarten Sap. 2024. https://doi.org/10.48550/arXiv.2401.06730 Relying on the Unreliable : The Impact of Language Models ' Reluctance to Express Uncertainty . arXiv preprint. ArXiv:2401.06730 [cs]

  137. [147]

    Quintana, Anita Delahay, and Xu Wang

    Xiaofei Zhou, Christopher Kok, Rebecca M. Quintana, Anita Delahay, and Xu Wang. 2023. https://doi.org/10.1145/3573051.3593388 How Learning Experience Designers Make Design Decisions : The Role of Data , the Reliance on Subject Matter Expertise , and the Opportunities for Data ...

  138. [148]

    Eva A. O. Zijlmans, Jesper Tijmstra, L. Andries van der Ark, and Klaas Sijtsma. 2018 a . https://doi.org/10.1177/0013164417728358 Item- Score Reliability in Empirical - Data Sets and Its Relationship With Other Item Indices . Educational and Psychological Measurement, 78(6):99...

  139. [149]

    Eva A. O. Zijlmans, L. Andries van der Ark, Jesper Tijmstra, and Klaas Sijtsma. 2018 b . https://doi.org/10.1177/0146621618758290 Methods for Estimating Item - Score Reliability . Applied Psychological Measurement, 42(7):553--570. Publisher: SAGE Publications Inc

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.