Pith. sign in

REVIEW 2 major objections 4 minor 58 references

Everyday human–LLM conversations carry measurable signs of informal learning — about 5% of user turns reach constructive engagement — and scaffolded assistant support, especially feedback and explanation, marks richer participation.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 17:23 UTC pith:SNXKMRF5

load-bearing objection A credible large-scale measurement paper whose descriptive prevalence result holds up, but whose scaffolding-engagement association is entangled with the annotation method and needs a paired audit before it convinces. the 2 major comments →

arxiv 2607.17643 v2 pith:SNXKMRF5 submitted 2026-07-20 cs.HC cs.CY

Informal Learning Emerges in Everyday Human-LLM Interaction

classification cs.HC cs.CY
keywords human-LLM interactioninformal learningconstructive engagementcognitive engagementscaffoldinglearner engagementconversational learninglarge language models
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper asks whether ordinary, unplanned use of large language models — writing, coding, troubleshooting — preserves opportunities to learn rather than simply offloading cognitive work to the machine. Analyzing 128,569 naturalistic conversations and 491,685 user turns, it finds that roughly a third of user turns show cognitive engagement and about 5% reach constructive engagement, the deepest observable level, where users test, revise, or extend ideas instead of just accepting an answer. These deeper moments are not scattered randomly: they cluster under explicit learning-oriented framing, in coding tasks, in longer exchanges, and after assistant turns that scaffold the user's next step rather than deliver a finished answer. If these behavioral signatures track learning, everyday AI use is a genuine — though selective — informal learning ecology, and preserving the cognitive opportunities people exercise while solving real problems becomes a measurable design goal.

Core claim

The central claim: everyday human–LLM interaction contains measurable, selective, and conditionally organized behavioral signatures of informal learning — not just answer delivery or cognitive offloading. Across 491,685 user turns from 128,569 conversations, 31.9% showed cognitive engagement and 4.9% reached constructive engagement, defined as users elaborating, testing, revising, or extending ideas. Constructive engagement was amplified by explicit learning-oriented framing, more common in coding than writing (5.4–15.2% vs 1.5–5.2% of turns), and more likely in longer exchanges. Conversations with at least one scaffolded assistant turn — support that leaves reasoning with the user — showed

What carries the argument

The paper's central instrument is a role-specific, turn-level annotation scheme that converts learning-science constructs into observable labels. User turns are coded on a depth hierarchy: passive receipt, active use or follow-up, and constructive engagement — the deepest level, reserved for turns where users generate, revise, test, or extend ideas. Assistant turns are coded into reference responses (which deliver the requested answer) versus scaffolded support (which diagnoses, structures, or redirects the user's process while preserving the user's reasoning role), and scaffolded turns are further tagged by support intent and by six non-exclusive support forms: feedback, hinting, instructin

Load-bearing premise

The load-bearing premise is that the LLM-assisted annotation pipeline — human-audited on 600 user turns and 180 assistant turns, using an annotator from the same model family as the systems under study — correctly measures constructive engagement and scaffolded support, and that constructive engagement is a valid observable proxy for informal learning; if the annotator over-labels user turns as constructive when they follow scaffolded assistant turns, the central support–enga

What would settle it

Mask the assistant turn in a sample of scaffolded-assistant→constructive-user pairs and have independent human coders label the user turns alone; if 'constructive' labels largely vanish without the scaffold context, the central association is an annotator artifact. A complementary test: randomly assign matched tasks to scaffolded versus direct-answer assistants and compare constructive-turn rates; if scaffolding shows no lift under random assignment, the observed association reflects selection of who receives scaffolding rather than an effect of support.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • If constructive engagement tracks informal learning, everyday LLM use is already a site of measurable learning behavior — roughly one in twenty user turns involves users testing, revising, or extending ideas rather than merely consuming answers.
  • Scaffolded assistant support, especially feedback and explanation, is consistently associated with richer constructive participation; the way an assistant responds, not just the accuracy of its answer, becomes a relevant dimension of its educational value.
  • Task ecology matters: constructive engagement ran at 5.4–15.2% of user turns in coding conversations versus 1.5–5.2% in writing, with coding roughly doubling the odds of a constructive turn in the context model — tasks that externalize errors and tests create the occasions where sense-making becomes visible.
  • Because support–engagement coupling is state- and form-dependent — explanation predicts constructive follow-up after any prior user state while feedback works mainly after passive turns — the timing of support within an exchange shapes whether users go deeper.
  • The findings shift evaluation of AI assistants from answer-delivery efficiency toward the preservation of cognitive opportunities: whether users reason, test ideas, and construct understanding during everyday problem-solving.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • My inference: if the support–engagement association is causal — which this observational design cannot establish — response design could deliberately deploy feedback and explanation to preserve cognitive opportunity, making 'constructive engagement rate' a candidate evaluation metric for assistants alongside task success.
  • My inference: the feedback label is defined as requiring a target in the user's preceding contribution, so feedback can only occur after users have put work on the table; part of the feedback–constructive association may therefore be definitional rather than a pure effect of the support form.
  • My inference: the open empirical question is whether these behavioral signatures predict actual learning outcomes such as retention, transfer, or skill growth; a longitudinal study linking constructive-turn rates to later performance would test whether the process indicators do double duty as proxies.
  • My inference: a direct contamination check — masking the assistant turn when independent human coders judge whether a user turn is constructive — would reveal whether the central support–engagement association is inflated by annotator priming.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 4 minor

Summary. This paper presents a large-scale observational study of informal learning in naturalistic human–LLM conversations. Using three public corpora (WildChat, LMSYS Chat, ShareChat) and restricting to coding and writing task settings, the authors develop an LLM-assisted annotation scheme for user engagement (passive/active/constructive) and assistant scaffolding (with intent and form labels). They report that 31.9% of 491,685 user turns exhibit cognitive engagement and 4.9% exhibit constructive engagement; that constructive engagement is more common under explicit learning-oriented framing, in coding tasks, and in longer conversations; and that scaffolded assistant support is positively associated with constructive engagement at conversation, support-form, and adjacent-turn levels (Poisson count ratios 1.57–2.49; adjacent-turn lifts up to +6.2 pp). The paper interprets these as measurable behavioral signatures of informal learning and discusses implications for AI evaluation.

Significance. If the core association survives scrutiny, this is a significant contribution to the human–AI interaction and learning sciences literature. The paper moves beyond laboratory or tutoring settings to a large naturalistic corpus and offers a reproducible measurement pipeline: the annotation protocol, validation tables, source data, and analysis code are archived. The convergence across three corpora and multiple robustness checks (metadata fixed effects, offset-rate sensitivities, FDR adjustments) is a strength. The construct of constructive engagement as a process-level indicator of informal learning is well grounded in ICAP and related theory. However, the central scaffolding–engagement claim currently rests on a measurement assumption that is not fully validated at the joint level, so the headline association should be treated as conditional until that is addressed.

major comments (2)
  1. [§4.4, Table A5, Table B5; Fig. 3b; Fig. 5a] Load-bearing measurement concern. The paper's central claim that scaffolded support is associated with constructive engagement (Section 2.3: Poisson count ratios 1.57–2.49; Section 2.5: adjacent-turn lifts up to +6.2 pp) depends on an LLM-assisted annotation pipeline in which the same GPT-5.1-family model labels both user engagement and assistant scaffolding. The user-turn prompt includes 'preceding local context' (Table A5), so the annotator is not blind to whether the previous assistant turn was scaffolded. The human audits (Table B5: N=600 user turns, N=180 assistant turns) validate the marginal quality of each label, but they are unpaired; they cannot detect differential misclassification in which a user turn is more likely to be coded constructive immediately after a scaffolded assistant turn. Such shared-method covariance could manufacture the support–engagement association indepen
  2. [§4.3.3, Table A4; §2.4–2.5] Construct-level dependency in support-form labels. The definitions in Table A4 create a potential mechanical tie between the user and assistant labels: M1 (feedback) requires 'a target in the user's preceding contribution,' and constructive engagement is defined as the user 'generat[ing], revis[ing] or extend[ing] ideas.' An assistant turn that evaluates a constructive user turn will automatically be feedback. While the state-conditioned analysis (Fig. 5d) partially accounts for the preceding user state, it does not remove the possibility that the LLM annotator applies the same codebook cue to both roles. This is especially relevant for the interpretation of M1 and M4 as 'leaving epistemic work to the user.' A sensitivity analysis re-annotating support forms without access to the user-turn engagement label (or vice versa) would clarify whether the support-form effects are independent of
minor comments (4)
  1. [Abstract; §2.1] The prevalence estimates (31.9% cognitive, 4.9% constructive) are based on conversations passing the minimum-turn screen (at least four message turns) and language/task filters. As written, the abstract may be read as applying to all everyday LLM use; please add a qualifier.
  2. [§2.3] The term 'reference conversations' is used before it is defined. Please define it at first use or point to the Methods section.
  3. [Fig. 5d; Table B19] The M1-feedback OR after prior passive turns (2.16) is based on a sparse subgroup. Please report cell counts or a small-sample robustness check to indicate stability.
  4. [Table B5] The M6 questioning label has F1=0.691 with precision 0.543. The text describes 'acceptable or better reliability' for all support forms; this label is weaker than the others and warrants a caveat.

Circularity Check

0 steps flagged

No circularity: engagement and scaffolding are independently defined empirical measures; the reported associations are observational analyses, not predictions forced by construction.

full rationale

The derivation chain is empirical, not definitional. Constructive engagement (Table A4: 'Higher-depth cognitive work in which the user generates, revises or extends ideas, explanations or constraints') and scaffolded support ('Assistant support that structures, diagnoses, regulates or redirects the user's process while leaving cognitive or task responsibility with the user') are distinct role-specific constructs drawn from external learning-science frameworks (refs 23, 24, 30, 31, 33, 43), annotated on separate turn types, and validated against human-confirmed labels (Table B5). The Section 2.3/2.5 support–engagement associations are observed covariances between these independently defined labels; no parameter is fitted to the target association, and no equation in the paper defines one construct in terms of the other. The same LLM annotator family labels both user engagement and assistant scaffolding, which is a genuine measurement-validity limitation (potential shared-method covariance), but it is not a circular reduction: the paper discloses this in the LLM use statement and Section 4.4, reports human audits and metadata robustness checks, and explicitly limits causal claims ('Assistant support was observed rather than assigned, capturing the naturally occurring covariation and sequencing central to our ecological analysis but limiting causal inference'). There are no self-citations carrying the argument, no imported uniqueness theorems, and no fitted parameter renamed as a prediction. Accordingly, no circular step is exhibited.

Axiom & Free-Parameter Ledger

4 free parameters · 4 axioms · 0 invented entities

No new natural-kind entities are postulated. The paper introduces behavioral labels (constructive engagement, scaffolded support) as measurement constructs, and the hand-chosen thresholds and pipeline choices listed above are the main free parameters. The axioms capture the theory-driven assumptions that link text annotations to informal learning.

free parameters (4)
  • Task-filter score threshold = 0.70
    Hand-chosen threshold for LLM-assisted semantic task filters (Supplementary Table A1). Determines which conversations are classified as coding/writing and thus affects the sample composition.
  • Minimum-turn screen = 4 message turns
    Retained conversations with at least four message turns (Methods 4.2). Excludes very short exchanges and shapes the analysis population.
  • Conversation-length buckets = 2-3, 4-6, 7+ user turns
    Hand-chosen buckets used in Fig. 2c and context models (Supplementary Table B13). The choice affects the magnitude of the length-gradient estimates.
  • Support-form non-exclusive labels = Multi-label M1-M6
    The decision to allow co-occurrence of support-form labels and to compare scaffolded conversations containing a form with those not containing it (Methods 4.3.3) affects the form-contrast interpretation.
axioms (4)
  • domain assumption ICAP framework's passive-active-constructive-interactive distinction is applicable to turn-level text in human-LLM conversation logs.
    The paper adapts Chi & Wylie's ICAP taxonomy (reference 24) to label user turns; this assumes the framework transfers to chat transcripts with coding/writing tasks.
  • domain assumption Public conversation logs from WildChat, LMSYS Chat, and ShareChat are representative of naturalistic everyday human-LLM interaction in English coding and writing.
    The analysis scope is defined by these three corpora and the English/coding/writing filters; generalization to other populations is assumed, with the limitation acknowledged.
  • domain assumption Constructive engagement, as operationalized here, is a valid observable proxy for informal learning behaviors.
    The paper explicitly states this is a process measure, not proof of durable learning. Still, the inference from behavior to 'informal learning' rests on this assumption.
  • ad hoc to paper LLM-assisted labels with human validation provide unbiased measurement.
    The annotation pipeline uses GPT-5.1-family models to generate the labels; human audits are small and may share the same theoretical priors. The entire measurement framework depends on this assumption.

pith-pipeline@v1.3.0-alltime-deepseek · 33385 in / 12119 out tokens · 111164 ms · 2026-08-01T17:23:50.950920+00:00 · methodology

0 comments
read the original abstract

As LLMs become increasingly capable of completing tasks for users, a central concern is that everyday AI use may become primarily cognitive offloading, eroding the opportunities through which people develop their own capabilities. We analyse large-scale human-LLM conversations to ask whether informal learning behaviors also emerge in this setting: whether users engage in exchanges in ways that preserve opportunities to learn. Across 128,569 naturalistic conversations, we translated learning-science constructs into turn-level behavioural signatures. Cognitive engagement, users' cognitive effort as reflected in the exchange, appeared in 31.9% of 491,685 user turns, whereas constructive engagement, the deepest observable form of learning-oriented engagement, appeared in 4.9%, showing that deeper sense-making was recurrent but selective. Our study further identifies factors associated with these forms of engagement. Scaffolded assistant support consistently marked richer constructive participation, with associations varying by user framing, task ecology, support form, timing and prior user state. Together, these findings show that everyday human-LLM interaction is not only answer delivery or cognitive offloading; it also contains measurable, selective and conditionally organized behavioural signatures of informal learning. They shift AI evaluation from answer-delivery efficiency toward the preservation of cognitive opportunities for users to reason, test ideas and construct understanding in the course of everyday problem-solving.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

58 extracted references · 1 canonical work pages

  1. [1]

    Chatterji, A.et al.How people use ChatGPT. Tech. Rep. Working Paper 34255, National Bureau of Economic Research (2025)

  2. [2]

    & Zhang, W

    Noy, S. & Zhang, W. Experimental evidence on the productivity effects of generative artificial intelligence.Science381, 187–192 (2023)

  3. [3]

    & Raymond, L

    Brynjolfsson, E., Li, D. & Raymond, L. Generative AI at work.Q. J. Econ.140, 889–942 (2025)

  4. [4]

    Costa-Gomes, B.et al.Public use of a generalist LLM chatbot for health queries. Nat. Health1, 689–696 (2026)

  5. [5]

    & Simon, H

    Anzai, Y. & Simon, H. A. The theory of learning by doing.Psychol. Rev.86, 124–140 (1979)

  6. [6]

    A.Experiential Learning: Experience as the Source of Learning and Development2nd edn (Pearson FT Press, Upper Saddle River, NJ, 2015)

    Kolb, D. A.Experiential Learning: Experience as the Source of Learning and Development2nd edn (Pearson FT Press, Upper Saddle River, NJ, 2015)

  7. [7]

    Risko, E. F. & Gilbert, S. J. Cognitive offloading.Trends Cogn. Sci.20, 676–688 (2016)

  8. [8]

    AI tools in society: Impacts on cognitive offloading and the future of critical thinking.Societies15, 6 (2025)

    Gerlich, M. AI tools in society: Impacts on cognitive offloading and the future of critical thinking.Societies15, 6 (2025)

  9. [9]

    Barke, S., James, M. B. & Polikarpova, N. Grounded copilot: How programmers interact with code-generating models.Proc. ACM Program. Lang.7, 85–111 (2023)

  10. [10]

    If the machine is as good as me, then what use am I?

    Kobiella, C., Flores L´ opez, Y. S., Waltenberger, F., Draxler, F. & Schmidt, A. “If the machine is as good as me, then what use am I?”—how the use of ChatGPT changes young professionals’ perception of productivity and accomplishment. Proc. CHI Conf. Hum. Factors Comput. Syst.1–16 (2024)

  11. [11]

    Informal learning in the workplace.Stud

    Eraut, M. Informal learning in the workplace.Stud. Contin. Educ.26, 247–273 (2004)

  12. [12]

    Marsick, V. J. & Watkins, K.Informal and incidental learning in the workplace (Routledge Revivals)(Routledge, London, 2015)

  13. [13]

    The forms of informal learning: Towards a conceptualization of the field

    Schugurensky, D. The forms of informal learning: Towards a conceptualization of the field. Tech. Rep. NALL Working Paper 19, Centre for the Study of Education and Work, OISE/UT, Toronto (2000). URL https://hdl.handle.net/1807/2733

  14. [14]

    Kasneci, E.et al.ChatGPT for good? On opportunities and challenges of large language models for education.Learn. Individ. Differ.103, 102274 (2023). 25

  15. [15]

    Meyer, J.et al.Using LLMs to bring evidence-based feedback into the classroom: AI-generated feedback increases secondary students’ text revision, motivation, and positive emotions.Comput. Educ. Artif. Intell.6, 100199 (2024)

  16. [16]

    & Collins-Thompson, K

    Arif, T., Asthana, S. & Collins-Thompson, K. Generation and assessment of multiple-choice questions from video transcripts using large language models. Proc. 11th ACM Conf. Learn. @ Scale530–534 (2024)

  17. [17]

    & Kim, J

    Jin, H., Lee, S., Shin, H. & Kim, J. Teach AI how to code: Using large language models as teachable agents for programming education.Proc. CHI Conf. Hum. Factors Comput. Syst.1–28 (2024)

  18. [18]

    & Hong, H

    Lim, H., Choi, D. & Hong, H. Identify design problems through questioning: Exploring role-playing interactions with large language models to foster design questioning skills.Companion Proc. CSCW598–602 (2024)

  19. [19]

    & Gaˇ sevi´ c, D

    Yan, L., Greiff, S., Teuber, Z. & Gaˇ sevi´ c, D. Promises and challenges of generative artificial intelligence for human learning.Nat. Hum. Behav.8, 1839–1850 (2024)

  20. [20]

    F¨ utterer, T.et al.Why using AI as a learning coach can improve education.Nat. Hum. Behav.(2026). URL https://doi.org/10.1038/s41562-026-02524-2

  21. [21]

    & Kasneci, E

    Terzimehi´ c, N., B¨ uhler, B. & Kasneci, E. Conversational AI as a catalyst for infor- mal learning: An empirical large-scale study on LLM use in everyday learning. Comput. Educ. Artif. Intell.11, 100634 (2026)

  22. [22]

    The relative effectiveness of human tutoring, intelligent tutoring systems, and other tutoring systems.Educ

    VanLehn, K. The relative effectiveness of human tutoring, intelligent tutoring systems, and other tutoring systems.Educ. Psychol.46, 197–221 (2011)

  23. [23]

    A., Blumenfeld, P

    Fredricks, J. A., Blumenfeld, P. C. & Paris, A. H. School engagement: Potential of the concept, state of the evidence.Rev. Educ. Res.74, 59–109 (2004)

  24. [24]

    Chi, M. T. & Wylie, R. The ICAP framework: Linking cognitive engagement to active learning outcomes.Educ. Psychol.49, 219–243 (2014)

  25. [25]

    Koedinger, K. R. & Aleven, V. Exploring the assistance dilemma in experiments with cognitive tutors.Educ. Psychol. Rev.19, 239–264 (2007)

  26. [26]

    12th Int

    Zhao, W.et al.WildChat: 1M ChatGPT interaction logs in the wild.Proc. 12th Int. Conf. Learn. Represent.(2024). URL https://openreview.net/forum? id=Bl8u7ZRlbM

  27. [27]

    12th Int

    Zheng, L.et al.LMSYS-Chat-1M: A large-scale real-world LLM conversation dataset.Proc. 12th Int. Conf. Learn. Represent.(2024). URL https://openreview. net/forum?id=BOfDKxfwt0

  28. [28]

    Yan, Y., Nguyen, T., Su, B., Lieffers, M. & Le, T. ShareChat: A dataset of chatbot conversations in the wild (2025). Preprint at https://doi.org/10.48550/ 26 arXiv.2512.17843

  29. [29]

    Preprint at https://doi.org/10.48550/ arXiv.2503.04761

    Handa, K.et al.Which economic tasks are performed with AI? Evidence from millions of Claude conversations (2025). Preprint at https://doi.org/10.48550/ arXiv.2503.04761

  30. [30]

    Wood, D., Bruner, J. S. & Ross, G. The role of tutoring in problem solving.J. Child Psychol. Psychiatry17, 89–100 (1976)

  31. [31]

    & Beishuizen, J

    Van de Pol, J., Volman, M. & Beishuizen, J. Scaffolding in teacher–student interaction: A decade of research.Educ. Psychol. Rev.22, 271–296 (2010)

  32. [32]

    & Leimeister, J

    Winkler, R., Hobert, S., Salovaara, A., S¨ ollner, M. & Leimeister, J. M. Sara, the lecturer: Improving learning in online education with a scaffolding-based conversational agent.Proc. CHI Conf. Hum. Factors Comput. Syst.1–14 (2020)

  33. [33]

    Quintana, C.et al.A scaffolding design framework for software to support science inquiry.J. Learn. Sci.13, 337–386 (2004)

  34. [34]

    & Heeren, B

    Keuning, H., Jeuring, J. & Heeren, B. A systematic literature review of automated feedback generation for programming exercises.ACM Trans. Comput. Educ.19, 1–43 (2018)

  35. [35]

    & Hayes, J

    Flower, L. & Hayes, J. R. A cognitive process theory of writing.Coll. Compos. Commun.32, 365–387 (1981)

  36. [36]

    Revision strategies of student writers and experienced adult writers

    Sommers, N. Revision strategies of student writers and experienced adult writers. Coll. Compos. Commun.31, 378–388 (1980)

  37. [37]

    & Timperley, H

    Hattie, J. & Timperley, H. The power of feedback.Rev. Educ. Res.77, 81–112 (2007)

  38. [38]

    Shute, V. J. Focus on formative feedback.Rev. Educ. Res.78, 153–189 (2008)

  39. [39]

    Nicol, D. J. & Macfarlane-Dick, D. Formative assessment and self-regulated learn- ing: A model and seven principles of good feedback practice.Stud. High. Educ. 31, 199–218 (2006)

  40. [40]

    Chi, M. T. H., Bassok, M., Lewis, M. W., Reimann, P. & Glaser, R. Self- explanations: How students study and use examples in learning to solve problems. Cogn. Sci.13, 145–182 (1989)

  41. [41]

    Learning from worked-out examples: A study on individual differences

    Renkl, A. Learning from worked-out examples: A study on individual differences. Cogn. Sci.21, 1–29 (1997)

  42. [42]

    & Wallace, R

    Aleven, V., Stahl, E., Schworm, S., Fischer, F. & Wallace, R. Help seeking and help design in interactive learning environments.Rev. Educ. Res.73, 277–320 (2003). 27

  43. [43]

    Reiser, B. J. Scaffolding complex learning: The mechanisms of structuring and problematizing student work.J. Learn. Sci.13, 273–304 (2004)

  44. [44]

    Learn.10, 17 (2025)

    Yin, J.et al.Effects of different AI-driven chatbot feedback on learning outcomes and brain activity.npj Sci. Learn.10, 17 (2025)

  45. [45]

    S., Collins, A

    Brown, J. S., Collins, A. & Duguid, P. Situated cognition and the culture of learning.Educ. Res.18, 32–42 (1989)

  46. [46]

    Collins, A., Brown, J. S. & Newman, S. E. Cognitive apprenticeship: Teaching the crafts of reading, writing, and mathematics. In Resnick, L. B. (ed.)Knowing, Learning, and Instruction: Essays in Honor of Robert Glaser, 453–494. Lawrence Erlbaum Associates, Hillsdale, NJ (1989)

  47. [47]

    W.et al.How large language models can reshape collective intelligence

    Burton, J. W.et al.How large language models can reshape collective intelligence. Nat. Hum. Behav.8, 1643–1655 (2024)

  48. [48]

    Lazer, D.et al.Computational social science.Science323, 721–723 (2009)

  49. [49]

    J.Bit by Bit: Social Research in the Digital Age(Princeton University Press, Princeton, NJ, 2017)

    Salganik, M. J.Bit by Bit: Social Research in the Digital Age(Princeton University Press, Princeton, NJ, 2017)

  50. [50]

    WildChat-4.8M

    Allen Institute for AI. WildChat-4.8M. https://huggingface.co/datasets/allenai/ WildChat-4.8M (2025)

  51. [51]

    Hern´ an, M. A. & Robins, J. M.Causal Inference: What If(Chapman & Hall/CRC, Boca Raton, 2020)

  52. [52]

    & Kubli, M

    Gilardi, F., Alizadeh, M. & Kubli, M. ChatGPT outperforms crowd workers for text-annotation tasks.Proc. Natl Acad. Sci. USA120, e2305016120 (2023)

  53. [53]

    Linguist.50, 237–291 (2024)

    Ziems, C.et al.Can large language models transform computational social science?Comput. Linguist.50, 237–291 (2024)

  54. [54]

    Matthews, B. W. Comparison of the predicted and observed secondary structure of T4 phage lysozyme.Biochim. Biophys. Acta405, 442–451 (1975)

  55. [55]

    Gwet, K. L. Computing inter-rater reliability and its variance in the presence of high agreement.Br. J. Math. Stat. Psychol.61, 29–48 (2008)

  56. [56]

    & Nelder, J

    McCullagh, P. & Nelder, J. A.Generalized Linear Models2nd edn, Vol. 37 of Monographs on Statistics and Applied Probability(Chapman & Hall, London, 1989)

  57. [57]

    Agresti, A.Categorical Data Analysis3rd edn (John Wiley & Sons, Hoboken, NJ, 2013). 28

  58. [58]

    Strict- English

    Benjamini, Y. & Hochberg, Y. Controlling the false discovery rate: a practical and powerful approach to multiple testing.J. R. Stat. Soc. B57, 289–300 (1995). 29 Supplementary information Supplementary information is organized to support reproducibility and interpretation. Appendix A records the corpus ledger, annotation ontology, corpus-scale annotation ...