Pith. sign in

REVIEW 3 major objections 1 minor 1 cited by

Physical laws are recovered from data by chaining simple, meaningful symbolic units rather than inventing one long expression at once.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-14 21:41 UTC pith:4BHBVVHU

load-bearing objection We only have the CoSR abstract; the supplied full text is an unrelated CHI UX paper, so every progressive-discovery claim is uncheckable. the 3 major comments →

arxiv 2603.13727 v2 pith:4BHBVVHU submitted 2026-03-14 cs.LG physics.data-an

Data-driven Progressive Discovery of Physical Laws

classification cs.LG physics.data-an
keywords symbolic regressionphysical law discoveryprogressive knowledge chainscaling lawsinterpretable machine learningscientific discoveryaerodynamic coefficients
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Conventional symbolic regression tries to invent a complete mathematical law in a single step and often produces long, unphysical formulas that fail to generalize. The paper argues that real physical discovery proceeds hierarchically, from simple relations to more complex ones, and therefore models discovery itself as a progressive chain of symbolic knowledge units that each carry clear physical meaning. By combining these units step by step along a logical path, the method recovers known laws (Kepler to Newton) and improves classical scaling relations in fluid and laser problems, while also extracting new scaling knowledge for aircraft aerodynamics. A sympathetic reader cares because the approach restores interpretability and generalization precisely where pure end-to-end search collapses into meaningless expressions.

Core claim

The authors claim that physical laws do not appear as monolithic expressions but as hierarchical chains; their Chain of Symbolic Regression therefore discovers them by successively assembling discrete knowledge units that already possess clear physical meanings, ultimately recovering the correct underlying law from data.

What carries the argument

Chain of Symbolic Regression (CoSR): a progressive sequence of symbolic knowledge units that are combined along a fixed logical path, each unit remaining physically interpretable.

Load-bearing premise

The claim rests on the premise that every physical law of interest can be broken into a fixed hierarchy of simple, physically meaningful symbolic units whose progressive combination is both necessary and sufficient to recover the true law.

What would settle it

Apply CoSR and a standard one-step symbolic regressor to a system whose progressive discovery path is already known (for example the Kepler-to-Newton sequence) and check whether CoSR alone recovers the intermediate units and the final law while the one-step method produces only lengthy unphysical expressions.

Watch this falsifier — get emailed when new claim-graph text bears on it.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 1 minor

Summary. The abstract proposes Chain of Symbolic Regression (CoSR), a framework that discovers physical laws by progressively combining discrete symbolic knowledge units with clear physical meanings, rather than one-shot end-to-end symbolic regression. It claims to recapitulate the historical path from Kepler’s third law to Newton’s law of universal gravitation, to improve classical scaling relations for turbulent Rayleigh–Bénard convection, pipe flow, and laser–metal interaction, and to discover new aerodynamic-coefficient scalings for aircraft. The supplied full manuscript body, however, is an entirely unrelated CHI ’26 multi-session HCI study of UX evaluators collaborating with novice versus experienced conversational AI assistants (arXiv:2603.13717). Consequently no methods, algorithms, intermediate expressions, datasets, error metrics, baselines, or ablations for CoSR are present.

Significance. If the progressive-chain idea were correctly implemented and validated, CoSR would address a recognized failure mode of conventional symbolic regression (lengthy, unphysical expressions with poor generalization) and would constitute a meaningful methodological contribution to data-driven scientific discovery. The historical Kepler-to-Newton reconstruction and the claimed improvements to classical scaling laws would be especially valuable if independently recovered from data. None of these contributions can be assessed from the manuscript as supplied.

major comments (3)
  1. The full text provided under the CoSR title and abstract is the complete, unrelated CHI ’26 paper “It Became My Buddy, But I’m Not Afraid to Disagree” (arXiv:2603.13717). No section, equation, algorithm, figure, or table belonging to CoSR appears. Every empirical claim in the abstract—recovery of the gravitational law, improved Rayleigh–Bénard / pipe-flow / laser-metal scalings, and new aircraft aerodynamic coefficients—is therefore uninspectable.
  2. Because the method body is absent, it is impossible to determine whether the “knowledge units” and the “specific logic” that combine them are recovered from data or are seeded by the same classical knowledge the method claims to rediscover. This circularity risk, already flagged by the abstract’s motivating premise, cannot be evaluated.
  3. No datasets, intermediate symbolic expressions, quantitative error metrics, baselines (e.g., standard genetic-programming or sparse-regression SR), or ablation studies are supplied. The central claim that progressive chaining overcomes the unphysical expressions of one-step SR therefore rests solely on an unsupported abstract.
minor comments (1)
  1. The arXiv identifier printed in the body (2603.13717) does not match the identifier under review (2603.13727), confirming a manuscript-swap error.

Circularity Check

0 steps flagged

No circularity can be assessed: supplied full text is an unrelated CHI UX/CA study; CoSR derivation chain is absent.

full rationale

The claimed paper (arXiv 2603.13727, CoSR / progressive symbolic discovery of physical laws) is represented only by its abstract. The body that follows under FULL MANUSCRIPT TEXT is an entirely different manuscript (CHI ’26 multi-session UX-evaluator study with conversational AI assistants, arXiv 2603.13717). Consequently there are no CoSR equations, intermediate symbolic units, fitting procedures, Kepler-to-Newton reconstruction steps, scaling-law improvements, or aircraft-coefficient results that can be inspected. Per the analyzer rules, circularity may be flagged only when a specific reduction can be quoted and exhibited (self-definitional identity, fitted parameter renamed as prediction, load-bearing self-citation, etc.). With no derivation present, no such reduction exists. The abstract’s narrative claim that CoSR “fully recapitulates” a known historical path is therefore uncheckable rather than circular. Score is 0; steps list is empty.

Axiom & Free-Parameter Ledger

0 free parameters · 2 axioms · 1 invented entities

With only the abstract, the ledger is necessarily incomplete. The central claim rests on the domain assumption that physical discovery is hierarchical and progressive, plus the ad-hoc construction of “knowledge units” and a “chain” logic whose precise definition is not given. No free parameters or invented physical entities (particles, forces, etc.) are named in the abstract.

axioms (2)
  • domain assumption Physical laws do not exist in a single form but follow a hierarchical and progressive pattern from simplicity to complexity.
    Stated as the fundamental motivation; if false for the target systems, the chain construction loses its justification.
  • ad hoc to paper A chain of symbolic knowledge units with clear physical meanings can be progressively combined along a specific logic to recover the true underlying law.
    Core modeling choice of CoSR; the abstract does not derive why this particular chaining is complete or unique.
invented entities (1)
  • Chain of Symbolic Regression (CoSR) / knowledge chain / knowledge units no independent evidence
    purpose: To structure symbolic discovery as progressive combination of simple units rather than one-shot search.
    These are the paper’s central constructs; no independent evidence outside the claimed recoveries is supplied in the abstract.

pith-pipeline@v1.1.0-grok45 · 40653 in / 2245 out tokens · 27222 ms · 2026-07-14T21:41:06.150457+00:00 · methodology

0 comments
read the original abstract

Symbolic regression is a powerful tool for knowledge discovery, enabling the extraction of interpretable mathematical expressions directly from data. However, conventional symbolic discovery typically follows an end-to-end, "one-step" process, which often generates lengthy and physically meaningless expressions when dealing with real physical systems, leading to poor model generalization. This limitation fundamentally stems from its deviation from the basic path of scientific discovery: physical laws do not exist in a single form but follow a hierarchical and progressive pattern from simplicity to complexity. Motivated by this principle, we propose Chain of Symbolic Regression (CoSR), a novel framework that models the discovery of physical laws as a chain of symbolic knowledge. This knowledge chain is formed by progressively combining multiple knowledge units with clear physical meanings along a specific logic, ultimately enabling the precise discovery of the underlying physical laws from data. CoSR fully recapitulates the progressive discovery path from Kepler's third law to the law of universal gravitation in classical mechanics, and is applied to three types of problems: turbulent Rayleigh-Benard convection, viscous flows in a circular pipe, and laser-metal interaction, demonstrating its ability to improve classical scaling theories. Finally, CoSR showcases its capability to discover new knowledge in the complex engineering problem of aerodynamic coefficients scaling for different aircraft.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Skin friction prediction for attached flows based on two-dimensional inviscid solutions

    physics.flu-dyn 2026-07 conditional novelty 6.0

    A progressive symbolic regression framework discovers an interpretable analytical formula chain for skin friction prediction across subsonic to hypersonic attached flows using inviscid surface flow features.

Reference graph

Works this paper leans on

162 extracted references · 41 canonical work pages · cited by 1 Pith paper · 3 internal anchors

  1. [1]

    Ajenaghughrure, Sonia C

    Ighoyota Ben. Ajenaghughrure, Sonia C. Sousa, Ilkka Johannes Kosunen, and David Lamas. 2019. Predictive model to assess user trust: a psycho-physiological approach. InProceedings of the 10th Indian Conference on Human-Computer Interaction (IndiaHCI ’19). Association for Computing Machinery, New York, NY, USA, 1–10. https://doi.org/10.1145/3364183.3364195

  2. [2]

    AltexSoft. 2019. UX Researcher: Methods, Skills and Process. https://www. altexsoft.com/blog/ux-researcher-methods-skills-process/

  3. [3]

    Theo Araujo and Nadine Bol. 2024. From speaking like a person to being personal: The effects of personalized, regular interactions with conversational agents.Computers in Human Behavior: Artificial Humans2, 1 (Jan. 2024), 100030. https://doi.org/10.1016/j.chbah.2023.100030

  4. [4]

    Andrea Batch, Yipeng Ji, Mingming Fan, Jian Zhao, and Niklas Elmqvist. 2023. uxSense: Supporting User Experience Analysis with Visualization and Com- puter Vision.IEEE Transactions on Visualization and Computer Graphics(2023), 1–15. https://doi.org/10.1109/TVCG.2023.3241581 Conference Name: IEEE Transactions on Visualization and Computer Graphics

  5. [5]

    Patricia Benner. 1982. From Novice To Expert.AJN The American Journal of Nursing82, 3 (1982), 402–407. https://journals.lww.com/ajnonline/citation/ 1982/82030/from_novice_to_expert.4.a

  6. [6]

    Erik Brynjolfsson, Danielle Li, and Lindsey R. Raymond. 2023. Generative AI at Work(Working Paper Series). National Bureau of Economic Research. https://doi.org/10.3386/w31161

  7. [7]

    Félix Buendía-García and Javier Piris-Ruano. 2025. Using Generative AI to Support UX Design Students in Web Development Courses.Applied Sciences15, 13 (June 2025), 7389. https://doi.org/10.3390/app15137389 Publisher: Multidis- ciplinary Digital Publishing Institute

  8. [8]

    Gajos, and Elena L

    Zana Buçinca, Phoebe Lin, Krzysztof Z. Gajos, and Elena L. Glassman. 2020. Proxy tasks and subjective measures can be misleading in evaluating explainable AI systems. InProceedings of the 25th International Conference on Intelligent User Interfaces (IUI ’20). Association for Computing Machinery, New York, NY, USA, 454–464. https://doi.org/10.1145/3377325.3377498

  9. [9]

    Zana Buçinca, Maja Barbara Malaya, and Krzysztof Z. Gajos. 2021. To Trust or to Think: Cognitive Forcing Functions Can Reduce Overreliance on AI in AI-assisted Decision-making.Proc. ACM Hum.-Comput. Interact.5, CSCW1 (April 2021), 188:1–188:21. https://doi.org/10.1145/3449287

  10. [10]

    2006.Constructing Grounded Theory: A Practical Guide Through Qualitative Analysis

    Kathy Charmaz. 2006.Constructing Grounded Theory: A Practical Guide Through Qualitative Analysis. SAGE. Google-Books-ID: v1qP1KbXz1AC. Multi-Session Study of UX Evaluators Using Conversational AI Agents CHI ’26, April 13–17, 2026, Barcelona, Spain

  11. [11]

    Chase, Doris B

    Catherine C. Chase, Doris B. Chin, Marily A. Oppezzo, and Daniel L. Schwartz

  12. [12]

    2009), 334–352

    Teachable Agents and the Protégé Effect: Increasing the Effort Towards Learning.Journal of Science Education and Technology18, 4 (Aug. 2009), 334–352. https://doi.org/10.1007/s10956-009-9180-4

  13. [13]

    Hsi-Jen Chen, Yan-Ting Chen, and Chia-Han Yang. 2022. Behaviors of Novice and Expert Designers in the Design Process: From Discovery to Design.Inter- national Journal of Design16, 3 (2022), 59–76. https://doi.org/10.57698/v16i3.04

  14. [14]

    Chilana, Jacob O

    Parmit K. Chilana, Jacob O. Wobbrock, and Andrew J. Ko. 2010. Understanding Usability Practices in Complex Domains. InProceedings of the 28th International Conference on Human Factors in Computing Systems - CHI ’10. ACM Press, Atlanta, Georgia, USA, 2337–2346. https://doi.org/10.1145/1753326.1753678

  15. [15]

    Dorothee Clasen and Marc Hassenzahl. 2024. Fostering people’s autonomy by foregrounding and questioning daily choices. InAdjunct Proceedings of the 2024 Nordic Conference on Human-Computer Interaction (NordiCHI ’24 Adjunct). Association for Computing Machinery, New York, NY, USA, 1–5. https://doi. org/10.1145/3677045.3685442

  16. [16]

    Andy Cockburn, Carl Gutwin, Joey Scarr, and Sylvain Malacria. 2014. Supporting Novice to Expert Transitions in User Interfaces.ACM Comput. Surv.47, 2 (Nov. 2014), 31:1–31:36. https://doi.org/10.1145/2659796

  17. [17]

    Tyler Corwin, Mehmet Kosa, Mahsa Nasri, Christoffer Holmgård, and Casper Harteveld. 2023. The Teaching Efficacy of the Protégé Effect in Gamified Edu- cation. In2023 IEEE Conference on Games (CoG). 1–8. https://doi.org/10.1109/ CoG57401.2023.10333166 ISSN: 2325-4289

  18. [18]

    Emmelyn A. J. Croes and Marjolijn L. Antheunis. 2021. Can we be friends with Mitsuku? A longitudinal study on the process of relationship formation between humans and a social chatbot - Emmelyn A. J. Croes, Marjolijn L. Antheunis, 2021.Journal of Social and Personal Relationships38, 1 (2021), 279–300. https: //doi.org/10.1177/0265407520959463

  19. [19]

    Fleur Deken, Maaike Kleinsmann, Marco Aurisicchio, Kristina Lauche, and Rob Bracewell. 2011. Tapping into past design experiences: knowledge sharing and creation during novice–expert design consultations.Research in Engineering Design23, 3 (Oct. 2011), 203–218. https://doi.org/10.1007/s00163-011-0123-8 Number: 3 Publisher: Springer

  20. [20]

    2022.Intelligent Adaptive Flight Training System- A Human Performance in the Loop for Real-Time Decision Making

    Jean-François Delisle. 2022.Intelligent Adaptive Flight Training System- A Human Performance in the Loop for Real-Time Decision Making. phd. Polytechnique Montréal. https://publications.polymtl.ca/10551/

  21. [21]

    Dietvorst, Joseph P

    Berkeley J. Dietvorst, Joseph P. Simmons, and Cade Massey. 2015. Algorithm aversion: People erroneously avoid algorithms after seeing them err.Journal of Experimental Psychology: General144, 1 (2015), 114–126. https://doi.org/10. 1037/xge0000033 Place: US Publisher: American Psychological Association

  22. [22]

    Shiying Ding, Xinyi Chen, Yan Fang, Wenrui Liu, Yiwu Qiu, and Chunlei Chai

  23. [23]

    In2023 16th Interna- tional Symposium on Computational Intelligence and Design (ISCID)

    DesignGPT: Multi-Agent Collaboration in Design. In2023 16th Interna- tional Symposium on Computational Intelligence and Design (ISCID). 204–208. https://doi.org/10.1109/ISCID59865.2023.00056 ISSN: 2473-3547

  24. [24]

    Peitong Duan, Jeremy Warner, Yang Li, and Bjoern Hartmann. 2024. Generating Automatic Feedback on UI Mockups with Large Language Models. InProceedings of the 2024 CHI Conference on Human Factors in Computing Systems (CHI ’24). Association for Computing Machinery, New York, NY, USA, 1–20. https://doi. org/10.1145/3613904.3642782

  25. [25]

    Mingming Fan, Yue Li, and Khai N. Truong. 2020. Automatic Detection of Usability Problem Encounters in Think-Aloud Sessions.ACM Transactions on Interactive Intelligent Systems10, 2 (June 2020), 1–24. https://doi.org/10.1145/ 3385732

  26. [26]

    Mingming Fan, Serina Shi, and Khai N Truong. 2020. Practices and Challenges of Using Think-Aloud Protocols in Industry: An International Survey.Journal of Usability Studies15, 2 (2020), 85–102

  27. [27]

    Mingming Fan, Ke Wu, Jian Zhao, Yue Li, Winter Wei, and Khai N. Truong

  28. [28]

    2020), 343–352

    VisTA: Integrating Machine Intelligence with Visualization to Support the Investigation of Think-Aloud Sessions.IEEE Transactions on Visualization and Computer Graphics26, 1 (Jan. 2020), 343–352. https://doi.org/10.1109/TVCG. 2019.2934797

  29. [29]

    Vera Liao, and Jian Zhao

    Mingming Fan, Xianyou Yang, TszTung Yu, Q. Vera Liao, and Jian Zhao. 2022. Human-AI Collaboration for UX Evaluation: Effects of Explanation and Synchro- nization. 6 (2022), 96:1–96:32. Issue CSCW1. https://doi.org/10.1145/3512943

  30. [30]

    James Fisher. 1991. Defining the novice user.Behaviour & Information Technology 10, 5 (1991), 437–441. https://doi.org/10.1080/01449299108924301 Place: United Kingdom Publisher: Taylor & Francis

  31. [31]

    Asbjørn Følstad, Effie Lai-Chong Law, and Kasper Hornbæk. 2010. Analysis in Usability Evaluations: An Exploratory Study. InProceedings of the 6th Nordic Conference on Human-Computer Interaction: Extending Boundaries (NordiCHI ’10). Association for Computing Machinery, New York, NY, USA, 647–650. https: //doi.org/10.1145/1868914.1868995

  32. [32]

    Asbjørn Følstad, Effie Lai-Chong Law, and Kasper Hornbæk. 2012. Analysis in Practical Usability Evaluation: A Survey Study. InProceedings of the 30th SIGCHI Conference on Human Factors in Computing Systems - CHI ’12. ACM Press, Austin, Texas, 2127–2136. https://doi.org/10.1145/2207676.2208365

  33. [33]

    Eureka Foong, Darren Gergle, and Elizabeth M. Gerber. 2017. Novice and Expert Sensemaking of Crowdsourced Design Feedback.Proc. ACM Hum.-Comput. Interact.1, CSCW (Dec. 2017), 45:1–45:18. https://doi.org/10.1145/3134680

  34. [34]

    2010.Longitudinal Research in Human-Computer Interaction

    Jens Gerken. 2010.Longitudinal Research in Human-Computer Interaction. Ph. D. Dissertation. Universität Konstanz, Konstanz, Germany. https://kops.uni- konstanz.de/entities/publication/ef10c26c-3f4c-41a9-8045-da53586fc360

  35. [35]

    Kaklamanos

    Kostis Giannakopoulos, Argyro Kavadella, Anas Aaqel Salim, Vassilis Stam- atopoulos, and Eleftherios G. Kaklamanos. 2023. Evaluation of the Perfor- mance of Generative AI Large Language Models ChatGPT, Google Bard, and Microsoft Bing Chat in Supporting Evidence-Based Dentistry: Comparative Mixed Methods Study.Journal of Medical Internet Research25, 1 (Dec...

  36. [36]

    Julián Grigera, Alejandra Garrido, José Matías Rivero, and Gustavo Rossi. 2017. Automatic detection of usability smells in web applications.International Journal of Human-Computer Studies97 (2017), 129–148

  37. [37]

    Siddharth Gulati, Joe McDonagh, Sonia Sousa, and David Lamas. 2024. Trust models and theories in human–computer interaction: A systematic literature review.Computers in Human Behavior Reports16 (Dec. 2024), 100495. https: //doi.org/10.1016/j.chbr.2024.100495

  38. [38]

    An Error Occurred!

    Kasper Hald, Katharina Weitz, Elisabeth André, and Matthias Rehm. 2021. “An Error Occurred!” - Trust Repair With Virtual Robot Using Levels of Mistake Explanation. InProceedings of the 9th International Conference on Human-Agent Interaction (HAI ’21). Association for Computing Machinery, New York, NY, USA, 218–226. https://doi.org/10.1145/3472307.3484170

  39. [39]

    Patrick Harms. 2019. Automated usability evaluation of virtual reality applica- tions.ACM Transactions on Computer-Human Interaction (TOCHI)26, 3 (2019), 1–36

  40. [40]

    Morten Hertzum and Niels Ebbe Jacobsen. 2001. The Evaluator Effect: A Chilling Fact About Usability Evaluation Methods.International Journal of Human-Computer Interaction15, 1 (2001), 183–204. https://doi.org/10.1207/ S15327590IJHC1501_14

  41. [41]

    Morten Hertzum, Niels Ebbe Jacobsen, and Bonnie E. John. 1998. The Evaluator Effect in Usability Tests. InCHI 98 Conference Summary on Human Factors in Computing Systems (CHI ’98). Association for Computing Machinery, New York, NY, USA, 255–256. https://doi.org/10.1145/286498.286737

  42. [42]

    Annabell Ho, Jeff Hancock, and Adam S Miner. 2018. Psychological, Relational, and Emotional Effects of Self-Disclosure After Conversations With a Chatbot. Journal of Communication68, 4 (Aug. 2018), 712–733. https://doi.org/10.1093/ joc/jqy026

  43. [43]

    Hoffman, Matthew Johnson, Jeffrey M

    Robert R. Hoffman, Matthew Johnson, Jeffrey M. Bradshaw, and Al Underbrink

  44. [44]

    2013), 84–88

    Trust in Automation.IEEE Intelligent Systems28, 1 (Jan. 2013), 84–88. https://doi.org/10.1109/MIS.2013.24

  45. [45]

    Jane Hogan, Gary Grant, Fiona Kelly, and Jennie O’Hare. 2020. Factors influ- encing acceptance of robotics in hospital pharmacy: a longitudinal study using the Extended Technology Acceptance Model.International Journal of Pharmacy Practice28, 5 (Oct. 2020), 483–490. https://doi.org/10.1111/ijpp.12637

  46. [46]

    Ironhack. 2023. The Evolving Role of a UX/UI Designer: Trends and Skills for Success. https://www.ironhack.com/gb/blog/the-evolving-role-of-a-ux-ui- designer-trends-and-skills-for-success

  47. [47]

    JongWook Jeong, NeungHoe Kim, and Hoh Peter In. 2020. Detecting usability problems in mobile applications on the basis of dissimilarity in user behavior. International Journal of Human-Computer Studies139 (2020), 102364

  48. [48]

    Kahr, Gerrit Rooks, Martijn C

    Patricia K. Kahr, Gerrit Rooks, Martijn C. Willemsen, and Chris C. P. Sni- jders. 2024. Understanding Trust and Reliance Development in AI Advice: Assessing Model Accuracy, Model Explanations, and Experiences from Pre- vious Interactions.ACM Trans. Interact. Intell. Syst.(Aug. 2024). https: //doi.org/10.1145/3686164 Just Accepted

  49. [49]

    It Requires Interest, Time, Patience and Struggle

    Mahmut Kalman. 2019. “It Requires Interest, Time, Patience and Struggle”: Novice Researchers’ Perspectives on and Experiences of the Qualitative Research Journey.Qualitative Research in Education8, 3 (Oct. 2019), 341–377. https: //doi.org/10.17583/qre.2019.4483 Number: 3

  50. [50]

    Takayuki Kanda, Rumi Sato, Naoki Saiwaki, and Hiroshi Ishiguro. 2007. A Two-Month Field Trial in an Elementary School for Long-Term Human–Robot Interaction.IEEE Transactions on Robotics23, 5 (Oct. 2007), 962–971. https: //doi.org/10.1109/TRO.2007.904904 Conference Name: IEEE Transactions on Robotics

  51. [51]

    Evangelos Karapanos, Jhilmil Jain, and Marc Hassenzahl. 2012. Theories, meth- ods and case studies of longitudinal HCI research. InCHI ’12 Extended Abstracts on Human Factors in Computing Systems (CHI EA ’12). Association for Com- puting Machinery, New York, NY, USA, 2727–2730. https://doi.org/10.1145/ 2212776.2212706

  52. [52]

    Abdolvahab Khademi. 2023. Can ChatGPT and Bard Generate Aligned Assess- ment Items? A Reliability Analysis against Human Performance.Journal of Applied Learning & Teaching6, 1 (May 2023). https://doi.org/10.37074/jalt.2023. 6.1.28 arXiv:2304.05372 [cs]. CHI ’26, April 13–17, 2026, Barcelona, Spain

  53. [53]

    Ahmad Khawaji, Jianlong Zhou, Fang Chen, and Nadine Marcus. 2015. Using Galvanic Skin Response (GSR) to Measure Trust and Cognitive Load in the Text- Chat Environment. InProceedings of the 33rd Annual ACM Conference Extended Abstracts on Human Factors in Computing Systems (CHI EA ’15). Association for Computing Machinery, New York, NY, USA, 1989–1994. htt...

  54. [54]

    Jay Kim. 2020. Rubber Ducking- What It Is and Why It Works. https://medium.com/@jkarma0920/rubber-ducking-what-it-is-and-why-it- works-5d026fd9ae58

  55. [55]

    Sunyoung Kim and Abhishek Choudhury. 2021. Exploring older adults’ per- ception and use of smart speaker-based voice assistants: A longitudinal study. Computers in Human Behavior124 (Nov. 2021), 106914. https://doi.org/10.1016/ j.chb.2021.106914

  56. [56]

    Skov, Peter Axel Nielsen, Jesper Kjeldskov, Jens Gerken, and Harald Reiterer

    Maria Kjærup, Mikael B. Skov, Peter Axel Nielsen, Jesper Kjeldskov, Jens Gerken, and Harald Reiterer. 2021. Longitudinal Studies in HCI Research: A Review of CHI Publications From 1982–2019. InAdvances in Longitudinal HCI Research, Evangelos Karapanos, Jens Gerken, Jesper Kjeldskov, and Mikael B. Skov (Eds.). Springer International Publishing, Cham, 11–39...

  57. [57]

    Eva Krapp, Robin Neuhaus, Marc Hassenzahl, and Matthias Laschke. 2024. In a Quasi-Social Relationship With ChatGPT. An Autoethnography on Engaging With Prompt-Engineered LLM Personas. InProceedings of the 13th Nordic Confer- ence on Human-Computer Interaction (NordiCHI ’24). Association for Computing Machinery, New York, NY, USA, 1–16. https://doi.org/10....

  58. [58]

    Emily Kuang. 2025. Evaluating Usability Challenges in VR Games for Older Adults: A Comparison With and Without AI Assistance. InAdjunct Proceedings of the 38th Annual ACM Symposium on User Interface Software and Technology (UIST Adjunct ’25). Association for Computing Machinery, New York, NY, USA, Article 92, 3 pages. https://doi.org/10.1145/3746058.3758343

  59. [59]

    Emily Kuang, Ehsan Jahangirzadeh Soure, Mingming Fan, Jian Zhao, and Kristen Shinohara. 2023. Collaboration with Conversational AI Assistants for UX Evaluation: Questions and How to Ask them (Voice vs. Text). InPro- ceedings of the 2023 CHI Conference on Human Factors in Computing Systems (CHI ’23). Association for Computing Machinery, New York, NY, USA, ...

  60. [60]

    Merging Results Is No Easy Task

    Emily Kuang, Xiaofu Jin, and Mingming Fan. 2022. "Merging Results Is No Easy Task": An International Survey Study of Collaborative Data Analysis Practices Among UX Practitioners. InProceedings of the 2022 CHI Conference on Human Factors in Computing Systems.Association for Computing Machinery, New York, NY, USA. https://doi.org/10.1145/3491102.3517647

  61. [61]

    Emily Kuang, Minghao Li, Mingming Fan, and Kristen Shinohara. 2024. Enhanc- ing UX Evaluation Through Collaboration with Conversational AI Assistants: Effects of Proactive Dialogue and Timing. InProceedings of the CHI Conference on Human Factors in Computing Systems (CHI ’24). Association for Computing Machinery, New York, NY, USA, 1–16. https://doi.org/1...

  62. [62]

    Vera Liao, Alison Smith-Renner, and Chenhao Tan

    Vivian Lai, Chacha Chen, Q. Vera Liao, Alison Smith-Renner, and Chenhao Tan

  63. [63]

    https://doi.org/10.48550/arXiv.2112.11471 arXiv:2112.11471 [cs]

    Towards a Science of Human-AI Decision Making: A Survey of Empirical Studies. https://doi.org/10.48550/arXiv.2112.11471 arXiv:2112.11471 [cs]

  64. [64]

    Richard Landis and Gary G

    J. Richard Landis and Gary G. Koch. 1977. The Measurement of Observer Agreement for Categorical Data.Biometrics33, 1 (March 1977), 159. https: //doi.org/10.2307/2529310

  65. [65]

    Claire Lauer, Danielle Storey, and Romit Soley. 2024. Vector Personas: How UX Researchers Can Use AI to Bring New Dimension to Traditional Persona Development. InProceedings of the 42nd ACM International Conference on Design of Communication (SIGDOC ’24). Association for Computing Machinery, New York, NY, USA, 151–157. https://doi.org/10.1145/3641237.3691663

  66. [66]

    Nadine Lessio and Alexis Morris. 2020. Toward Design Archetypes for Conver- sational Agent Personality. In2020 IEEE International Conference on Systems, Man, and Cybernetics (SMC). 3221–3228. https://doi.org/10.1109/SMC42975. 2020.9283254 ISSN: 2577-1655

  67. [67]

    Tobias Lieberei, Virginia Deborah Elaine Welter, Leroy Großmann, and Moritz Krell. 2023. Findings from the expert-novice paradigm on differential response behavior among multiple-choice items of a pedagogical content knowledge test – implications for test development. 14 (2023). https://doi.org/10.3389/fpsyg. 2023.1240120

  68. [68]

    Yiren Liu, Pranav Sharma, Mehul Oswal, Haijun Xia, and Yun Huang. 2025. PersonaFlow: Designing LLM-Simulated Expert Perspectives for Enhanced Research Ideation. InProceedings of the 2025 ACM Designing Interactive Systems Conference (DIS ’25). Association for Computing Machinery, New York, NY, USA, 506–534. https://doi.org/10.1145/3715336.3735789

  69. [69]

    Zhe Liu, Chunyang Chen, Junjie Wang, Mengzhuo Chen, Boyu Wu, Xing Che, Dandan Wang, and Qing Wang. 2023. Chatting with GPT-3 for Zero-Shot Human-Like Mobile Automated GUI Testing. https://doi.org/10.48550/arXiv. 2305.09434 arXiv:2305.09434 [cs]

  70. [70]

    Zhe Liu, Chunyang Chen, Junjie Wang, Mengzhuo Chen, Boyu Wu, Xing Che, Dandan Wang, and Qing Wang. 2024. Make LLM a Testing Expert: Bringing Human-like Interaction to Mobile GUI Testing via Functionality-aware Deci- sions. InProceedings of the IEEE/ACM 46th International Conference on Software Engineering (ICSE ’24). Association for Computing Machinery, N...

  71. [71]

    Tao Long, Katy Ilonka Gero, and Lydia B Chilton. 2024. Not Just Novelty: A Longitudinal Study on Utility and Customization of an AI Workflow. In Proceedings of the 2024 ACM Designing Interactive Systems Conference (DIS ’24). Association for Computing Machinery, New York, NY, USA, 782–803. https: //doi.org/10.1145/3643834.3661587

  72. [72]

    Yuwen Lu, Yuewen Yang, Qinyi Zhao, Chengzhi Zhang, and Toby Jia-Jun Li

  73. [73]

    AI Assistance for UX: A Literature Review Through Human-Centered AI

    AI Assistance for UX: A Literature Review Through Human-Centered AI. https://doi.org/10.48550/arXiv.2402.06089 arXiv:2402.06089 [cs]

  74. [74]

    Like Having a Really Bad PA

    Ewa Luger and Abigail Sellen. 2016. "Like Having a Really Bad PA": The Gulf between User Expectation and Experience of Conversational Agents. In Proceedings of the 2016 CHI Conference on Human Factors in Computing Systems (New York, NY, USA)(CHI ’16). Association for Computing Machinery, 5286–

  75. [75]

    https://doi.org/10.1145/2858036.2858288

  76. [76]

    Kim Man Lui and Keith C. C. Chan. 2006. Pair programming productivity: Novice–novice vs. expert–expert.International Journal of Human-Computer Studies64, 9 (Sept. 2006), 915–925. https://doi.org/10.1016/j.ijhcs.2006.04.010

  77. [77]

    Shuai Ma, Qiaoyi Chen, Xinru Wang, Chengbo Zheng, Zhenhui Peng, Ming Yin, and Xiaojuan Ma. 2025. Towards Human-AI Deliberation: Design and Evalua- tion of LLM-Empowered Deliberative AI for AI-Assisted Decision-Making. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems (CHI ’25). Association for Computing Machinery, New York, N...

  78. [78]

    Ellie Martin. 2016. Why 5 is the magic number for UX usability testing | Inside Design Blog. https://www.invisionapp.com/inside-design/ux-usability- research-testing/

  79. [79]

    McGrath, Oliver Lack, James Tisch, and Andreas Duenser

    Melanie J. McGrath, Oliver Lack, James Tisch, and Andreas Duenser. 2025. Measuring trust in artificial intelligence: validation of an established scale and its short form.Frontiers in Artificial Intelligence8 (May 2025). https: //doi.org/10.3389/frai.2025.1582880 Publisher: Frontiers

  80. [80]

    Valerie Mendoza and David G. Novick. 2005. Usability over time. InPro- ceedings of the 23rd annual international conference on Design of communi- cation: documenting & designing for pervasive information (SIGDOC ’05). As- sociation for Computing Machinery, New York, NY, USA, 151–158. https: //doi.org/10.1145/1085313.1085348

Showing first 80 references.