Pith. sign in

REVIEW 5 major objections 5 minor 100 references

Interaction as Intelligence: Deep Research With Human-AI Partnership

T0 review · 5 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read Deep Cognition claims interaction itself is intelligence, and a pause-and-steer research tool outperforms black-box systems by up to 50 points.

desk verdict A real interactive deep-research system and a plausible design thesis, but the headline empirical claims collapse under a confounded benchmark and a demand-laden user protocol. read the letter →

arxiv 2507.15759 v1 pith:CZNLTBCL submitted 2025-07-21 cs.CL

classification cs.CL
keywords human-AIinteractioncognitiveoversightdeepresearchasintelligencehuman-in-the-loopagentsBrowseComp-ZHuserevaluationtransparentreasoning
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that for long, open-ended research tasks, interaction is not a wrapper around AI capability but a core part of the intelligence doing the work. It identifies the standard 'input-wait-output' pattern of today's deep research systems—send a query, wait through a black box, read a finished report—as the cause of compounding errors, rigid research boundaries, and lost chances to inject expertise. The authors build Deep Cognition, a multi-agent system whose interface exposes the AI's reasoning live and lets a user pause, redirect, add domain knowledge, and edit details at any point, while a preference agent adapts to the user's behavior. Their evaluation claims large gains: six interaction metrics improve by 8.8 to 29.2 percentage points over the strongest baseline, and on 22 hard BrowseComp-ZH questions the system with graduate-level oversight reaches 72.73% accuracy versus 40.91% for two leading commercial systems and 22.73% for a third. The underlying thesis is that human cognitive oversight plus AI execution outperforms either alone.

What carries the argument

The load-bearing mechanism is the Deep Cognition multi-agent architecture plus its interaction surface. A research agent plans searches, clarifies ambiguity through targeted questions, drafts and self-critiques the report; a browsing agent fetches and filters web pages; and a preference agent treats user actions and corrections as in-context reward signals to adapt search, source, and report preferences within a session. The interaction surface supplies the four features that make 'cognitive oversight' concrete: transparent display of reasoning and search strategy, a pause-and-interrupt control that works while the research is running, fine-grained editing and clarification at the level of individual claims or sources, and adaptive behavior driven by prior user choices. The system's claim that these features matter is carried by the contrast between the combined condition and the two ablations.

What would settle it

A preregistered evaluation with new, blinded participants who have not read the study protocol would settle the benchmark claim: run the same 22 BrowseComp-ZH questions under three conditions (expert oversight with interaction, expert knowledge without interaction, and non-expert interaction), with independent graders who do not know the condition, and check whether the combined condition still exceeds both ablations and the commercial baselines. If the gap shrinks to noise, the cognitive-oversight thesis loses its empirical support.

Watch

Extended reading notes

Core claim

The paper's central claim is that meaningful interaction during an AI's long thinking process is itself a component of intelligence, not merely a user-interface feature. On the system's own terms, Deep Cognition implements 'cognitive oversight': the human watches the model's search strategies, reasoning, and report drafts as they form, interrupts at critical moments, and steers with domain knowledge, while the system observes those interventions and adapts. The decisive evidence is the ablation on BrowseComp-ZH: with both cognition and interaction the system scores 72.73% accuracy, but with only background knowledge (no interaction) it scores 45.45%, and with interaction by non-expert participants it scores 40.91%—the same as the strongest commercial baselines. The paper reads this as showing that neither expert knowledge alone nor interactive control alone is sufficient; the combination is what creates the jump. In user evaluation, the same design claims 63% average report-quality improvement over the system with interaction disabled, and top scores on transparency, interruptibility, fine-grained interaction, real-time intervention, ease of collaboration, and results-worth-effort.

Load-bearing premise

The reported results assume that the 13 graduate-student participants and the 4 benchmark participants gave authentic behavior and ratings, rather than following the study protocol's instruction to aim for 4-5 points across all dimensions before stopping generation; if that demand characteristic drove the scores, the claimed interaction benefits would not generalize.

Editorial extensions

If this is right

  • Deep research products should expose their reasoning and allow mid-flight intervention rather than returning a finished report after black-box processing.
  • System capability and human oversight are not alternatives: the ablation result implies that comparable models can differ by more than 30 points depending on whether an expert can steer them live.
  • Interaction quality—transparency, interruptibility, fine-grained control—should be reported as a first-class evaluation dimension for research assistants, alongside output quality.
  • Users delegate mechanical phases such as browsing and summarizing and re-engage at decision points, so systems should support dynamic switches between hands-on and hands-off collaboration.
  • If the accuracy result replicates, benchmark comparisons of deep research systems should state the human interaction condition, since a model-only comparison would understate what the system can do with oversight.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A testable extension would vary participant expertise continuously—novice, graduate, domain expert—to map where cognitive oversight adds the most value and where it begins to slow or bias the research.
  • The small 22-question sample and four participants per condition mean the headline 31.8 to 50.0 point margins should be read as an existence proof until replicated on a larger, independently selected question set.
  • The thesis predicts that as long-task models improve, the bottleneck will shift from correcting model errors to injecting tacit knowledge and refining goals, which is measurable by comparing expert and novice oversight on the same questions.
  • Treating user actions as reward signals could make systems adapt to compliance rather than to true preferences; a useful check is whether systems built this way generalize across users with different interaction styles.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 5 minor

Summary. The paper argues that human-AI interaction is itself a dimension of intelligence, and introduces Deep Cognition, a multi-agent deep-research system that implements transparent, interruptible, fine-grained interaction and a preference-adaptation mechanism. The authors report a 13-participant user study in which Deep Cognition outperforms Gemini, OpenAI Deep Research, and Grok 3 on six interaction-oriented metrics, and a BrowseComp-ZH benchmark study in which Deep Cognition with graduate-level human interaction reaches 72.73% accuracy versus 40.91% for Gemini and OpenAI and 22.73% for Grok 3. The paper also describes design suggestions and dynamic collaboration patterns derived from the user study.

Significance. The conceptual framing is timely: making the human a cognitive overseer rather than an end-user of black-box research agents is a productive design direction, and the interface features described in Section 3.2 are concrete and well-motivated. The detailed user-study protocol in Appendix B is a useful artifact for reproducibility, and the BrowseComp-ZH evaluation addresses a relevant Chinese-language benchmark. If the reported accuracy gains were causally attributable to interaction itself, the result would be important for the deep-research and human-AI collaboration communities. However, the current evidence base does not support the headline claims: the user-study protocol introduces a demand characteristic, the Table 5 benchmark condition confounds interaction with participant expertise and question selection, and the abstract reports numbers that do not appear in Table 4. These are load-bearing threats to the central claim that interaction constitutes intelligence.

major comments (5)
  1. [Appendix B.1] The participant instructions state: "you should aim to achieve 4-5 points across all dimensions before stopping generation." This is a direct instruction to give high ratings, so the user-study scores in Table 4 and the abstract's percentage improvements likely reflect compliance rather than authentic assessments. The manuscript does not acknowledge or control for this demand characteristic, and it is therefore not possible to interpret the Transparency, Interruptibility, or Fine-Grained Interaction gains as genuine user experience differences.
  2. [Table 5 and Section 5.2] The benchmark ablation confounds interaction with participant expertise. The three Deep Cognition conditions vary two independent factors at once: DC (cog+int) used 4 graduate-level participants, DC (non cog) used 4 middle-school-level participants, and DC (non int) had no human in the loop. The 31.8 and 50.0 percentage-point gaps relative to Gemini/OpenAI and Grok could therefore be selection effects of expert knowledge or of the author-selected 22-question subset, not effects of transparent, interruptible interaction. With only 4 participants per condition and no confidence intervals, clustering-aware error bars, or significance tests, the abstract's "31.8% to 50.0% points of improvements" claim is not supported as an effect of interaction. An experiment that holds participant expertise fixed across conditions, samples questions at random, and reports inferential statistics is needed.
  3. [Abstract vs. Table 4] The abstract reports Transparency +20.0%, Fine-Grained Interaction +29.2%, Real-Time Intervention +18.5%, Ease of Collaboration +27.7%, Results-Worth-Effort +8.8%, and Interruptibility +20.7%. Table 4 shows +25.0%, +44.6%, +24.4%, +43.0%, +10.8%, and +31.4%, respectively. These numbers disagree, and the abstract values appear nowhere in the manuscript. The authors should correct the abstract and verify all reported values against their data.
  4. [Sections 4.1.2 and 5.1] The six headline interaction metrics are defined by the authors and describe capabilities that, per Table 1, only Deep Cognition implements. For example, Transparency is defined as visible decision-making and Deep Cognition is the only system with a visible reasoning process; Interruptibility is defined as pause/resume and Deep Cognition is the only system with a pause feature. Under these definitions, high scores on these metrics are near-tautological and do not test the paper's thesis that interaction is a source of intelligence. A more meaningful test would compare Deep Cognition against a control with comparable interaction affordances, or would tie the interaction metrics to an external task-performance measure.
  5. [Table 3 and Section 5.1] Table 3 reports a "63%" improvement of Deep Cognition over "DC (non)" without specifying the number of raters, the expertise of raters, or how the non-interactive condition was generated and evaluated. The text does not state whether these are expert annotations, participant ratings, or automated metrics, and no error bars are provided. The claim of a 63% average improvement is therefore not verifiable from the current description.
minor comments (5)
  1. [Abstract] The phrase "31.8% to 50.0% points" should read "31.8 to 50.0 percentage points," since these are differences in accuracy, not relative improvements.
  2. [Table 4] The header "Fine-Grained Interation" contains a typo and should be "Fine-Grained Interaction."
  3. [Section 3.1.3] The claim that the preference agent implements "In-Context Reinforcement Learning" is asserted without architectural detail or experimental evidence; if ICRL is intended only as an analogy, that should be stated explicitly.
  4. [Appendix B.2.2] The rubric for "Inspirational Perspectives" and the question about "Long-term Collaboration Willingness" each appear twice in the appendix; the duplicated lines should be removed.
  5. [Naming conventions] The notation for ablated conditions is inconsistent: Table 3 uses "DC (non)" while Table 5 uses "DC (non int)" and "DC (non cog)"; the notation should be unified and defined at first use.

Circularity Check

3 steps flagged · score 6.0 of 10

Partial circularity: the headline interaction metrics are defined into the system, the user-study protocol instructs the target scores, and the benchmark 'ablation' is confounded with participant expertise.

  1. self definitional [Section 4.1.2, Table 2; Section 2.2, Table 1]
    "Transparency ... Assesses the interpretability and explainability of the model's decision-making processes and reasoning mechanisms. ... Interruptibility ... Assesses the system's ability to tolerate pauses or context switches and to resume smoothly without loss of state or progress. ... Fine-Grained Interaction ... Evaluates the system's capacity to incorporate user feedback and enable precise, granular control over output generation."

    These interaction metrics are defined by features that Table 1 assigns exclusively to Deep Cognition (OpenAI, Gemini, and Grok 3 are marked × on Transparency, Real-Time Intervention, Fine-Grained Interaction, and Cognitive Oversight; DC is marked ✓). Asking users to rate exactly the features only one system possesses makes the reported advantages (e.g., Transparency 5.00, +25.0%; Fine-Grained Interaction +44.6%) restatements of the system design rather than independent tests of the interaction hypothesis.

  2. fitted input called prediction [Appendix B.1, Pre-Study]
    "You need to guide the model to improve report writing depth and information retrieval efficiency through various interaction methods during the model's research process (interruption, adding expert prior knowledge, reviewing model-retrieved information, auditing the model's self-evaluation process, new thinking, strategic guidance, or personal files). Please note that you should aim to achieve 4-5 points across all dimensions before stopping generation."

    The protocol makes reaching 4-5 on the evaluation dimensions the participant's stopping criterion. The same participants then supply the 1-5 ratings reported in Tables 3 and 4 (for example, Organization +97%, Cutting-Edge +79%, Transparency 5.00). The high scores are therefore produced by the instruction itself—a target score is written into the procedure and then reported as an evaluation outcome, rather than measured as an independent result.

1 more flagged steps
  1. self definitional [Section 4.2 Benchmark Evaluation; Section 5.2, Table 5]
    "we selected 22 questions (top two from each of 11 categories) for comprehensive assessment. ... DC (non cog). represents the baseline condition with participants possessing foundational knowledge levels (n=4 participants with middle school-level knowledge); DC (non int). represents the autonomous system condition without human intervention; DC (cog+int). represents the interactive condition with graduate-level participants engaging in real-time collaboration with the system (n=4 participants). ..."

    The claimed ablation does not isolate the interaction mechanism: the 'cognition + interaction' condition uses graduate-level participants, the 'without cognition' condition uses middle-school-level participants, and the 'without interaction' condition has no participant at all. The conclusion that the combination of cognition and interaction is what drives the 72.73% accuracy is therefore forced by the way the conditions are defined; expertise level and human presence vary together with the intended independent variable. A cleaner design would hold participant expertise fixed and toggle only the interaction mechanism.

full rationale

The paper's central empirical claims are only partially independent of their inputs. The headline interaction metrics (Transparency, Interruptibility, Fine-Grained Interaction, Real-Time Intervention) are defined by features that only Deep Cognition implements, so its superiority on those scales is close to a restatement of the design. The Appendix B.1 instruction to aim for 4-5 points on all dimensions before stopping turns the reported 1-5 user ratings into a compliance measure rather than an independent evaluation. The BrowseComp-ZH result is grounded in an external benchmark, but the Table 5 'ablation' confounds the interaction mechanism with participant expertise and question selection, so the 31.8-50.0pp claim cannot be attributed to the interactive cognitive-oversight design. No load-bearing self-citation chain or imported uniqueness theorem was found; self-citations to prior same-group work (e.g., DeepResearcher [97][98]) are not used to justify the empirical claims. The paper would be substantially strengthened by a protocol that does not pre-specify target scores, and by a benchmark ablation that holds expertise constant while toggling only interaction.

Assumptions & free parameters 3 free parameters · 3 assumptions · 0 invented entities

The claimed improvements depend on hand-chosen protocol values (k=5, 22-question benchmark subset, 4 benchmark participants), on the assumption that Likert self-reports measure the target constructs, on the representativeness of the selected benchmark subset, and on an unverified ICRL description of the preference agent.

free parameters (3)
  • k (number of search results per query) = 5
    Set in Section 3.1.2 (footnote 2: 'We set k=5 in this work.') and affects the information available to the system.
  • Benchmark subset size from BrowseComp-ZH = 22 questions (top 2 from each of 11 categories)
    Selected in Section 4.2 without stated criteria; the choice can change reported accuracy.
  • Participants in benchmark main condition = 4 graduate students
    Table 5 reports n=4 for DC (cog+int); the 72.73% accuracy rests on this small sample.
assumptions (3)
  • domain assumption Self-reported Likert ratings accurately measure transparency, interruptibility, and collaboration quality
    The user evaluation relies on 5-point scales (Table 2, Appendix B) without validation against objective measures.
  • domain assumption The 22-question BrowseComp-ZH subset is representative of deep research performance
    Section 4.2 selects 'top two from each of 11 categories' with no justification; claims of 31.8-50.0 point improvements depend on this subset.
  • ad hoc to paper The preference agent's behavior is correctly described as In-Context Reinforcement Learning
    Section 3.1.3 invokes ICRL references but provides no algorithm, training, or measured adaptation; the description is not independently evidenced.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Interaction as Intelligence: Deep Research With Human-AI Partnership." pith.science (2026). https://pith.science/paper/CZNLTBCL

@misc{pith2026250715759,
  author       = {Pith},
  title        = {Pith review of: Interaction as Intelligence: Deep Research With Human-AI Partnership},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/CZNLTBCL}},
  note         = {Machine review of arXiv:2507.15759}
}
read the original abstract

This paper introduces "Interaction as Intelligence" research series, presenting a reconceptualization of human-AI relationships in deep research tasks. Traditional approaches treat interaction merely as an interface for accessing AI capabilities-a conduit between human intent and machine output. We propose that interaction itself constitutes a fundamental dimension of intelligence. As AI systems engage in extended thinking processes for research tasks, meaningful interaction transitions from an optional enhancement to an essential component of effective intelligence. Current deep research systems adopt an "input-wait-output" paradigm where users initiate queries and receive results after black-box processing. This approach leads to error cascade effects, inflexible research boundaries that prevent question refinement during investigation, and missed opportunities for expertise integration. To address these limitations, we introduce Deep Cognition, a system that transforms the human role from giving instructions to cognitive oversight-a mode of engagement where humans guide AI thinking processes through strategic intervention at critical junctures. Deep cognition implements three key innovations: (1)Transparent, controllable, and interruptible interaction that reveals AI reasoning and enables intervention at any point; (2)Fine-grained bidirectional dialogue; and (3)Shared cognitive context where the system observes and adapts to user behaviors without explicit instruction. User evaluation demonstrates that this cognitive oversight paradigm outperforms the strongest baseline across six key metrics: Transparency(+20.0%), Fine-Grained Interaction(+29.2%), Real-Time Intervention(+18.5%), Ease of Collaboration(+27.7%), Results-Worth-Effort(+8.8%), and Interruptibility(+20.7%). Evaluations on challenging research problems show 31.8% to 50.0% points of improvements over deep research systems.

Figures

Figures reproduced from arXiv: 2507.15759 by the authors.

Figure 1
Figure 1. Overall evaluation results. We present the user evaluation (seven metrics on the left part), report quality (six metrics in the middle), and evaluation results on deep research problems (the right part) with three conditions: Without Cognition, With Cognition & Interaction, and Without Interaction. These results demonstrate the cognitive amplification effect of deep cognition when users collaborate with AI to perfor… view at source ↗
Figure 2
Figure 2. The evolution of human machine interaction from [PITH_FULL_IMAGE:figures/full_fig_p002_2.png] view at source ↗
Figure 3
Figure 3. Deep cognition framework overview. This human-in-the-loop research assistant system breaks down [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figures from the paper (4 more)
Figure 4
Figure 4. Figure 4: Deep cognition interface design showcasing key interactive features: (A) Research scope clarification [PITH_FULL_IMAGE:figures/full_fig_p007_4.png]
Figure 5
Figure 5. Figure 5: Left: Distribution of participant ratings (1–5) indicating the extent to which each system feature benefited [PITH_FULL_IMAGE:figures/full_fig_p010_5.png]
Figure 6
Figure 6. Figure 6: Changes in users’ behavioral tendencies when using the deep research system to perform complex [PITH_FULL_IMAGE:figures/full_fig_p011_6.png]
Figure 7
Figure 7. Figure 7: Human–AI collaboration code book B User Study Protocol B.1 Pre-Study Study Overview This protocol evaluates four AI research systems: deep cognition, OpenAI Deep Research (O3), Grok 3 Deeper Search, and Gemini Deep Research (default). Participants complete authentic re…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

100 extracted references · 23 canonical work pages

  1. [1]

    URL: " 'urlintro :=

    ENTRY address author booktitle chapter edition editor howpublished institution journal key month note number organization pages publisher school series title type volume year eprint doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRINGS urlintro eprinturl eprintpr...

  2. [2]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...

  3. [3]

    Anthropic. 2025. https://www.anthropic.com/claude/sonnet Claude sonnet 4: Hybrid reasoning model with superior intelligence for high-volume use cases, and 200k context window

  4. [4]

    Akari Asai, Jacqueline He, Rulin Shao, Weijia Shi, Amanpreet Singh, Joseph Chee Chang, Kyle Lo, Luca Soldaini, Sergey Feldman, Mike D'arcy, David Wadden, Matt Latzke, Minyang Tian, Pan Ji, Shengyan Liu, Hao Tong, Bohao Wu, Yanyu Xiong, Luke Zettlemoyer, Graham Neubig, Dan Weld, Doug Downey, Wen tau Yih, Pang Wei Koh, and Hannaneh Hajishirzi. 2024. http://...

  5. [5]

    Lisanne Bainbridge. 1983 a . Ironies of automation. In Analysis, design and evaluation of man--machine systems, pages 129--135. Elsevier

  6. [6]

    Lisanne Bainbridge. 1983 b . https://doi.org/https://doi.org/10.1016/0005-1098(83)90046-8 Ironies of automation . Automatica, 19(6):775--779

  7. [7]

    Gagan Bansal, Jennifer Wortman Vaughan, Saleema Amershi, Eric Horvitz, Adam Fourney, Hussein Mozannar, Victor Dibia, and Daniel S. Weld. 2024. http://arxiv.org/abs/2412.10380 Challenges in human-agent communication

  8. [8]

    Alexandra Bremers and Wendy Ju. 2024. Can machines tell what people want? bringing situated intelligence to generative ai. In Proceedings of the Halfway to the Future Symposium, pages 1--6

Show all 100 references
  1. [9]

    Le, Christopher Ré, and Azalia Mirhoseini

    Bradley Brown, Jordan Juravsky, Ryan Ehrlich, Ronald Clark, Quoc V. Le, Christopher Ré, and Azalia Mirhoseini. 2024. http://arxiv.org/abs/2407.21787 Large language monkeys: Scaling inference compute with repeated sampling

  2. [10]

    Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffr...

  3. [11]

    ByteDance . 2024. https://github.com/bytedance/deer-flow Deerflow . Community-driven deep research framework combining LLMs with web search, crawling, and code execution tools

  4. [12]

    Runze Cai, Nuwan Janaka, Hyeongcheol Kim, Yang Chen, Shengdong Zhao, Yun Huang, and David Hsu. 2025. https://doi.org/10.1145/3706598.3713953 Aiget: Transforming everyday moments into hidden knowledge discovery with ai assistance on smart glasses . In Proceedings of the 2025 CH...

  5. [13]

    Nicolas Camara. 2025. https://github.com/nickscamara/open-deep-research Open deep research . Open-source clone of OpenAI's Deep Research using Firecrawl for web data extraction and AI reasoning

  6. [14]

    Heloisa Candello and Claudio Pinhanez. 2016. Designing conversational interfaces. Simp \'o sio Brasileiro sobre Fatores Humanos em Sistemas Computacionais-IHC , 100

  7. [15]

    Yining Cao, Peiling Jiang, and Haijun Xia. 2025. https://doi.org/10.1145/3706598.3713285 Generative and malleable user interfaces with generative and evolving task-driven data model . In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems, CHI '25, New...

  8. [16]

    Pan, Shuyi Yang, Lakshya A

    Mert Cemri, Melissa Z. Pan, Shuyi Yang, Lakshya A. Agrawal, Bhavya Chopra, Rishabh Tiwari, Kurt Keutzer, Aditya Parameswaran, Dan Klein, Kannan Ramchandran, Matei Zaharia, Joseph E. Gonzalez, and Ion Stoica. 2025. http://arxiv.org/abs/2503.13657 Why do multi-agent llm systems fail?

  9. [17]

    Mingyang Chen, Tianpeng Li, Haoze Sun, Yijie Zhou, Chenzheng Zhu, Fan Yang, Zenan Zhou, Weipeng Chen, Haofen Wang, Jeff Z Pan, et al. 2025 a . Learning to reason with search for llms via reinforcement learning. arXiv preprint arXiv:2503.19470

  10. [18]

    Si Chen, Haocong Cheng, and Yun Huang. 2024. https://doi.org/10.1007/978-3-031-64487-0_9 Emotion Recognition in Self-Regulated Learning: Advancing Metacognition Through AI-Assisted Reflections , pages 185--212. Springer Nature Switzerland, Cham

  11. [19]

    Valerie Chen, Alan Zhu, Sebastian Zhao, Hussein Mozannar, David Sontag, and Ameet Talwalkar. 2025 b . Need help? designing proactive ai assistants for programming. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems, pages 1--18

  12. [20]

    Collins, Ilia Sucholutsky, Umang Bhatt, Kartik Chandra, Lionel Wong, Mina Lee, Cedegao E

    Katherine M. Collins, Ilia Sucholutsky, Umang Bhatt, Kartik Chandra, Lionel Wong, Mina Lee, Cedegao E. Zhang, Tan Zhi-Xuan, Mark Ho, Vikash Mansinghka, Adrian Weller, Joshua B. Tenenbaum, and Thomas L. Griffiths. 2024. http://arxiv.org/abs/2408.03943 Building machines that lea...

  13. [21]

    Zhang, Han Bao, Hanwei Xu, Haocheng Wang, Haowei Zhang, Honghui Ding, Huajian Xin, Huazuo Gao, Hui Li, Hui Qu, J

    DeepSeek-AI, Aixin Liu, Bei Feng, Bing Xue, Bingxuan Wang, Bochao Wu, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, Damai Dai, Daya Guo, Dejian Yang, Deli Chen, Dongjie Ji, Erhang Li, Fangyun Lin, Fucong Dai, Fuli Luo, Guangbo Hao, Guanting Chen, Guowei L...

  14. [23]

    Assaf Elovic. 2025. https://github.com/assafelovic/gpt-researcher Gpt researcher . Open deep research agent for web and local research with detailed report generation and citations

  15. [25]

    Raymond Fok and Daniel S. Weld. 2024. https://doi.org/10.1002/aaai.12182 In search of verifiability: Explanations rarely enable complementary performance in ai‑advised decision making . AI Magazine, 45(3):317--332

  16. [26]

    George Fragiadakis, Christos Diou, George Kousiouris, and Mara Nikolaidou. 2025. http://arxiv.org/abs/2407.19098 Evaluating human-ai collaboration: A review and methodological framework

  17. [27]

    Melinda Gervasio, Pedro Sequeira, Eric Yeh, Nicholas Marion, Sarah Bakst, and Helen Gent. 2025. Ai as collaborative partner: Rethinking human-ai teaming for the real world. In Proceedings of the AAAI Symposium Series, volume 5, pages 63--66

  18. [28]

    Catalina Gomez, Sue Min Cho, Shichang Ke, Chien-Ming Huang, and Mathias Unberath. 2025. https://doi.org/10.3389/fcomp.2024.1521066 Human-ai collaboration is not very collaborative yet: a taxonomy of interaction patterns in ai-assisted decision making from a systematic review ....

  19. [29]

    Google . 2025. https://gemini.google/overview/deep-research/ Gemini deep research - your personal research assistant . Accessed: April 14, 2025

  20. [30]

    Jake Grigsby, Linxi Fan, and Yuke Zhu. 2024. http://arxiv.org/abs/2310.09971 Amago: Scalable in-context reinforcement learning for adaptive agents

  21. [31]

    Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, et al. 2025. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning. arXiv preprint arXiv:2501.12948

  22. [32]

    Rae, and Laurent Sifre

    Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, Tom Hennigan, Eric Noland, Katie Millican, George van den Driessche, Bogdan Damoc, Aurelia Guy, Simon Osin...

  23. [33]

    Sili Huang, Jifeng Hu, Hechang Chen, Lichao Sun, and Bo Yang. 2024. http://arxiv.org/abs/2405.20692 In-context decision transformer: Reinforcement learning via hierarchical chain-of-thought

  24. [34]

    Edwin Hutchins. 1995. Cognition in the Wild. MIT press

  25. [35]

    Folasade Olubusola Isinkaye, Yetunde O Folajimi, and Bolande Adefowoke Ojokoh. 2015. Recommendation systems: Principles, methods and evaluation. Egyptian informatics journal, 16(3):261--273

  26. [36]

    Semnani, and Monica S

    Yucheng Jiang, Yijia Shao, Dekun Ma, Sina J. Semnani, and Monica S. Lam. 2024. http://arxiv.org/abs/2408.15232 Into the unknown unknowns: Engaged human learning through participation in language model agent conversations

  27. [37]

    Bowen Jin, Hansi Zeng, Zhenrui Yue, Dong Wang, Hamed Zamani, and Jiawei Han. 2025 a . Search-r1: Training llms to reason and leverage search engines with reinforcement learning. arXiv preprint arXiv:2503.09516

  28. [38]

    Mingyu Jin, Weidi Luo, Sitao Cheng, Xinyi Wang, Wenyue Hua, Ruixiang Tang, William Yang Wang, and Yongfeng Zhang. 2025 b . http://arxiv.org/abs/2411.13504 Disentangling memory and reasoning ability in large language models

  29. [39]

    Jina AI . 2025. https://github.com/jina-ai/node-DeepResearch node-deepresearch . Iterative search, reading, and reasoning system for deep research queries with focus on concise answers

  30. [40]

    Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei

    Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. 2020. http://arxiv.org/abs/2001.08361 Scaling laws for neural language models

  31. [41]

    help me help the ai

    Sunnie S. Y. Kim, Elizabeth Anne Watkins, Olga Russakovsky, Ruth Fong, and Andr\' e s Monroy-Hern\' a ndez. 2023. https://doi.org/10.1145/3544548.3581001 "help me help the ai": Understanding how explainability can support human-ai interaction . In Proceedings of the 2023 CHI C...

  32. [42]

    Ziegler, Elizabeth Barnes, and Lawrence Chan

    Thomas Kwa, Ben West, Joel Becker, Amy Deng, Katharyn Garcia, Max Hasin, Sami Jawhar, Megan Kinniment, Nate Rush, Sydney Von Arx, Ryan Bloom, Thomas Broadley, Haoxing Du, Brian Goodrich, Nikola Jurkovic, Luke Harold Miles, Seraphina Nix, Tao Lin, Neev Parikh, David Rein, Lucas...

  33. [43]

    LangChain AI . 2025. https://github.com/langchain-ai/open_deep_research Open deep research . Open-source research assistant for automated deep research and report generation

  34. [44]

    Michael Laskin, Luyu Wang, Junhyuk Oh, Emilio Parisotto, Stephen Spencer, Richie Steigerwald, DJ Strouse, Steven Hansen, Angelos Filos, Ethan Brooks, Maxime Gazeau, Himanshu Sahni, Satinder Singh, and Volodymyr Mnih. 2022. http://arxiv.org/abs/2210.14215 In-context reinforceme...

  35. [45]

    Lee, Annie Xie, Aldo Pacchiano, Yash Chandak, Chelsea Finn, Ofir Nachum, and Emma Brunskill

    Jonathan N. Lee, Annie Xie, Aldo Pacchiano, Yash Chandak, Chelsea Finn, Ofir Nachum, and Emma Brunskill. 2023. http://arxiv.org/abs/2306.14892 Supervised pretraining can learn in-context reinforcement learning

  36. [46]

    Wang, Minae Kwon, Joon Sung Park, Hancheng Cao, Tony Lee, Rishi Bommasani, Michael Bernstein, and Percy Liang

    Mina Lee, Megha Srivastava, Amelia Hardy, John Thickstun, Esin Durmus, Ashwin Paranjape, Ines Gerard-Ursin, Xiang Lisa Li, Faisal Ladhak, Frieda Rong, Rose E. Wang, Minae Kwon, Joon Sung Park, Hancheng Cao, Tony Lee, Rishi Bommasani, Michael Bernstein, and Percy Liang. 2024. h...

  37. [47]

    Vera Liao, Junti Zhang, and Yi-Chieh Lee

    Jingshu Li, Yitian Yang, Q. Vera Liao, Junti Zhang, and Yi-Chieh Lee. 2025. http://arxiv.org/abs/2501.12868 As confidence aligns: Exploring the effect of ai confidence on human self-confidence in human-ai decision making

  38. [48]

    Licong Lin, Yu Bai, and Song Mei. 2024. http://arxiv.org/abs/2310.08566 Transformers as decision makers: Provable in-context reinforcement learning via supervised pretraining

  39. [49]

    Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig. 2023. Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing. ACM computing surveys, 55(9):1--35

  40. [50]

    Xingyu Bruce Liu, Shitao Fang, Weiyan Shi, Chien-Sheng Wu, Takeo Igarashi, and Xiang'Anthony' Chen. 2025 a . Proactive conversational agents with inner thoughts. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems, pages 1--19

  41. [51]

    Xingyu Bruce Liu, Haijun Xia, and Xiang Anthony Chen. 2025 b . http://arxiv.org/abs/2502.18676 Interacting with thoughtful ai

  42. [52]

    Yiren Liu, Si Chen, Haocong Cheng, Mengxia Yu, Xiao Ran, Andrew Mo, Yiliu Tang, and Yun Huang. 2024 a . https://doi.org/10.1145/3613904.3642698 How ai processing delays foster creativity: Exploring research question co-creation with an llm-based agent . In Proceedings of the 2...

  43. [53]

    Yiren Liu, Pranav Sharma, Mehul Jitendra Oswal, Haijun Xia, and Yun Huang. 2024 b . http://arxiv.org/abs/2409.12538 Personaflow: Boosting research ideation with llm-simulated expert personas

  44. [54]

    Liangchen Luo, Yinxiao Liu, Rosanne Liu, Samrat Phatale, Harsh Lara, Yunxuan Li, Lei Shu, Yun Zhu, Lei Meng, Jiao Sun, et al. 2024. https://arxiv.org/abs/2406.06592 Improve mathematical reasoning in language models by automated process supervision . ArXiv preprint, abs/2406.06592

  45. [55]

    Michael Frederick McTear, Zoraida Callejas, and David Griol. 2016. The conversational interface, volume 6. Springer

  46. [56]

    Meta AI . 2025. https://ai.meta.com/blog/llama-4-multimodal-intelligence/ The llama 4 herd: The beginning of a new era of natively multimodal ai innovation . Accessed:

  47. [57]

    Bryan Min, Allen Chen, Yining Cao, and Haijun Xia. 2025. https://doi.org/10.1145/3706598.3714164 Malleable overview-detail interfaces . In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems, CHI '25, New York, NY, USA. Association for Computing Machinery

  48. [58]

    Bryan Min and Haijun Xia. 2025. http://arxiv.org/abs/2502.14229 Feedforward in generative ai: Opportunities for a design space

  49. [59]

    MiniMax, Aonian Li, Bangwei Gong, Bo Yang, Boji Shan, Chang Liu, Cheng Zhu, Chunhao Zhang, Congchao Guo, Da Chen, Dong Li, Enwei Jiao, Gengxin Li, Guojun Zhang, Haohai Sun, Houze Dong, Jiadai Zhu, Jiaqi Zhuang, Jiayuan Song, Jin Zhu, Jingtao Han, Jingyang Li, Junbin Xie, Junha...

  50. [60]

    Marvin Minsky. 1987. http://www.jstor.org/stable/20708493 The society of mind . The Personalist Forum, 3(1):19--32

  51. [61]

    OpenAI, :, Aaron Jaech, Adam Kalai, Adam Lerer, Adam Richardson, Ahmed El-Kishky, Aiden Low, Alec Helyar, Aleksander Madry, Alex Beutel, Alex Carney, Alex Iftimie, Alex Karpenko, Alex Tachard Passos, Alexander Neitz, Alexander Prokofiev, Alexander Wei, Allison Tam, Ally Bennet...

  52. [62]

    OpenAI . 2022. https://openai.com/index/chatgpt/ Introducing chatgpt

  53. [63]

    OpenAI. 2024. https://openai.com/index/learning-to-reason-with-llms/ Learning to reason with llms, september 2024

  54. [64]

    OpenAI . 2025 a . https://cdn.openai.com/deep-research-system-card.pdf Deep research system card . Accessed: April 14, 2025

  55. [65]

    OpenAI . 2025 b . https://openai.com/index/introducing-deep-research/ Introducing deep research . Accessed: April 14, 2025

  56. [66]

    Michael J Pazzani and Daniel Billsus. 2007. Content-based recommendation systems. In The adaptive web: methods and strategies of web personalization, pages 325--341. Springer

  57. [67]

    Perplexity AI . 2025. https://www.perplexity.ai/hub/blog/introducing-perplexity-deep-research Introducing perplexity deep research . Accessed: April 14, 2025

  58. [68]

    Peter Pirolli. 2009. Powers of 10: Modeling complex information-seeking systems at multiple scales. Computer, 42(3):33--40

  59. [69]

    Michael Poli, Armin W Thomas, Eric Nguyen, Pragaash Ponnusamy, Björn Deiseroth, Kristian Kersting, Taiji Suzuki, Brian Hie, Stefano Ermon, Christopher Ré, Ce Zhang, and Stefano Massaroli. 2024. http://arxiv.org/abs/2403.17844 Mechanistic design and scaling of hybrid architectures

  60. [70]

    Kevin Pu, Daniel Lazaro, Ian Arawjo, Haijun Xia, Ziang Xiao, Tovi Grossman, and Yan Chen. 2025. https://doi.org/10.1145/3706598.3713357 Assistance or disruption? exploring and evaluating the design and trade-offs of proactive ai programming support . In Proceedings of the 2025...

  61. [71]

    Peinuan Qin, Chi-Lan Yang, Jingshu Li, Jing Wen, and Yi-Chieh Lee. 2025. https://doi.org/10.1145/3706598.3713146 Timing matters: How using llms at different timings influences writers' perceptions and ideation outcomes in ai-assisted ideation . In Proceedings of the 2025 CHI C...

  62. [72]

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021. http://arxiv.org/abs/2103.00020 Learning transferable visual models from natural lang...

  63. [73]

    Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. 2018. Improving language understanding by generative pre-training

  64. [74]

    Jude Rayan, Dhruv Kanetkar, Yifan Gong, Yuewen Yang, Srishti Palani, Haijun Xia, and Steven P. Dow. 2024. https://doi.org/10.1145/3635636.3656184 Exploring the potential for generative ai-based conversational cues for real-time collaborative ideation . In Proceedings of the 16...

  65. [75]

    Paul Resnick and Hal R Varian. 1997. Recommender systems. Communications of the ACM, 40(3):56--58

  66. [77]

    Aymeric Roucher, Albert Villanova del Moral, Thomas Wolf, Leandro von Werra, and Erik Kaunismäki. 2025. `smolagents`: a smol library to build great agentic systems. https://github.com/huggingface/smolagents

  67. [78]

    Kanell, Peter Xu, Omar Khattab, and Monica S

    Yijia Shao, Yucheng Jiang, Theodore A. Kanell, Peter Xu, Omar Khattab, and Monica S. Lam. 2024. http://arxiv.org/abs/2402.14207 Assisting in writing wikipedia-like articles from scratch with large language models

  68. [79]

    Yijia Shao, Humishka Zope, Yucheng Jiang, Jiaxin Pei, David Nguyen, Erik Brynjolfsson, and Diyi Yang. 2025. http://arxiv.org/abs/2506.06576 Future of work with ai agents: Auditing automation and augmentation potential across the u.s. workforce

  69. [80]

    Wenxuan Shi, Haochen Tan, Chuqiao Kuang, Xiaoguang Li, Xiaozhe Ren, Chen Zhang, Hanting Chen, Yasheng Wang, Lifeng Shang, Fisher Yu, and Yunhe Wang. 2025. http://arxiv.org/abs/2505.24332 Pangu deepdiver: Adaptive search intensity scaling via open-web reinforcement learning

  70. [81]

    Huatong Song, Jinhao Jiang, Yingqian Min, Jie Chen, Zhipeng Chen, Wayne Xin Zhao, Lei Fang, and Ji-Rong Wen. 2025. R1-searcher: Incentivizing the search capability in llms via reinforcement learning. arXiv preprint arXiv:2503.05592

  71. [82]

    Sakhinana Sagar Srinivas and Venkataramana Runkana. 2025. http://arxiv.org/abs/2504.01281 Scaling test-time inference with policy-optimized, dynamic retrieval-augmented generation via kv caching and decoding

  72. [83]

    Sangho Suh, Meng Chen, Bryan Min, Toby Jia-Jun Li, and Haijun Xia. 2024. https://doi.org/10.1145/3613904.3642400 Luminate: Structured generation and exploration of design space with large language models for human-ai co-creation . In Proceedings of the 2024 CHI Conference on H...

  73. [84]

    Kimi Team, Angang Du, Bofei Gao, Bowei Xing, Changjiu Jiang, Cheng Chen, Cheng Li, Chenjun Xiao, Chenzhuang Du, Chonghua Liao, Chuning Tang, Congcong Wang, Dehao Zhang, Enming Yuan, Enzhe Lu, Fengxiang Tang, Flood Sung, Guangda Wei, Guokun Lai, Haiqing Guo, Han Zhu, Hao Ding, ...

  74. [85]

    Xinru Wang, Mengjie Yu, Hannah Nguyen, Michael Iuzzolino, Tianyi Wang, Peiqi Tang, Natasha Lynova, Co Tran, Ting Zhang, Naveen Sendhilnathan, Hrvoje Benko, Haijun Xia, and Tanya R. Jonker. 2025. https://doi.org/10.1145/3708359.3712074 Less or more: Towards glanceable explanati...

  75. [86]

    Elizabeth Anne Watkins, Emanuel Moss, Giuseppe Raffa, and Lama Nachman. 2025. http://arxiv.org/abs/2503.05926 What's so human about human-ai collaboration, anyway? generative ai and human-computer interaction

  76. [87]

    Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus

    Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, Ed H. Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus. 2022. https://openreview.net/forum?id=yzkSU5zd...

  77. [88]

    Yom-Tov, and Anat Rafaeli

    Monika Westphal, Michael Vössing, Gerhard Satzger, Galit B. Yom-Tov, and Anat Rafaeli. 2023. https://doi.org/https://doi.org/10.1016/j.chb.2023.107714 Decision control and explanations in human-ai collaboration: Improving user perceptions and compliance . Computers in Human Be...

  78. [89]

    Ryen W. White. 2024. http://arxiv.org/abs/2311.01235 Advancing the search frontier with ai agents

  79. [90]

    Chabris, Alex Pentland, Nada Hashmi, and Thomas W

    Anita Williams Woolley, Christopher F. Chabris, Alex Pentland, Nada Hashmi, and Thomas W. Malone. 2010. https://doi.org/10.1126/science.1193147 Evidence for a collective intelligence factor in the performance of human groups . Science, 330(6004):686--688

  80. [91]

    xAI . 2025. https://x.ai/news/grok-3 Grok 3 beta — the age of reasoning agents . Accessed: April 14, 2025

  81. [92]

    Shijie Xia, Yiwei Qin, Xuefeng Li, Yan Ma, Run-Ze Fan, Steffi Chern, Haoyang Zou, Fan Zhou, Xiangkun Hu, Jiahe Jin, et al. 2025. Generative ai act ii: Test time scaling drives cognition engineering. arXiv preprint arXiv:2504.13828

  82. [93]

    Zhengtao Xu, Tianqi Song, and Yi-Chieh Lee. 2025. https://doi.org/https://doi.org/10.1016/j.ijhcs.2025.103455 Confronting verbalized uncertainty: Understanding how llm’s verbalized uncertainty influences users in ai-assisted decision-making . International Journal of Human-Com...

  83. [94]

    An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, Chujie Zheng, Dayiheng Liu, Fan Zhou, Fei Huang, Feng Hu, Hao Ge, Haoran Wei, Huan Lin, Jialong Tang, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jia...

  84. [95]

    Jingyi Yang, Shuai Shao, Dongrui Liu, and Jing Shao. 2025 b . http://arxiv.org/abs/2506.00618 Riosworld: Benchmarking the risk of multimodal computer-use agents

  85. [96]

    Ryan Yen, Jiawen Stefanie Zhu, Sangho Suh, Haijun Xia, and Jian Zhao. 2024. https://doi.org/10.1145/3654777.3676357 Coladder: Manipulating code generation via multi-level blocks . In Proceedings of the 37th Annual ACM Symposium on User Interface Software and Technology, UIST '...

  86. [97]

    Ming Yin. 2025. Bridging the gap between machine confidence and human perceptions. Nature Machine Intelligence, pages 1--2

  87. [98]

    David Zhang. 2025. https://github.com/dzhng/deep-research Deep research . AI-powered research assistant for iterative, deep research using search engines, web scraping, and LLMs

  88. [100]

    Yuxiang Zheng, Dayuan Fu, Xiangkun Hu, Xiaojie Cai, Lyumanshan Ye, Pengrui Lu, and Pengfei Liu. 2025 b . http://arxiv.org/abs/2504.03160 Deepresearcher: Scaling deep research via reinforcement learning in real-world environments

  89. [101]

    Peilin Zhou, Bruce Leon, Xiang Ying, Can Zhang, Yifan Shao, Qichen Ye, Dading Chong, Zhiling Jin, Chenxuan Xie, Meng Cao, Yuxin Gu, Sixin Hong, Jing Ren, Jian Chen, Chao Liu, and Yining Hua. 2025. http://arxiv.org/abs/2504.19314 Browsecomp-zh: Benchmarking web browsing ability...

  90. [102]

    @esa (Ref

    \@ifxundefined[1] #1\@undefined \@firstoftwo \@secondoftwo \@ifnum[1] #1 \@firstoftwo \@secondoftwo \@ifx[1] #1 \@firstoftwo \@secondoftwo [2] @ #1 \@temptokena #2 #1 @ \@temptokena \@ifclassloaded agu2001 natbib The agu2001 class already includes natbib coding, so you should ...

  91. [103]

    \@lbibitem[] @bibitem@first@sw\@secondoftwo \@lbibitem[#1]#2 \@extra@b@citeb \@ifundefined br@#2\@extra@b@citeb \@namedef br@#2 \@nameuse br@#2\@extra@b@citeb \@ifundefined b@#2\@extra@b@citeb @num @parse #2 @tmp #1 NAT@b@open@#2 NAT@b@shut@#2 \@ifnum @merge>\@ne @bibitem@firs...

  92. [104]

    collective intelligence

    @open @close @open @close and [1] URL: #1 \@ifundefined chapter * \@mkboth \@ifxundefined @sectionbib * \@mkboth * \@mkboth\@gobbletwo \@ifclassloaded amsart * \@ifclassloaded amsbook * \@ifxundefined @heading @heading NAT@ctr thebibliography [1] @ \@biblabel @NAT@ctr \@bibset...

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.