Pith. sign in

REVIEW 3 major objections 5 minor 89 references

Facets, Taxonomies, and Syntheses: Navigating Structured Representations in LLM-Assisted Literature Review

T0 review · 3 major / 5 minor · reviewed 2026-08-16 · deepseek-v4-flash

Pith's one-line read DimInd shows that guiding literature review through linked table, taxonomy, and synthesis views reduces the mental effort of organizing a 50-paper collection compared with a ChatGPT-assisted workflow.

desk verdict Genuinely new integration of table, taxonomy, and synthesis for LLM-assisted literature review, honestly evaluated — but the less-effort claim rests on unscored self-reports and one weak task-comparability assumption. read the letter →

arxiv 2504.18496 v1 pith:BHHGLSOZ submitted 2025-04-25 cs.HC

classification cs.HC
keywords literaturereviewLLMassistancesensemakingstructuredrepresentationsfacetedtablestaxonomiesnarrativesynthesiscognitiveload
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to show that literature review over a large paper collection can be supported by guiding researchers through a chain of LLM-generated structured representations rather than by chat alone. It presents DimInd, in which researchers first define facets of interest and the system fills a comparison table with short evidence snippets from each paper's full text; those columns are then organized into hierarchical taxonomies of concepts, and selected branches can be summarized into a cited narrative. In a within-subjects study, 23 researchers used DimInd on one 50-paper collection and a ChatGPT-assisted baseline on another. Participants reported significantly less mental effort extracting and organizing information with DimInd, and rated it higher for categorizing papers and verifying generated information, though not for control or confidence. If right, the result means that the bottleneck in LLM-assisted review shifts from extraction to verification and synthesis, and that scalable review tools should provide several linked levels of compression rather than a single table or a free-form chat.

What carries the argument

The central object is a chain of linked structured representations: a paper collection list, a faceted comparison table whose cells are LLM-generated evidence snippets drawn from each paper's full text, a hierarchical facet taxonomy per column that clusters those snippets into themes, and a facet synthesis that summarizes selected branches with inline citations. Each level retains provenance to the source text, so a click can take a reader from a snippet to the highlighted passage in the PDF. This chain carries the argument by externalizing the schema as it forms: the table reduces foraging cost, the taxonomy supports pattern-finding, and the synthesis supports presentation, while provenance makes verification possible.

What would settle it

Run a matched study in which the same 50-paper collection is reviewed under both DimInd and the ChatGPT baseline, with expert raters scoring the resulting outlines blind; if the effort advantage shrinks, disappears, or reverses when task difficulty is held fixed and output quality is scored, the central claim would be refuted.

Watch

Extended reading notes

Core claim

The paper's central claim is that successive, linked structured representations of paper information can lower the cognitive cost of literature review at scale. DimInd transforms a raw paper collection into a faceted comparison table whose cells are LLM-generated evidence snippets traced to the source full text, then clusters each facet's snippets into an editable facet taxonomy, and finally turns selected taxonomy branches into a narrative synthesis with inline citations. In the evaluation, DimInd outperformed a ChatGPT-assisted manual workflow on self-reported effort for extraction and organization (median 6 vs. 5, Wilcoxon p < .01), on meaningful paper categorization (median 6 vs. 5, p < .01), and on ease of verifying system-generated information (median 6 vs. 5, p < .05), while showing no significant difference on user control or confidence in outline quality. The authors interpret these results as evidence that multiple levels of compression, with provenance connecting each level to the papers, better support the sensemaking loop than a single representation or an unstructured conversation.

Load-bearing premise

The two tasks used in the within-subjects comparison are similar enough in difficulty that the measured reduction in effort can be attributed to DimInd rather than to the task; the paper states this but does not report an objective difficulty check or scoring of the outlines.

Editorial extensions

If this is right

  • A reviewer can shift limited attention from reading and extracting to checking, organizing, and deciding what to include.
  • Reviewing dozens of papers in a single sitting becomes feasible for an individual researcher rather than requiring a multi-author team.
  • Verification becomes cheaper because every generated snippet traces back to a highlighted passage in the source PDF.
  • Users can steer LLM output by editing taxonomies, so the review reflects their own categories rather than the model's default framing.
  • LLM assistance is unlikely to replace the researcher's own synthesis; participants did not trust final narrative generation, so tools should focus on pre-writing stages.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial: The same linked-representation recipe should transfer to other document-heavy synthesis domains such as clinical trial review, legal discovery, or competitive intelligence, since the requirements are a collection, definable facets, and source text for verification.
  • Editorial: The observed gap between table-level detail and taxonomy-level themes suggests a two-facet pivot view (for example, application domain by risk) as a natural next mechanism; the authors mention this only as future work, so it remains untested.
  • Editorial: Self-reported effort may partly reflect interface novelty or the participants' familiarity with chat tools; longer-term use could either deepen the benefit or make the fixed structure feel constraining, so longitudinal deployment would be informative.
  • Editorial: The effort reduction likely concentrates in the foraging and extraction stage rather than the writing stage, since participants explicitly resisted delegating the final synthesis; this suggests that future systems should measure value at the early outline stage, not at the finished review.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper presents DimInd, an interactive system that scaffolds literature review over large paper collections through four linked LLM-generated structured representations: a faceted literature table, per-facet taxonomies, and narrative syntheses with provenance to source text. The design is motivated by three goals drawn from sensemaking and information foraging theory. A within-subjects study with 23 CS researchers compared DimInd with a ChatGPT-assisted baseline on two 50-paper outline-creation tasks. Post-task Likert ratings showed significant advantages for DimInd on reduced mental effort for extraction/organization (Q1), meaningful categorization (Q2), and ease of verification (Q5), with no significant difference on perceived control (Q4) or confidence in outline quality (Q6). Qualitative analysis describes how participants used tables as information scent, taxonomies as navigational hubs, and retained scholarly agency by avoiding fully automated synthesis. The paper concludes that the structured representations reduce cognitive load and improve the effectiveness of exploratory literature review.

Significance. If the central claim is accepted, the paper makes a useful contribution to HCI and LLM-assisted knowledge work: it demonstrates a concrete workflow for moving beyond flat tables toward multi-level, provenance-linked representations, and it reports a carefully counterbalanced within-subjects comparison against a realistic ChatGPT baseline. The paper is also transparent in valuable ways: it documents the full LLM prompts, implementation details, the participant demographics, and several limitations. The empirical foundation, however, currently rests almost entirely on self-reported effort and perceived support; no objective measure of the produced outlines is reported, and the two tasks differ in a way that is directly relevant to the effort comparison. These issues are fixable, but they are load-bearing for the strength of the headline claim.

major comments (3)
  1. [Section 5.2] The comparability of the two tasks is load-bearing for the effort comparison, yet the text states that one dimension "could be reasonably developed from abstracts" while the other "required deeper engagement with the full texts." Because DimInd's automatic value extraction is most advantageous in the full-text condition, the measured advantage on Q1 and Q2 could be partly a task effect rather than a system effect. The within-subjects counterbalancing controls order but not task difficulty. Please report per-task and per-condition means, an interaction test, or a manipulation check of perceived task difficulty, and adjust the headline claim until this is resolved.
  2. [Section 5.5] The paper explicitly states, "We did not analyze the content of the outlines participants created during their tasks," which removes the only objective outcome measure. Without blind or rubric-based scoring of outline completeness, structural organization, and citation accuracy, the reduced-effort result is consistent with an effort-quality tradeoff in which DimInd offloads extraction and categorization to the LLM at the cost of output quality or user understanding. The adjacent self-reported confidence item (Q6) was not significant (W=34.5, p=.08), so even perceived quality parity is not established. I recommend adding expert scoring of the outlines, or at least structural metrics such as number of subsections, coverage of the provided sections, and citation counts, and narrowing claims such as "increased paper organization effectiveness" until such evidence is available.
  3. [Section 5.4] The baseline condition did not receive a comparable onboarding session: participants were given a hands-on tutorial for DimInd, while no tutorial was provided for ChatGPT despite the study's own screening criterion of prior familiarity. This asymmetry in training and task novelty is a plausible source of bias in favor of DimInd's perceived-effort ratings. A brief structured practice task for the baseline, or inclusion of prior ChatGPT experience as a covariate, would strengthen the fairness of the comparison.
minor comments (5)
  1. [Section 5.5] Section 5.5 says Bonferroni corrections were applied, but the Results section reports raw p-values without specifying which values are adjusted; please clarify the correction procedure for each test.
  2. [Figure 7] The caption of Figure 7 refers to Q1 through Q6 but does not map these labels to the survey statements in Table 3; please add the mapping to the caption or the figure itself.
  3. [Appendix C.3.1 and C.2.1] There are minor typos in the prompts: "wihtout" (C.3.1), "an facet" (C.2.1), and Section 2.2 "between of two papers" should read "between two papers."
  4. [Section 4.2.1] The text says k=4 and n=4 were chosen "empirically" but does not report the supporting exploration; please either provide a brief summary of that comparison or describe the choice as a design heuristic.
  5. [Section 7.1] The limitation paragraph asserts "we believe this is a sufficient scale" for generalizability without supporting evidence; this is acceptable as a statement of intent, but the phrasing could be softened to "we believe this scale captures the relevant cognitive challenges."

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the headline claim rests on a new empirical comparison against an external ChatGPT baseline, not on a self-derived fit or self-citation chain.

full rationale

The paper's central claim, that DimInd reduced the mental effort of extracting and organizing information during literature review, is an empirical result from a within-subjects user study, not a formal derivation. The quantitative evidence is a direct comparison against a ChatGPT-assisted baseline workflow: participants rated the effort of each condition on a post-task Likert scale (Fig. 7, Q1: W = 21.5, p < .01). This measures the claimed construct directly rather than defining it in terms of the outcome, so there is no self-definitional reduction. The acknowledged decision not to score outline content (Section 5.5: 'We did not analyze the content of the outlines participants created during their tasks due to the significant diversity in form and content, complicating an unbiased expert evaluation') creates a validity threat that the reported effort advantage could reflect an effort-quality tradeoff, but that is a correctness and interpretation concern, not circularity. The design goals are motivated by sensemaking and information foraging theory and by prior systems, including several authored by the same research groups, but those citations motivate the design rather than constitute evidence for the headline effectiveness claim. No fitted parameter is renamed as a prediction, no uniqueness theorem is imported from the authors' prior work, and no equation or construction reduces to its own inputs. The paper is self-contained in the sense that its main claim is tested against an external baseline in a new study, so the circularity score is 0.

Assumptions & free parameters 4 free parameters · 4 assumptions · 0 invented entities

This is an empirical systems paper, not a derivation, so the ledger records hand-set design parameters and domain assumptions. The central claim depends on the comparability of the study tasks, the reliability of LLM extraction, and the validity of self-report measures. No new physical or theoretical entities are introduced.

free parameters (4)
  • facet discovery subset size k = 4
    Chosen empirically in pilot exploration; larger subsets produced generic facets (Section 4.2.1). Affects which facets are suggested and therefore the user experience.
  • facet discovery subset count n = 4
    Chosen empirically because more subsets rarely added unique facets (Section 4.2.1).
  • maximum taxonomy depth = 5
    LLM prompt cap for taxonomy generation (Section 4.2.3); shapes the level of abstraction users can reach.
  • task collection size = 50 papers
    Selected for study feasibility and consistency across the two tasks (Section 5.2); not a fitted value but a hand-set study parameter.
assumptions (4)
  • domain assumption Sensemaking and information foraging models describe the literature review process well enough to support the design goals.
    Design goals DG1-DG3 are explicitly based on Pirolli and Card's sensemaking model and information foraging theory (Section 3). If these models do not apply, the scaffolding rationale weakens.
  • domain assumption LLM-generated snippets, taxonomies, and syntheses reflect the source papers accurately enough for users to organize and verify information.
    The pipeline relies on GPT-4o-mini and o3-mini outputs (Section 4.2). No accuracy audit beyond user self-report and interaction is reported.
  • domain assumption Self-reported Likert ratings and interaction logs are valid measures of literature review support.
    The main effectiveness claims use perceived effort and perceived categorization ability; outline content was deliberately not analyzed (Section 5.5).
  • ad hoc to paper The two review tasks are comparable in difficulty and nature.
    The paper asserts comparability (Section 5.2) but uses different survey papers, topics, and taxonomy dimensions, one relying more on abstracts and the other on full texts.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Facets, Taxonomies, and Syntheses: Navigating Structured Representations in LLM-Assisted Literature Review." pith.science (2026). https://pith.science/paper/BHHGLSOZ

@misc{pith2026250418496,
  author       = {Pith},
  title        = {Pith review of: Facets, Taxonomies, and Syntheses: Navigating Structured Representations in LLM-Assisted Literature Review},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/BHHGLSOZ}},
  note         = {Machine review of arXiv:2504.18496}
}
read the original abstract

Comprehensive literature review requires synthesizing vast amounts of research -- a labor intensive and cognitively demanding process. Most prior work focuses either on helping researchers deeply understand a few papers (e.g., for triaging or reading), or retrieving from and visualizing a vast corpus. Deep analysis and synthesis of large paper collections (e.g., to produce a survey paper) is largely conducted manually with little support. We present DimInd, an interactive system that scaffolds literature review across large paper collections through LLM-generated structured representations. DimInd scaffolds literature understanding with multiple levels of compression, from papers, to faceted literature comparison tables with information extracted from individual papers, to taxonomies of concepts, to narrative syntheses. Users are guided through these successive information transformations while maintaining provenance to source text. In an evaluation with 23 researchers, DimInd supported participants in extracting information and conceptually organizing papers with less effort compared to a ChatGPT-assisted baseline workflow.

Figures

Figures reproduced from arXiv: 2504.18496 by the authors.

Figure 1
Figure 1. We present an LLM-assisted workflow aimed at scaffolding literature review over large paper collections, and instantiate [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Columns can be added to the literature review table in two ways: A) [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. In DimInd, users review large paper collections by navigating and analyzing information across various structured representations. Each cell in the literature review table is a snippet of faceted information from a paper (evidence snippet). Clicking on a snippet shows a popover with additional detail (evidence summary), with a button that can further open the paper PDF in an integrated paper reader with attributed p… view at source ↗
Figures from the paper (5 more)
Figure 4
Figure 4. Figure 4: The facet taxonomy. Each category shows the num [PITH_FULL_IMAGE:figures/full_fig_p006_4.png]
Figure 5
Figure 5. Figure 5: Selecting specific categories in the facet taxonomy: 1) highlights cells for the included papers in the literature review [PITH_FULL_IMAGE:figures/full_fig_p007_5.png]
Figure 6
Figure 6. Figure 6: Users can view additional detail while exploring the synthesized representations: 1) Clicking an evidence snippet in [PITH_FULL_IMAGE:figures/full_fig_p007_6.png]
Figure 7
Figure 7. Figure 7: Participants’ post-task ratings (7-point Likert scale) of literature review utility, control, information verifiability, and [PITH_FULL_IMAGE:figures/full_fig_p010_7.png]
Figure 8
Figure 8. Figure 8: Paper collections in DimInd can be created interactively from a research question or search query [PITH_FULL_IMAGE:figures/full_fig_p017_8.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

89 extracted references · 25 canonical work pages

  1. [1]

    Loulwah AlSumait, Daniel Barbará, James Gentle, and Carlotta Domeniconi. 2009. Topic Significance Ranking of LDA Generative Models. In Machine Learning and Knowledge Discovery in Databases. Springer Berlin Heidelberg, Berlin, Heidelberg, 67–82

  2. [2]

    Nouf Ibrahim Altmami and Mohamed El Bachir Menai. 2022. Automatic summa- rization of scientific articles: A survey.Journal of King Saud University - Computer and Information Sciences 34, 4 (2022), 1011–1028. doi:10.1016/j.jksuci.2020.04.020 arXiv, April, 2025 Raymond Fok, Joseph Chee Chang, Marissa Radensky, Pao Siangliulue, Jonathan Bragg, Amy X. Zhang, ...

  3. [3]

    Anderson, David R

    Lorin W. Anderson, David R. Krathwohl, Peter W. Airasian, Kathleen A. Cruik- shank, Richard E. Mayer, Paul R. Pintrich, James D. Raths, and Merlin C. Wittrock

  4. [4]

    Hearst, Andrew Head, and Kyle Lo

    Tal August, Lucy Lu Wang, Jonathan Bragg, Marti A. Hearst, Andrew Head, and Kyle Lo. 2023. Paper Plain: Making Medical Research Papers Approachable to Healthcare Consumers with Natural Language Processing. ACM Trans. Comput.- Hum. Interact. 30, 5, Article 74 (Sept. 2023), 38 pages. doi:10.1145/3589955

  5. [5]

    Belem, Pouya Pezeskhpour, Hayate Iso, Seiji Maekawa, Nikita Bhutani, and Estevam Hruschka

    Catarina G. Belem, Pouya Pezeskhpour, Hayate Iso, Seiji Maekawa, Nikita Bhutani, and Estevam Hruschka. 2024. From Single to Multi: How LLMs Halluci- nate in Multi-Document Summarization. arXiv:2410.13961 [cs.CL]

  6. [6]

    David M Blei, Andrew Y Ng, and Michael I Jordan. 2003. Latent dirichlet allocation. Journal of machine Learning research 3, Jan (2003), 993–1022

  7. [7]

    Francisco Bolanos, Angelo Salatino, Francesco Osborne, and Enrico Motta. 2024. Artificial intelligence for literature reviews: Opportunities and challenges. Artifi- cial Intelligence Review 57, 10 (2024), 259

  8. [8]

    Rohit Borah, Andrew W Brown, Patrice L Capers, and Kathryn A Kaiser. 2017. Analysis of the time and workers needed to conduct systematic reviews of medical interventions using data from the PROSPERO registry. BMJ open 7, 2 (2017), e012545

Show all 89 references
  1. [9]

    Lutz Bornmann and Rüdiger Mutz. 2015. Growth rates of modern science: A bibliometric analysis based on the number of publications and cited references. Journal of the association for information science and technology 66, 11 (2015), 2215–2222

  2. [10]

    Virginia Braun and Victoria Clarke. 2006. Using thematic analysis in psy- chology. Qualitative Research in Psychology 3, 2 (2006), 77–101. doi:10.1191/ 1478088706qp063oa

  3. [11]

    Jackie Chandler and Sally Hopewell. 2013. Cochrane methods - twenty years experience in developing systematic review methods. Systematic reviews 2 (2013), 1–6

  4. [12]

    Jonathan Chang, Jordan Boyd-Graber, Sean Gerrish, Chong Wang, and David M. Blei. 2009. Reading tea leaves: how humans interpret topic models. InProceedings of the 23rd International Conference on Neural Information Processing Systems (Vancouver, British Columbia, Canada) (Nips...

  5. [13]

    Joseph Chee Chang, Nathan Hahn, and Aniket Kittur. 2020. Mesh: Scaffolding Comparison Tables for Online Decision Making. InProceedings of the 33rd Annual ACM Symposium on User Interface Software and Technology . Acm, Virtual Event USA, 391–405. doi:10.1145/3379337.3415865

  6. [14]

    Joseph Chee Chang, Nathan Hahn, Adam Perer, and Aniket Kittur. 2019. Search- Lens: composing and capturing complex user interests for exploratory search. In Proceedings of the 24th International Conference on Intelligent User Interfaces . Acm, Marina del Ray California, 498–50...

  7. [16]

    Der-Thanq “Victor” Chen, Yu-Mei Wang, and Wei Ching Lee and. 2016. Chal- lenges confronting beginning researchers in conducting literature reviews. Stud- ies in Continuing Education 38, 1 (2016), 47–60. doi:10.1080/0158037x.2015.1030335 arXiv:https://doi.org/10.1080/0158037X.2...

  8. [17]

    Yung-Sung Chuang, Benjamin Cohen-Wang, Shannon Zejiang Shen, Zhaofeng Wu, Hu Xu, Xi Victoria Lin, James Glass, Shang-Wen Li, and Wen tau Yih. 2025. SelfCite: Self-Supervised Alignment for Context Attribution in Large Language Models. arXiv:2502.09604 [cs.CL]

  9. [18]

    Harris Cooper. 2015. Research synthesis and meta-analysis: A step-by-step ap- proach (5 ed.). Applied Social Research Methods Series, Vol. 2. Sage Publications, Thousand Oaks, CA

  10. [19]

    Juliet M Corbin and Anselm Strauss. 1990. Grounded theory research: Procedures, canons, and evaluative criteria. Qualitative sociology 13, 1 (1990), 3–21

  11. [20]

    Ben Daniel. 2022. Common challenges postgraduate students and early-career academics face when engaging with the scholarly literature. Electronic Journal of Business Research Methods 20, 3 (2022), 142–152

  12. [21]

    Martinez, Iain J

    Jay DeYoung, Stephanie C. Martinez, Iain J. Marshall, and Byron C. Wallace. 2024. Do Multi-Document Summarization Models Synthesize? Transactions of the Association for Computational Linguistics 12 (2024), 1043–1062. doi:10.1162/tacl_ a_00687

  13. [22]

    Caitlin Doogan and Wray Buntine. 2021. Topic Model or Topic Twaddle? Re- evaluating Semantic Interpretability Measures. In Proceedings of the 2021 Confer- ence of the North American Chapter of the Association for Computational Linguis- tics: Human Language Technologies. Associ...

  14. [23]

    Eamon Duede, William Dolan, André Bauer, Ian Foster, and Karim Lakhani

  15. [24]

    Elicit. 2023. Elicit: The AI Research Assistant. https://elicit.com

  16. [25]

    Shai Erera, Michal Shmueli-Scheuer, Guy Feigenblat, Ora Peled Nakash, Odellia Boni, Haggai Roitman, Doron Cohen, Bar Weiner, Yosi Mass, Or Rivlin, Guy Lev, Achiya Jerbi, Jonathan Herzig, Yufang Hou, Charles Jochim, Martin Gleize, Francesca Bonin, Francesca Bonin, and David Kon...

  17. [26]

    Feuston and Jed R

    Jessica L. Feuston and Jed R. Brubaker. 2021. Putting Tools in Their Place: The Role of Time and Perspective in Human-AI Collaboration for Qualitative Analysis. Proc. ACM Hum.-Comput. Interact. 5, Cscw2, Article 469 (Oct. 2021), 25 pages. doi:10.1145/3479856

  18. [27]

    Raymond Fok, Hita Kambhamettu, Luca Soldaini, Jonathan Bragg, Kyle Lo, Marti Hearst, Andrew Head, and Daniel S Weld. 2023. Scim: Intelligent Skimming Support for Scientific Papers. In Proceedings of the 28th International Conference on Intelligent User Interfaces (Iui ’23) . A...

  19. [28]

    Raymond Fok, Nedim Lipka, Tong Sun, and Alexa F Siu. 2024. Marco: Supporting Business Document Workflows via Collection-Centric Information Foraging with Large Language Models. In Proceedings of the CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI ...

  20. [29]

    Siu, and Daniel S

    Raymond Fok, Alexa F. Siu, and Daniel S. Weld. 2025. Toward Living Narrative Reviews: An Empirical Study of the Processes and Challenges in Updating Survey Articles in Computing Research. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems (Yokohama...

  21. [30]

    Jie Gao, Kenny Tsu Wei Choo, Junming Cao, Roy Ka-Wei Lee, and Simon Perrault

  22. [31]

    Charlie George and Andreas Stuhlmueller. 2023. Factored Verification: Detecting and Reducing Hallucination in Summaries of Academic Papers. In Proceedings of the Second Workshop on Information Extraction from Scientific Publications . Association for Computational Linguistics,...

  23. [32]

    Darcy Haag Granello. 2001. Promoting cognitive complexity in graduate writ- ten work: Using Bloom’s taxonomy as a pedagogical tool to improve literature reviews. Counselor Education and Supervision 40, 4 (2001), 292–307

  24. [33]

    Griffiths and Mark Steyvers

    Thomas L. Griffiths and Mark Steyvers. 2004. Finding scientific topics.Proceedings of the National Academy of Sciences 101, suppl_1 (2004), 5228–5235. doi:10.1073/ pnas.0307752101

  25. [34]

    2008–2025

    Grobid. 2008–2025. Grobid. https://github.com/kermitt2/grobid. swh:1:dir:dab86b296e3c3216e2241968f0d63b68e8209d3c

  26. [35]

    Han, Junhang Yu, Raphael Bournet, Alexandre Ciorascu, Wendy E

    Han L. Han, Junhang Yu, Raphael Bournet, Alexandre Ciorascu, Wendy E. Mackay, and Michel Beaudouin-Lafon. 2022. Passages: Interacting with Text Across Documents. In Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems (New Orleans, LA, USA)(CHI ’22). As...

  27. [36]

    Hayato Hashimoto, Kazutoshi Shinoda, Hikaru Yokono, and Akiko Aizawa. 2017. Automatic Generation of Review Matrices as Multi-document Summarization of Scientific Papers. In Birndl @ Sigir. Association for Computing Machinery, New York, NY, USA, 69–82

  28. [37]

    Weld, and Marti A

    Andrew Head, Kyle Lo, Dongyeop Kang, Raymond Fok, Sam Skjonsberg, Daniel S. Weld, and Marti A. Hearst. 2021. Augmenting Scientific Papers with Just-in-Time, Position-Sensitive Definitions of Terms and Symbols. In Proceedings of the 2021 CHI Conference on Human Factors in Compu...

  29. [39]

    Eric Horvitz. 1999. Principles of mixed-initiative user interfaces. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems (Pittsburgh, Pennsylvania, USA) (CHI ’99). Association for Computing Machinery, New York, NY, USA, 159–166. doi:10.1145/302979.303030

  30. [40]

    Alexander Miserlis Hoyle, Pranav Goel, Rupak Sarkar, and Philip Resnik. 2022. Are Neural Topic Models Broken?. InFindings of the Association for Computational Linguistics: EMNLP 2022. Association for Computational Linguistics, Abu Dhabi, United Arab Emirates, 5321–5344. doi:10...

  31. [41]

    Chao-Chun Hsu, Erin Bransom, Jenna Sparks, Bailey Kuehl, Chenhao Tan, David Wadden, Lucy Wang, and Aakanksha Naik. 2024. CHIME: LLM-Assisted Hierar- chical Organization of Scientific Studies for Literature Review Support. InFindings Facets, Taxonomies, and Syntheses: Navigatin...

  32. [42]

    Brubaker

    Jialun Aaron Jiang, Kandrea Wade, Casey Fiesler, and Jed R. Brubaker. 2021. Sup- porting Serendipity: Opportunities and Challenges for Human-AI Collaboration in Qualitative Analysis. Proc. ACM Hum.-Comput. Interact. 5, Cscw1, Article 94 (April 2021), 23 pages. doi:10.1145/3449168

  33. [43]

    Arif E Jinha. 2010. Article 50 million: an estimate of the number of scholarly articles in existence. Learned publishing 23, 3 (2010), 258–263

  34. [44]

    Hyeonsu Kang, Joseph Chee Chang, Yongsung Kim, and Aniket Kittur. 2022. Threddy: An Interactive System for Personalized Thread-based Exploration and Organization of Scientific Literature. In Proceedings of the 35th Annual ACM Symposium on User Interface Software and Technology...

  35. [46]

    Tetsu Kasanishi, Masaru Isonuma, Junichiro Mori, and Ichiro Sakata. 2023. SciRe- viewGen: A Large-scale Dataset for Automatic Literature Review Generation. In Findings of the Association for Computational Linguistics: ACL 2023 , Anna Rogers, Jordan Boyd-Graber, and Naoaki Okaz...

  36. [47]

    Hanan Khalil, Lotfi Tamara, Gabriel Rada, and Elie A. Akl. 2022. Challenges of evidence synthesis during the 2020 COVID pandemic: a scoping review. Journal of Clinical Epidemiology 142 (2022), 10–18. doi:10.1016/j.jclinepi.2021.10.017

  37. [48]

    Zhang, and Joseph Chee Chang

    Tae Soo Kim, Matt Latzke, Jonathan Bragg, Amy X. Zhang, and Joseph Chee Chang. 2023. Papeos: Augmenting Research Papers with Talk Videos. In Proceed- ings of the 36th Annual ACM Symposium on User Interface Software and Technology (San Francisco, CA, USA) (UIST ’23). Associatio...

  38. [49]

    Knight, Max L

    Ian A. Knight, Max L. Wilson, David F. Brailsford, and Natasa Milic-Frayling. 2019. Enslaved to the Trapped Data: A Cognitive Work Analysis of Medical Systematic Reviews. In Proceedings of the 2019 Conference on Human Information Interaction and Retrieval (Glasgow, Scotland UK...

  39. [50]

    Lam, Janice Teoh, James A

    Michelle S. Lam, Janice Teoh, James A. Landay, Jeffrey Heer, and Michael S. Bernstein. 2024. Concept Induction: Analyzing Unstructured Text with High- Level Concepts Using LLooM. In Proceedings of the CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CH...

  40. [51]

    Alghamdi, Tal August, Avinash Bhat, Madiha Zahrah Choksi, Senjuti Dutta, Jin L.C

    Mina Lee, Katy Ilonka Gero, John Joon Young Chung, Simon Buckingham Shum, Vipul Raheja, Hua Shen, Subhashini Venugopalan, Thiemo Wambsganss, David Zhou, Emad A. Alghamdi, Tal August, Avinash Bhat, Madiha Zahrah Choksi, Senjuti Dutta, Jin L.C. Guo, Md Naimul Hoque, Yewon Kim, S...

  41. [52]

    Yoonjoo Lee, Hyeonsu B Kang, Matt Latzke, Juho Kim, Jonathan Bragg, Joseph Chee Chang, and Pao Siangliulue. 2024. PaperWeaver: Enriching Topical Paper Alerts by Contextualizing Recommended Papers with User-collected Pa- pers. In Proceedings of the 2024 CHI Conference on Human ...

  42. [53]

    Zhehui Liao, Maria Antoniak, Inyoung Cheong, Evie Yu-Yen Cheng, Ai-Heng Lee, Kyle Lo, Joseph Chee Chang, and Amy X. Zhang. 2024. LLMs as Re- search Tools: A Large Scale Survey of Researchers’ Usage and Perceptions. arXiv:2411.05025 [cs.CL]

  43. [54]

    Michael Xieyang Liu, Jane Hsieh, Nathan Hahn, Angelina Zhou, Emily Deng, Shaun Burley, Cynthia Taylor, Aniket Kittur, and Brad A. Myers. 2019. Unakite: Scaffolding Developers’ Decision-Making Using the Web. In Proceedings of the 32nd Annual ACM Symposium on User Interface Soft...

  44. [55]

    Michael Xieyang Liu, Aniket Kittur, and Brad A. Myers. 2022. Crystalline: Low- ering the Cost for Developers to Collect and Organize Information for Decision Making. In CHI Conference on Human Factors in Computing Systems . Acm, New Orleans LA USA, 1–16. doi:10.1145/3491102.3501968

  45. [56]

    Yao Lu, Yue Dong, and Laurent Charlin. 2020. Multi-XScience: A Large-scale Dataset for Extreme Multi-document Summarization of Scientific Articles. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). Association for Computational ...

  46. [57]

    Hearst, and Dongyeop Kang

    Anna Martin-Boyle, Aahan Tyagi, Marti A. Hearst, and Dongyeop Kang. 2024. Shallow Synthesis of Knowledge in GPT-Generated Texts: A Case Study in Auto- matic Related Work Composition. arXiv:2402.12255 [cs.CL]

  47. [58]

    Matthew Michelson and Katja Reuter. 2019. The significant cost of systematic reviews and meta-analyses: A call for greater involvement of machine learning to assess the promise of clinical trials. Contemporary Clinical Trials Communications 16 (2019), 100443. doi:10.1016/j.con...

  48. [59]

    Eliot Moss. 2021. Connected Papers. https://www.connectedpapers.com/ Ac- cessed: 2025-04-04

  49. [60]

    Sonia Murthy, Kyle Lo, Daniel King, Chandra Bhagavatula, Bailey Kuehl, Sophie Johnson, Jonathan Borchardt, Daniel Weld, Tom Hope, and Doug Downey. 2022. ACCoRD: A Multi-Document Approach to Generating Diverse Descriptions of Sci- entific Concepts. InProceedings of the 2022 Con...

  50. [61]

    Arpit Narechania, Alireza Karduni, Ryan Wesslen, and Emily Wall. 2022. VI- TALITY: Promoting Serendipitous Discovery of Academic Literature with Trans- formers & Visual Analytics. IEEE Transactions on Visualization and Computer Graphics 28, 1 (Jan. 2022), 486–496. doi:10.1109/...

  51. [62]

    Benjamin Newman, Yoonjoo Lee, Aakanksha Naik, Pao Siangliulue, Raymond Fok, Juho Kim, Daniel S Weld, Joseph Chee Chang, and Kyle Lo. 2024. Arx- ivDIGESTables: Synthesizing Scientific Literature into Tables using Language Models. In Proceedings of the 2024 Conference on Empiric...

  52. [63]

    Zhang, Jonathan Bragg, and Joseph Chee Chang

    Srishti Palani, Aakanksha Naik, Doug Downey, Amy X. Zhang, Jonathan Bragg, and Joseph Chee Chang. 2023. Relatedly: Scaffolding Literature Reviews with Existing Related Work Sections. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems . Acm, Hamburg...

  53. [64]

    Rock Yuren Pang, Hope Schroeder, Kynnedy Simone Smith, Solon Barocas, Ziang Xiao, Emily Tseng, and Danielle Bragg. 2025. Understanding the LLM-ification of CHI: Unpacking the Impact of LLMs at CHI through a Systematic Literature Review. arXiv:2501.12557 [cs.HC]

  54. [65]

    Chau Minh Pham, Alexander Hoyle, Simeng Sun, Philip Resnik, and Mo- hit Iyyer. 2024. TopicGPT: A Prompt-based Topic Modeling Framework. arXiv:2311.01449 [cs.CL]

  55. [66]

    Peter Pirolli and Stuart Card. 1999. Information foraging. Psychological review 106, 4 (1999), 643

  56. [67]

    Peter Pirolli and Stuart Card. 2005. The sensemaking process and leverage points for analyst technology as identified through cognitive task analysis. Proceedings of International Conference on Intelligence Analysis 5 (2005), 2–4

  57. [68]

    the answer

    Riaz Qureshi, Daniel Shaughnessy, Kayden AR Gill, Karen A Robinson, Tianjing Li, and Eitan Agai. 2023. Are ChatGPT and large language models “the answer” to bringing us closer to systematic review automation? Systematic Reviews 12, 1 (2023), 72

  58. [69]

    Zhang, and Daniel S Weld

    Napol Rachatasumrit, Jonathan Bragg, Amy X. Zhang, and Daniel S Weld. 2022. CiteRead: Integrating Localized Citation Contexts into Scientific Paper Reading. In 27th International Conference on Intelligent User Interfaces (Iui ’22) . Association for Computing Machinery, New Yor...

  59. [70]

    Tim Rietz and Alexander Maedche. 2021. Cody: An AI-Based System to Semi- Automate Coding for Qualitative Research. In Proceedings of the 2021 CHI Con- ference on Human Factors in Computing Systems (Yokohama, Japan) (CHI ’21). Association for Computing Machinery, New York, NY, ...

  60. [71]

    Russell, Mark J

    Daniel M. Russell, Mark J. Stefik, Peter Pirolli, and Stuart K. Card. 1993. The cost structure of sensemaking. In Proceedings of the INTERACT ’93 and CHI ’93 Conference on Human Factors in Computing Systems (CHI ’93) . Association for Computing Machinery, New York, NY, USA, 26...

  61. [72]

    Kaveh G Shojania, Margaret Sampson, Mohammed T Ansari, Jun Ji, Steve Doucette, and David Moher. 2007. How quickly do systematic reviews go out of date? A survival analysis. Annals of internal medicine 147, 4 (2007), 224–233

  62. [73]

    Aviv Slobodkin, Eran Hirsch, Arie Cattan, Tal Schuster, and Ido Dagan. 2024. Attribute First, then Generate: Locally-attributable Grounded Text Generation. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Assoc...

  63. [74]

    Hannah Snyder. 2019. Literature review as a research methodology: An overview and guidelines. Journal of Business Research 104 (2019), 333–339. doi:10.1016/j. jbusres.2019.07.039

  64. [75]

    Michael Spenke, Christian Beilken, and Thomas Berlage. 1996. FOCUS: the interactive table for product comparison and selection. In Proceedings of the 9th annual ACM symposium on User interface software and technology (UIST ’96) . Association for Computing Machinery, New York, ...

  65. [76]

    Sangho Suh, Meng Chen, Bryan Min, Toby Jia-Jun Li, and Haijun Xia. 2024. Luminate: Structured Generation and Exploration of Design Space with Large Language Models for Human-AI Co-Creation. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems (Honolu...

  66. [77]

    Teo Susnjak, Peter Hwang, Napoleon Reyes, Andre L. C. Barczak, Timothy McIn- tosh, and Surangika Ranathunga. 2025. Automating Research Synthesis with Domain-Specific Large Language Model Fine-Tuning. ACM Trans. Knowl. Discov. Data 19, 3, Article 68 (March 2025), 39 pages. doi:...

  67. [78]

    James Thomas, Anna Noel-Storr, Iain Marshall, Byron Wallace, Steven McDonald, Chris Mavergames, Paul Glasziou, Ian Shemilt, Anneliese Synnot, Tari Turner, Julian Elliott, Thomas Agoritsas, John Hilton, Caroline Perron, Elie Akl, Rebecca Hodder, Charlotte Pestridge, Lauren Albr...

  68. [79]

    David Wadden, Kyle Lo, Bailey Kuehl, Arman Cohan, Iz Beltagy, Lucy Lu Wang, and Hannaneh Hajishirzi. 2022. SciFact-Open: Towards open-domain scientific claim verification. In Findings of the Association for Computational Linguistics: EMNLP 2022, Yoav Goldberg, Zornitsa Kozarev...

  69. [80]

    David Wadden, Kyle Lo, Lucy Lu Wang, Arman Cohan, Iz Beltagy, and Hannaneh Hajishirzi. 2022. MultiVerS: Improving scientific claim verification with weak supervision and full-document context. In Findings of the Association for Compu- tational Linguistics: NAACL 2022, Marine C...

  70. [81]

    Han Wang, Nirmalendu Prakash, Nguyen Khoi Hoang, Ming Shan Hee, Usman Naseem, and Roy Ka-Wei Lee. 2023. Prompting Large Language Models for Topic Modeling. In 2023 IEEE International Conference on Big Data (BigData) . IEEE Computer Society, Los Alamitos, CA, USA, 1236–1241. do...

  71. [82]

    Huey, Rui Sheng, Saurabh Mehta, and Fei Wang

    Xingbo Wang, Samantha L. Huey, Rui Sheng, Saurabh Mehta, and Fei Wang. 2024. SciDaSynth: Interactive Structured Knowledge Extraction and Synthesis from Scientific Literature with Large Language Model. arXiv:2404.13765 [cs.HC]

  72. [83]

    Zihan Wang, Jingbo Shang, and Ruiqi Zhong. 2023. Goal-Driven Explainable Clustering via Language Descriptions. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, Singapore, 10626–10649. doi:10.1...

  73. [84]

    Xianjun Yang, Kaiqiang Song, Sangwoo Cho, Xiaoyang Wang, Xiaoman Pan, Linda Petzold, and Dong Yu. 2023. OASum: Large-Scale Open Domain Aspect- based Summarization. In Findings of the Association for Computational Linguistics: ACL 2023. Association for Computational Linguistics...

  74. [85]

    Jiajie Zhang, Yushi Bai, Xin Lv, Wanjun Gu, Danqing Liu, Minhao Zou, Shulin Cao, Lei Hou, Yuxiao Dong, Ling Feng, and Juanzi Li. 2024. LongCite: Enabling LLMs to Generate Fine-grained Citations in Long-context QA. arXiv:2409.02897 [cs.CL]

  75. [86]

    Yuwei Zhang, Zihan Wang, and Jingbo Shang. 2023. ClusterLLM: Large Language Models as a Guide for Text Clustering. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, Singapore, 13903–13920. doi:...

  76. [87]

    Evaluation Metrics and Results

    Kun Zhu, Xiaocheng Feng, Xiachong Feng, Yingsheng Wu, and Bing Qin. 2023. Hierarchical Catalogue Generation for Literature Review: A Benchmark. In Find- ings of the Association for Computational Linguistics: EMNLP 2023 . Association for Computational Linguistics, Singapore, 67...

  77. [2001]

    Longman, New York

    A Taxonomy for Learning, Teaching, and Assessing: A Revision of Bloom’s Taxonomy of Educational Objectives. Longman, New York

  78. [2023]

    ACM Trans

    CoAIcoder: Examining the Effectiveness of AI-assisted Human-to-Human Collaboration in Qualitative Analysis. ACM Trans. Comput.-Hum. Interact. 31, 1, Article 6 (Nov. 2023), 38 pages. doi:10.1145/3617362

  79. [2024]

    arXiv:2405.15828 [cs.DL]

    Oil & Water? Diffusion of AI Within and Across Scientific Fields. arXiv:2405.15828 [cs.DL]

  80. [4401]

    doi:10.18653/v1/2023.findings-acl.268

  81. [8074]

    doi:10.18653/v1/2020.emnlp-main.648

Pith tools

Reviewed August 16, 2026 · model on record in the stance chip above.