{"total":23,"items":[{"citing_arxiv_id":"2607.08348","ref_index":35,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"Works on My QPU: Reproducibility in Quantum Computing Research","primary_cat":"quant-ph","submitted_at":"2026-07-09T10:53:02+00:00","verdict":"CONDITIONAL","verdict_confidence":"MODERATE","novelty_score":6.0,"formal_verification":"none","one_line_summary":"A large-scale audit finds that only ~25% of quantum-computing papers provide code artifacts and ~65% of those artifacts fail to execute in a clean environment.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.02416","ref_index":32,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"The Future of NLP may not be at NLP Conferences: Scholarly Migration Patterns in Natural Language Processing","primary_cat":"cs.CL","submitted_at":"2026-07-02T16:47:14+00:00","verdict":"CONDITIONAL","verdict_confidence":"HIGH","novelty_score":7.0,"formal_verification":"none","one_line_summary":"Established and new NLP authors are migrating from *ACL venues to general ML venues, with ML venues conferring a large content-matched citation premium.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.28120","ref_index":13,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"The Reciprocal Impact of Science and Software: A Cross-Corpus Analysis of How Research Shapes Software and Software Enables Research","primary_cat":"cs.DL","submitted_at":"2026-06-26T14:20:33+00:00","verdict":null,"verdict_confidence":null,"novelty_score":null,"formal_verification":null,"one_line_summary":null,"context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.18874","ref_index":29,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Externalizing Research Synthesis and Validation in AI Scientists through a Research Harness","primary_cat":"cs.AI","submitted_at":"2026-06-17T09:52:14+00:00","verdict":"CONDITIONAL","verdict_confidence":"MODERATE","novelty_score":6.0,"formal_verification":"none","one_line_summary":"Xcientist is a research harness that externalizes an AI scientist's literature grounding, idea evolution, experiments, and repairs into auditable artifacts, demonstrated on memory, traffic forecasting, and PDE-solving tasks.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.05443","ref_index":17,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"MIRAI: Prediction and Generation of High-Impact Academic Research","primary_cat":"cs.DL","submitted_at":"2026-06-03T21:06:01+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":3.0,"formal_verification":"none","one_line_summary":"MIRAI predicts 5-year PageRank and citation impact from paper title/abstract/date with Spearman's ρ 0.47/0.62, and generates ideas judged 4:3 more impactful by LLM.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.29522","ref_index":50,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"DeepSurvey: Enhancing Analytical Depth and Citation Reliability in Automated Survey Generation","primary_cat":"cs.AI","submitted_at":"2026-05-28T07:40:10+00:00","verdict":"UNVERDICTED","verdict_confidence":"UNKNOWN","novelty_score":6.0,"formal_verification":"none","one_line_summary":"DeepSurvey introduces an agentic system for automated survey generation that improves depth through full-text keynotes, cross-paper clustering, and code analysis, while boosting citation reliability via graph expansion, hybrid filtering, and evidence-constrained assignment, with reported gains over ","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.29234","ref_index":1,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Rethinking Literature Search Evaluation: Deep Research Helps, and Human Citation Lists Are Not a Ground Truth","primary_cat":"cs.AI","submitted_at":"2026-05-28T01:50:52+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"Deep bibliography expansion in literature search achieves high recall while human citations are found to have only 51% moderate relevance compared to 86-88% for AI methods.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.28187","ref_index":21,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Whose Name Comes Up? III: Persona Prompting Effects in LLM-Based Scholar Recommendation","primary_cat":"cs.IR","submitted_at":"2026-05-27T09:09:30+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"Audits of 43 LLMs show that varying persona prompts (language, location, role-and-task) and context affects technical quality and social representativeness of scholar recommendations, with location impacting diversity and factuality.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.23694","ref_index":27,"ref_count":2,"confidence":0.9,"is_internal_anchor":false,"paper_title":"ChartFI: Benchmarking Faithfulness and Insightfulness of Chart Descriptions from Multimodal Large Language Models","primary_cat":"cs.CL","submitted_at":"2026-05-22T14:49:48+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"ChartFI-Bench supplies 896 chart-description pairs from visually complex charts and defines four metrics (Faithfulness, Coverage, Informativeness, Acuity) aligned to four quality dimensions to evaluate MLLM-generated descriptions.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.20833","ref_index":22,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"MemGym: a Long-Horizon Memory Environment for LLM Agents","primary_cat":"cs.CL","submitted_at":"2026-05-20T07:25:33+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":7.0,"formal_verification":"none","one_line_summary":"MemGym unifies agent gyms into a memory benchmark with isolated scoring across tool-use, research, coding, and computer-use regimes plus a lightweight reward model for tractable coding evaluation.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.15474","ref_index":23,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Jobs' AI Exposure Should Be Measured from Evidence, Not Model Priors","primary_cat":"cs.IR","submitted_at":"2026-05-14T23:29:42+00:00","verdict":"CONDITIONAL","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"The authors propose a retrieval-augmented framework that grounds AI exposure labels for 18,796 O*NET occupation-task pairs in retrieved news and academic abstracts, outperforming zero-shot prompting in 72% of disagreements and aligning better with observed real-world usage.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.14790","ref_index":15,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Graphs of Research: Citation Evolution Graphs as Supervision for Research Idea Generation","primary_cat":"cs.CL","submitted_at":"2026-05-14T12:57:56+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":7.0,"formal_verification":"none","one_line_summary":"GoR extracts citation DAGs using position, frequency, predecessor links and time, then fine-tunes Qwen2.5-7B on 498 seed papers to generate ideas, claiming SOTA over gpt-4o baselines via LLM judges.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.11258","ref_index":23,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Unlocking LLM Creativity in Science through Analogical Reasoning","primary_cat":"cs.AI","submitted_at":"2026-05-11T21:35:44+00:00","verdict":"CONDITIONAL","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"Analogical reasoning increases LLM solution diversity by 90-173% and novelty rate to over 50%, delivering up to 13-fold gains on biomedical tasks including perturbation prediction and cell communication.","context_count":1,"top_context_role":"other","top_context_polarity":"unclear","context_text":"Foster, Cyril Zhang, and Aleksandrs Slivkins. Can large language models explore in-context?, 2024. URL https://arxiv.org/abs/2403. 15371. [22] Akarsh Kumar, Ryan Bahlous-Boldi, Prafull Sharma, Phillip Isola, Sebastian Risi, Yujin Tang, and David Ha. Digital red queen: Adversarial program evolution in core war with llms, 2026. URLhttps://arxiv.org/abs/2601.03335. [23] Jack Lanchantin, Angelica Chen, Shehzaad Dhuliawala, Ping Yu, Jason Weston, Sainbayar Sukhbaatar, and Ilia Kulikov. Diverse preference optimization, 2025. URL https://arxiv. org/abs/2501.18101. [24] Chris Lu, Cong Lu, Robert Tjarko Lange, Jakob Foerster, Jeff Clune, and David Ha. The ai scientist: Towards fully automated open-ended scientific discovery, 2024."},{"citing_arxiv_id":"2604.23430","ref_index":23,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Automating Categorization of Scientific Texts with In-Context Learning and Prompt-Chaining in Large Language Models","primary_cat":"cs.IR","submitted_at":"2026-04-25T19:52:21+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":5.0,"formal_verification":"none","one_line_summary":"Prompt chaining with off-the-shelf LLMs outperforms in-context learning and BERT for 1st- and 2nd-level classification on the ORKG taxonomy using the FORC dataset, but struggles at the 3rd level.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2604.18874","ref_index":51,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"How Adversarial Environments Mislead Agentic AI?","primary_cat":"cs.AI","submitted_at":"2026-04-20T21:53:39+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"Adversarial compromise of tool outputs misleads agentic AI via breadth and depth attacks, revealing that epistemic and navigational robustness are distinct and often trade off against each other.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2604.12498","ref_index":19,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Lit2Vec: A Reproducible Workflow for Building a Legally Screened Chemistry Corpus from S2ORC for Downstream Retrieval and Text Mining","primary_cat":"cs.DB","submitted_at":"2026-04-14T09:26:11+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":5.0,"formal_verification":"none","one_line_summary":"Lit2Vec delivers a documented, reproducible pipeline that extracts and annotates a large licensed chemistry paper corpus from S2ORC with paragraph embeddings and subfield labels.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2604.08898","ref_index":24,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Omakase: proactive assistance with actionable suggestions for evolving scientific research projects","primary_cat":"cs.HC","submitted_at":"2026-04-10T03:05:02+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":4.0,"formal_verification":"none","one_line_summary":"Omakase monitors project documents to infer timely queries and distills research reports into actionable suggestions that users rated significantly more useful than raw reports.","context_count":1,"top_context_role":"dataset","top_context_polarity":"use_dataset","context_text":"1 Technology Probe: Living Research Interests Document The first iteration of our system is a technology probe consisted of three parts: (1) a shared Google Doc as aresearch interests document that can be freely edited by the participants; (2) apaper recom- mendations pipelinethat periodically read the research interests document, retrieved papers from Semantic Scholar API [24, 50] and an open-sourced deep research system [ 49], and generated brief explanations of their relevance to the user interests; and (3)recom- mendation emailsdelivered twice a week containing the paper rec- ommendations, from which users could add the paper back to their research interests document, asked an AI assistant about them, or 3 Siangliulue et al."},{"citing_arxiv_id":"2604.07530","ref_index":11,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"The Shrinking Lifespan of LLMs in Science","primary_cat":"cs.DL","submitted_at":"2026-04-08T19:12:09+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":7.0,"formal_verification":"none","one_line_summary":"LLM adoption in science follows a compressing inverted-U trajectory where release year predicts time-to-peak and lifespan better than model attributes.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2603.04459","ref_index":50,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Benchmark of Benchmarks: Unpacking Influence and Code Repository Quality in LLM Safety Benchmarks","primary_cat":"cs.CR","submitted_at":"2026-03-03T09:10:45+00:00","verdict":"CONDITIONAL","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"Only 39% of LLM safety benchmark repositories run without modification, 6% include ethical warnings, and adoption tracks author prominence and runnability rather than code quality metrics.","context_count":1,"top_context_role":"background","top_context_polarity":"background","context_text":"Following rules of thumb in Statistics [42, 66, 117], we set a minimum sample size of 25 per group, requiring us to analyze all benchmark papers collectively (see Table 1) rather than segmenting them by topic. We apply the Mann- Whitney U (M-W U) test [70, 76] to assess statistical signifi- cance [40,84,124] and use Cliff's delta [41,60,71] to measure practical significance (effect size) [50, 64, 86, 92] in distribu- tion differences across the five metrics. 7 The methods used are robust to sample size, accommodate unequal group sizes, and require no assumptions on data distribution. The detailed results of the M-W U test and Cliff's delta are presented in Table 3. While non-benchmark papers usually have higher average values on influence metrics (descriptive"},{"citing_arxiv_id":"2510.07037","ref_index":16,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Beyond Monolingual Assumptions: A Survey of Code-Switched NLP in the Era of Large Language Models across Modalities","primary_cat":"cs.CL","submitted_at":"2025-10-08T14:04:14+00:00","verdict":"ACCEPT","verdict_confidence":"LOW","novelty_score":7.0,"formal_verification":"none","one_line_summary":"A comprehensive survey of code-switched NLP research with LLMs across modalities, covering 327 studies, 15+ tasks, 30+ datasets, and 80+ languages while outlining challenges and a future roadmap.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2510.00361","ref_index":32,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Attribution Gradients: Incrementally Unfolding Citations for Critical Examination of Attributed AI Answers","primary_cat":"cs.HC","submitted_at":"2025-10-01T00:07:28+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"Attribution gradients consolidate citation evidence and enable incremental unfolding of secondary sources, leading to deeper engagement in a lab study of critical reading tasks for AI answers.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2508.20765","ref_index":42,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Looking Beyond the Obvious: A Survey on Abstract Concept Recognition for Video Understanding","primary_cat":"cs.CV","submitted_at":"2025-08-28T13:19:49+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":3.0,"formal_verification":"none","one_line_summary":"A literature survey on abstract concept recognition in videos that catalogs prior tasks and datasets while advocating for foundation models and reuse of decades of community experience.","context_count":1,"top_context_role":"background","top_context_polarity":"background","context_text":"various domains, including detection, segmenta- tion, and generation. As explained in this section, filtering out literature focusing specifically on abstract concepts requires developing a detailed methodology resulting in this pipeline. Figure 4 highlights the automatic and manual steps for the literature survey pipeline. We use the data indexed until September 2024 from Semantic Scholar[42] as it contains a large corpus of papers from vari- ous A∗, A, and B rated conferences according to [43]. Semantic Scholar, besides offering informa- tion about the venue, authors, and the title of the manuscript, also includes abstracts that aid in the automatic filtering process. We filter out papers that do not have the word 'video' in the title or"},{"citing_arxiv_id":"2409.14634","ref_index":37,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Human-LLM Compound System for Scientific Ideation through Facet Recombination and Novelty Evaluation","primary_cat":"cs.HC","submitted_at":"2024-09-23T00:09:34+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":7.0,"formal_verification":"none","one_line_summary":"Scideator enables facet-based scientific ideation through LLM-driven extraction, human-guided recombination, analogous retrieval, and facet-grounded novelty verification, showing significantly higher creativity support than a baseline LLM in a user study with CS researchers.","context_count":1,"top_context_role":"background","top_context_polarity":"background","context_text":"Shen, Amanpreet Singh, Luca Soldaini, Shivashankar Subramanian, A. Tanaka, Alex D Wade, Linda M. Wagner, Lucy Lu Wang, Christopher Wilhelm, Caroline Wu, Jiangjiang Yang, Angele Zamarron, Madeleine van Zuylen, and Daniel S. Weld. 2023. The Semantic Scholar Open Data Platform. ArXiv abs/2301.10140 (2023). https://api.semanticscholar.org/CorpusID:256194545 [37] Dan Lahav, Jon Saad Falcon, Bailey Kuehl, Sophie Johnson, Sravanthi Parasa, Noam Shomron, Duen Horng Chau, Diyi Yang, Eric Horvitz, Daniel S Weld, et al. 2022. A search engine for discovery of scientific challenges and directions. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 36. 11982-11990. [38] Pier Luca Lanzi and Daniele Loiacono."}],"limit":50,"offset":0}