REVIEW 4 major objections 1 cited by
One continuous agent trajectory with shared memory produces consistent hierarchical docs for whole codebases.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-13 07:54 UTC pith:WTC3UJSY
load-bearing objection Solid systems paper: single-agent long-horizon docs with dependency order + hierarchical memory actually beats the usual multi-agent baselines on consistency and regeneration, with real ablations and code. the 4 major comments →
Remember Your Trace: Memory-Guided Long-Horizon Agentic Framework for Consistent and Hierarchical Repository-Level Code Documentation
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Treating repository documentation as one long-horizon agentic process—with dependency-aware traversal and a shared RepoMemory that accumulates retrievals, drafts, and verified documents via READ, WRITE, VERIFY, and FINISH—yields hierarchical documentation (component, module, repository) that is more complete, truthful, helpful, and sufficient for code regeneration than systems that document components independently.
What carries the argument
RepoMemory plus dependency-aware traversal: a shared store of prior work traces that the agent accesses through adaptive READ, hierarchical WRITE, self- and cross-document VERIFY (NLI conflict check plus self-scores), and FINISH commit, ordered so each unit is documented only after its dependencies and children.
Load-bearing premise
The multi-criteria evaluation stack—LLM judges, NLI conflict scores, and regenerating code from documentation alone—faithfully measures practical documentation quality and that strong results on twenty selected Python repositories generalize.
What would settle it
On a held-out set of large multi-language repositories, measure whether MemDocAgent still eliminates redundant retrieval, reduces cross-document inconsistency by a large margin, and improves Pass@k/CodeBLEU when models regenerate functions from its documentation alone, relative to independent per-component baselines under the same judges.
If this is right
- Hierarchical docs from one trajectory can serve as the shared context both developers and coding agents use for navigation and change.
- Eliminating repeated source-file retrieval and verifying against committed memory reduces contradiction rates across documents.
- Information-sufficiency via code regeneration becomes a practical yardstick for whether docs preserve implementable detail.
- Dependency-first ordering plus memory commit supports covering components, modules, and the whole repository without isolated pipelines.
Where Pith is reading between the lines
- The same READ/WRITE/VERIFY memory pattern could transfer to other long-horizon software tasks that must keep many sequential artifacts consistent, such as multi-file refactor plans or test suites.
- If information sufficiency tracks real usefulness, documentation systems may be scored less by surface style and more by whether agents can rebuild behavior from the text alone.
- Weak spots in LLM judges or NLI conflict detection would systematically bias any ranking that relies on them, so human-in-the-loop spot checks remain necessary for deployment.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes MemDocAgent, a long-horizon agentic framework for repository-level hierarchical code documentation. It documents every unit (component, module, repository) in one continuous trajectory ordered by Dependency-Aware Traversal Guiding, and reuses prior retrievals and outputs via RepoMemory through READ/WRITE/VERIFY/FINISH. Against open-source (Prompting, RepoAgent, DocAgent, CodeWiki) and closed-source (DeepWiki, Claude-Code) baselines on 20 stratified Python DevEval repositories, with Qwen3-Coder and GPT-5-mini backbones, it reports best completeness, truthfulness, and helpfulness (Table 2), strongest information sufficiency via code regeneration Pass@k/CodeBLEU (Figure 3), large reductions in redundant retrieval and cross-document inconsistency (Figure 1), plus ablations, efficiency, verification-attempt, and scalability analyses.
Significance. If the results hold under fairer controls, this is a solid systems contribution to automated software documentation and long-horizon coding agents: it reframes repository documentation as a single persistent trajectory rather than independent per-component jobs, ships a concrete memory protocol and dependency-aware order, and evaluates with a multi-criteria stack that includes a practically motivated information-sufficiency metric (564 DevEval tests). Strengths include dual backbones, open- and closed-source baselines, per-granularity breakdowns, ablations of think/memory/conflict-verify, efficiency measurements, and public code/data. The work is relevant to both human onboarding and agent tooling over large codebases.
major comments (4)
- §5.1, §6.1, Figure 3: Information sufficiency is a load-bearing pillar for the claim of “practical applicability,” yet regeneration scores can be driven by documentation format richness and hierarchical volume rather than memory-guided consistency. WRITE prompts and formats (Appendix A.2, Tables 4–12) mandate Args/Returns/Control-Flow/Mermaid/examples and multi-level synthesis; scores rise monotonically from (C) to (C+M) to (C+M+R), and generated docs beat short human Ground-truth docstrings. Without length-, section-density-, or format-matched controls against a non-memory baseline that uses the same hierarchical templates, the causal link from RepoMemory/VERIFY to regeneration gains remains under-identified.
- Table 3 ablations (w/o think, w/o memory system, w/o conflict verify) show large drops but do not hold output structure fixed: a stateless cache cleared per unit still uses the same WRITE formats and VERIFY rubric. A format-matched Prompting or DocAgent-style run with identical section schemas (and, ideally, matched token budgets) is needed to show that dependency-aware π and persistent RepoMemory—not template richness—drive completeness/truthfulness/helpfulness and Pass@k/CodeBLEU. As written, the ablations support “the full system helps” more than “memory-guided long-horizon is the mechanism.”
- Figure 1 and §1: The claimed ~75.5% relative cut in cross-document inconsistency and 0% redundant retrieval are central to the motivation, but the manuscript does not fully specify how inconsistency was measured for baselines that lack shared state (claim extraction, pairing policy, NLI thresholds, sampling). Appendix A.3 details VERIFY for MemDocAgent; the same protocol should be applied uniformly to baseline outputs and reported with inter-annotator or model agreement so the 13%→3.1% comparison is auditable.
- §5.3 / Appendix B.3: Evaluation is restricted to 20 Python repositories from DevEval. The central claim is framed as repository-level documentation for modern workflows; without at least a small multi-language or out-of-DevEval stress set (or a clear scope limitation in the abstract/conclusion), generalization beyond this stratified Python sample is overstated relative to the evidence.
Circularity Check
Empirical systems paper with independent external benchmarks; only minor self-citation of an evaluation framework, not a by-construction derivation.
specific steps
-
self citation load bearing
[§5.1 Truthfulness; Refs. [59], [60]]
"We adopt a fine-grained documentation evaluation framework that decomposes generated documentation and evaluates each segment using predefined evaluation criteria [59, 60] to compute these two scores."
Reference [59] (Referee) is prior work by overlapping authors (Bae, Lee, Choi, Lee et al.) used for the truthfulness pipeline. This is a minor methodological self-citation, not a uniqueness theorem or a prediction forced by a fitted parameter; truthfulness is only one of four reported criteria and still conditions claims on source code with an external judge model.
full rationale
MemDocAgent is an empirical agentic systems paper, not a first-principles derivation. The load-bearing claims are comparative performance on completeness, truthfulness, helpfulness, and information sufficiency against open- and closed-source baselines on 20 DevEval Python repositories. Those metrics are grounded outside the method’s free parameters: section/entity completeness via AST and pattern matching against source; truthfulness via claim-level checks against code (with Claude Haiku 4.5 as judge, deliberately different from generation backbones); helpfulness via LLM-as-judge; information sufficiency via code regeneration measured by unit-test Pass@k and CodeBLEU. Ablations (Table 3) and hierarchical variants (C / C+M / C+M+R) further separate components without redefining the targets as fitted constants. VERIFY’s self-scores and NLI checks against previously committed docs are part of the generation loop, not the external evaluation, so reporting lower cross-document inconsistency is optimization success rather than circular prediction. The only mild circularity-adjacent element is adopting the authors’ own Referee framework [59] alongside FactScore [60] for fine-grained truthfulness scoring—a non-load-bearing methodological self-citation that does not force the superiority claim across the other three criteria. No self-definitional equations, fitted-input-as-prediction, uniqueness theorems imported from the authors, or renamed known results appear in the derivation chain.
Axiom & Free-Parameter Ledger
free parameters (4)
- verify_threshold =
0.9
- max_revisions_per_unit =
2
- max_steps_per_subtask =
10
- NLI_entailment_threshold_tau_nli
axioms (5)
- domain assumption High-quality repository documentation must be hierarchical (component, module, repository) and mutually consistent across dependency-related units.
- domain assumption LLM-as-judge section scores and fine-grained claim consistency/relevance against code are valid proxies for helpfulness and truthfulness.
- ad hoc to paper If a model can regenerate a function body from signature plus documentation and pass unit tests, the documentation is information-sufficient for practical use.
- domain assumption A dependency graph from static relations (calls, inheritance, imports, containment) plus SCC condensation yields a valid documentation order.
- domain assumption Pretrained NLI models can detect cross-document factual conflicts among atomic claims about code behavior.
invented entities (2)
-
RepoMemory (hierarchical documentation + external search stores)
independent evidence
-
MemDocAgent action protocol (READ/WRITE/VERIFY/FINISH) with dependency-aware π
independent evidence
read the original abstract
Automated code documentation is essential for modern software development, providing the contextual grounding that both human developers and coding agents rely on to navigate large codebases. Existing repository-level approaches process components independently, causing redundant retrieval and conflicting descriptions across documents while producing outputs that lack hierarchical structure. Therefore, we propose MemDocAgent, a long-horizon agentic framework that generates documentation within a single, integrated context spanning the entire repository. It combines two components: (i) Dependency-Aware Traversal Guiding that predetermines a traversal order respecting dependency and granularity hierarchies; (ii) Memory-Guided Agentic Interaction, in which the agent interacts with RepoMemory, a shared memory accumulating prior work traces through read, write, and verify operations. Through an in-depth multi-criteria evaluation, MemDocAgent achieves the best performance over both open and closed-source baselines and demonstrates practical applicability in real software development workflows.
Figures
Forward citations
Cited by 1 Pith paper
-
Shared Organizational Memory for Enterprise Coding Agents: System Design and Deployment Snapshot
A production deployment snapshot shows hook-based automatic capture and LLM curation turning 900 agent learnings into 1,144 shared question-answer memories, with no evidence yet of retrieval or coding-task benefit.
Reference graph
Works this paper leans on
-
[1]
DocAgent: A multi-agent system for automated code documentation generation
Dayu Yang, Antoine Simoulin, Xin Qian, Xiaoyi Liu, Yuwei Cao, Zhaopu Teng, and Grey Yang. DocAgent: A multi-agent system for automated code documentation generation. In Pushkar Mishra, Smaranda Muresan, and Tao Yu, editors,Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (ACL) (Volume 3: System Demonstrations), pages...
2025
-
[2]
Anh Nguyen Hoang, Minh Le-Anh, Bach Le, and Nghi DQ Bui. Codewiki: Evaluating ai’s ability to generate holistic documentation for large-scale codebases.arXiv preprint arXiv:2510.24428, 2025
Pith/arXiv arXiv 2025
-
[3]
RepoAgent: An LLM-powered open-source framework for repository-level code documentation generation
Qinyu Luo, Yining Ye, Shihao Liang, Zhong Zhang, Yujia Qin, Yaxi Lu, Yesai Wu, Xin Cong, Yankai Lin, Yingli Zhang, Xiaoyin Che, Zhiyuan Liu, and Maosong Sun. RepoAgent: An LLM-powered open-source framework for repository-level code documentation generation. In Delia Irazu Hernandez Farias, Tom Hope, and Manling Li, editors,Proceedings of the 2024 Conferen...
2024
-
[4]
Precise documentation: The key to better software
David Lorge Parnas. Precise documentation: The key to better software. InThe Future of Soft- ware Engineering, 2010. URLhttps://api.semanticscholar.org/CorpusID:38934599
2010
-
[5]
Golara Garousi, Vahid Garousi-Yusifo˘glu, Guenther Ruhe, Junji Zhi, Mahmoud Moussavi, and Brian Smith. Usage and usefulness of technical software documentation: An industrial case study.Information and Software Technology, 57:664–682, 2015. ISSN 0950-5849. doi: https://doi.org/10.1016/j.infsof.2014.08.003. URL https://www.sciencedirect.com/ science/articl...
-
[6]
Ai-driven chatbot as a support tool for developers during the onboarding process
Lea Katalina Kivinen. Ai-driven chatbot as a support tool for developers during the onboarding process. 2023
2023
-
[7]
Software engineering (extended abstract) an unconsummated marriage
David Lorge Parnas. Software engineering (extended abstract) an unconsummated marriage. ACM SIGSOFT Software Engineering Notes, 22(6):1–3, 1997
1997
-
[8]
The relevance of software documentation, tools and technologies: a survey
Andrew Forward and Timothy C Lethbridge. The relevance of software documentation, tools and technologies: a survey. InProceedings of the 2002 ACM symposium on Document engineering, pages 26–33, 2002
2002
-
[9]
A study of the docu- mentation essential to software maintenance
Sergio Cozzetti B De Souza, Nicolas Anquetil, and Káthia M De Oliveira. A study of the docu- mentation essential to software maintenance. InProceedings of the 23rd annual international conference on Design of communication: documenting & designing for pervasive information, pages 68–75, 2005
2005
-
[10]
Software documentation issues unveiled.2019 IEEE/ACM 41st International Conference on Software Engineering (ICSE), pages 1199–1210,
Emad Aghajani, Csaba Nagy, Olga Lucero Vega-Márquez, Mario Linares-Vásquez, Laura Moreno, Gabriele Bavota, and Michele Lanza. Software documentation issues unveiled.2019 IEEE/ACM 41st International Conference on Software Engineering (ICSE), pages 1199–1210,
2019
-
[11]
URLhttps://api.semanticscholar.org/CorpusID:174800564
-
[12]
Cost, benefits and quality of software development documentation: A systematic mapping
Junji Zhi, Vahid Garousi-Yusifo˘glu, Bo Sun, Golara Garousi, Shawn Shahnewaz, and Guenther Ruhe. Cost, benefits and quality of software development documentation: A systematic mapping. Journal of Systems and Software, 99:175–198, 2015
2015
-
[13]
Mea- suring program comprehension: A large-scale field study with professionals.IEEE Transactions on Software Engineering, 44(10):951–976, 2017
Xin Xia, Lingfeng Bao, David Lo, Zhenchang Xing, Ahmed E Hassan, and Shanping Li. Mea- suring program comprehension: A large-scale field study with professionals.IEEE Transactions on Software Engineering, 44(10):951–976, 2017
2017
-
[14]
DeepWiki.https://deepwiki.com/, 2025
Cognition AI. DeepWiki.https://deepwiki.com/, 2025
2025
-
[15]
Claude Code.https://www.anthropic.com/claude-code, 2025
Anthropic. Claude Code.https://www.anthropic.com/claude-code, 2025
2025
-
[16]
Evalu- ating usage and quality of technical software documentation: an empirical study
Golara Garousi, Vahid Garousi, Mahmoud Moussavi, Guenther Ruhe, and Brian Smith. Evalu- ating usage and quality of technical software documentation: an empirical study. InProceedings of the 17th international conference on evaluation and assessment in software engineering, pages 24–35, 2013. 10
2013
-
[17]
Lutfi Eren Erdogan, Nicholas Lee, Sehoon Kim, Suhong Moon, Hiroki Furuta, Gopala Anu- manchipalli, Kurt Keutzer, and Amir Gholami. Plan-and-act: Improving planning of agents for long-horizon tasks.arXiv preprint arXiv:2503.09572, 2025
Pith/arXiv arXiv 2025
-
[18]
Swe-agent: Agent-computer interfaces enable automated software engineering
John Yang, Carlos E Jimenez, Alexander Wettig, Kilian Lieret, Shunyu Yao, Karthik Narasimhan, and Ofir Press. Swe-agent: Agent-computer interfaces enable automated software engineering. Advances in Neural Information Processing Systems, 37:50528–50652, 2024
2024
-
[19]
Xingyao Wang, Boxuan Li, Yufan Song, Frank F Xu, Xiangru Tang, Mingchen Zhuge, Jiayi Pan, Yueqi Song, Bowen Li, Jaskirat Singh, et al. Openhands: An open platform for ai software developers as generalist agents.arXiv preprint arXiv:2407.16741, 2024
Pith/arXiv arXiv 2024
-
[20]
Codeagent: Enhancing code generation with tool-integrated agent systems for real-world repo-level coding challenges
Kechi Zhang, Jia Li, Ge Li, Xianjie Shi, and Zhi Jin. Codeagent: Enhancing code generation with tool-integrated agent systems for real-world repo-level coding challenges. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 13643–13658, 2024
2024
-
[21]
Huy Nhat Phan, Tien N Nguyen, Phong X Nguyen, and Nghi DQ Bui. Hyperagent: Generalist software engineering agents to solve coding tasks at scale.arXiv preprint arXiv:2409.16299, 2024
Pith/arXiv arXiv 2024
-
[22]
Repairagent: An autonomous, llm-based agent for program repair
Islem Bouzenia, Premkumar Devanbu, and Michael Pradel. Repairagent: An autonomous, llm-based agent for program repair. In2025 IEEE/ACM 47th International Conference on Software Engineering (ICSE), pages 2188–2200. IEEE, 2025
2025
-
[23]
Xiang Deng, Jeff Da, Edwin Pan, Yannis Yiming He, Charles Ide, Kanak Garg, Niklas Lauffer, Andrew Park, Nitin Pasari, Chetan Rane, et al. Swe-bench pro: Can ai agents solve long-horizon software engineering tasks?arXiv preprint arXiv:2509.16941, 2025
Pith/arXiv arXiv 2025
-
[24]
Haotian Luo, Huaisong Zhang, Xuelin Zhang, Haoyu Wang, Zeyu Qin, Wenjie Lu, Guozheng Ma, Haiying He, Yingsha Xie, Qiyang Zhou, et al. Ultrahorizon: Benchmarking agent capabili- ties in ultra long-horizon scenarios.arXiv preprint arXiv:2509.21766, 2025
arXiv 2025
-
[25]
Rui Ye, Zhongwang Zhang, Kuan Li, Huifeng Yin, Zhengwei Tao, Yida Zhao, Liangcai Su, Liwen Zhang, Zile Qiao, Xinyu Wang, et al. Agentfold: Long-horizon web agents with proactive context management.arXiv preprint arXiv:2510.24699, 2025
arXiv 2025
-
[26]
Shukai Liu, Jian Yang, Bo Jiang, Yizhi Li, Jinyang Guo, Xianglong Liu, and Bryan Dai. Context as a tool: Context management for long-horizon swe-agents.arXiv preprint arXiv:2512.22087, 2025
arXiv 2025
-
[27]
React: Synergizing reasoning and acting in language models
Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik R Narasimhan, and Yuan Cao. React: Synergizing reasoning and acting in language models. InThe eleventh international conference on learning representations, 2022
2022
-
[28]
Reflexion: Language agents with verbal reinforcement learning.Advances in neural information processing systems, 36:8634–8652, 2023
Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao. Reflexion: Language agents with verbal reinforcement learning.Advances in neural information processing systems, 36:8634–8652, 2023
2023
-
[29]
Metagpt: Meta programming for a multi-agent collaborative framework
Sirui Hong, Mingchen Zhuge, Jonathan Chen, Xiawu Zheng, Yuheng Cheng, Jinlin Wang, Ceyao Zhang, Zili Wang, Steven Ka Shing Yau, Zijuan Lin, et al. Metagpt: Meta programming for a multi-agent collaborative framework. InThe twelfth international conference on learning representations, 2023
2023
-
[30]
Chatdev: Communicative agents for software development
Chen Qian, Wei Liu, Hongzhang Liu, Nuo Chen, Yufan Dang, Jiahao Li, Cheng Yang, Weize Chen, Yusheng Su, Xin Cong, et al. Chatdev: Communicative agents for software development. InProceedings of the 62nd annual meeting of the association for computational linguistics (volume 1: Long papers), pages 15174–15186, 2024
2024
-
[31]
Autogen: Enabling next-gen llm applications via multi-agent conversations
Qingyun Wu, Gagan Bansal, Jieyu Zhang, Yiran Wu, Beibin Li, Erkang Zhu, Li Jiang, Xiaoyun Zhang, Shaokun Zhang, Jiale Liu, et al. Autogen: Enabling next-gen llm applications via multi-agent conversations. InFirst conference on language modeling, 2024. 11
2024
-
[32]
Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang
Nelson F. Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang. Lost in the middle: How language models use long contexts.Transactions of the Association for Computational Linguistics, 12:157–173, 2024. doi: 10.1162/tacl_a_00638. URLhttps://aclanthology.org/2024.tacl-1.9/
-
[33]
The illusion of diminishing returns: Measuring long horizon execution in LLMs
Akshit Sinha, Arvindh Arun, Shashwat Goel, Steffen Staab, and Jonas Geiping. The illusion of diminishing returns: Measuring long horizon execution in LLMs. InThe Fourteenth Inter- national Conference on Learning Representations, 2026. URL https://openreview.net/ forum?id=3lm8lWYxiq
2026
-
[34]
Guangya Wan, Mingyang Ling, Xiaoqi Ren, Rujun Han, Sheng Li, and Zizhao Zhang. Compass: Enhancing agent long-horizon reasoning with evolving context.arXiv preprint arXiv:2510.08790, 2025
arXiv 2025
-
[35]
Generative agents: Interactive simulacra of human behavior
Joon Sung Park, Joseph O’Brien, Carrie Jun Cai, Meredith Ringel Morris, Percy Liang, and Michael S Bernstein. Generative agents: Interactive simulacra of human behavior. InProceed- ings of the 36th annual acm symposium on user interface software and technology, pages 1–22, 2023
2023
-
[36]
Patil, Kevin Lin, Sarah Wooders, and Joseph E
Charles Packer, Vivian Fang, Shishir G. Patil, Kevin Lin, Sarah Wooders, and Joseph E. Gon- zalez. Memgpt: Towards llms as operating systems.CoRR, abs/2310.08560, 2023. URL https://doi.org/10.48550/arXiv.2310.08560
-
[37]
V oyager: An open-ended embodied agent with large language models
Guanzhi Wang, Yuqi Xie, Yunfan Jiang, Ajay Mandlekar, Chaowei Xiao, Yuke Zhu, Linxi Fan, and Anima Anandkumar. V oyager: An open-ended embodied agent with large language models. arXiv preprint arXiv:2305.16291, 2023
Pith/arXiv arXiv 2023
-
[38]
Memorybank: Enhancing large language models with long-term memory
Wanjun Zhong, Lianghong Guo, Qiqi Gao, He Ye, and Yanlin Wang. Memorybank: Enhancing large language models with long-term memory. InProceedings of the AAAI conference on artificial intelligence, volume 38, pages 19724–19731, 2024
2024
-
[39]
Hipporag: Neurobio- logically inspired long-term memory for large language models.Advances in neural information processing systems, 37:59532–59569, 2024
Bernal J Gutiérrez, Yiheng Shu, Yu Gu, Michihiro Yasunaga, and Yu Su. Hipporag: Neurobio- logically inspired long-term memory for large language models.Advances in neural information processing systems, 37:59532–59569, 2024
2024
-
[40]
A-mem: Agentic memory for llm agents.Advances in Neural Information Processing Systems, 2025
Wujiang Xu, Zujie Liang, Kai Mei, Hang Gao, Juntao Tan, and Yongfeng Zhang. A-mem: Agentic memory for llm agents.Advances in Neural Information Processing Systems, 2025
2025
-
[41]
Hiagent: Hierarchical working memory management for solving long-horizon agent tasks with large lan- guage model
Mengkang Hu, Tianxing Chen, Qiguang Chen, Yao Mu, Wenqi Shao, and Ping Luo. Hiagent: Hierarchical working memory management for solving long-horizon agent tasks with large lan- guage model. InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 32779–32798, 2025
2025
-
[42]
Xixi Wu, Kuan Li, Yida Zhao, Liwen Zhang, Litu Ou, Huifeng Yin, Zhongwang Zhang, Xinmiao Yu, Dingchu Zhang, Yong Jiang, et al. Resum: Unlocking long-horizon search intelligence via context summarization.arXiv preprint arXiv:2509.13313, 2025
arXiv 2025
-
[43]
Minki Kang, Wei-Ning Chen, Dongge Han, Huseyin A Inan, Lukas Wutschitz, Yanzhi Chen, Robert Sim, and Saravan Rajmohan. Acon: Optimizing context compression for long-horizon llm agents.arXiv preprint arXiv:2510.00615, 2025
Pith/arXiv arXiv 2025
-
[44]
Mo Li, LH Xu, Qitai Tan, Long Ma, Ting Cao, and Yunxin Liu. Sculptor: Empowering llms with cognitive agency via active context management.arXiv preprint arXiv:2508.04664, 2025
arXiv 2025
-
[45]
Automatic generation of natural language summaries for java classes
Laura Moreno, Jairo Aponte, Giriprasad Sridhara, Andrian Marcus, Lori Pollock, and K Vijay- Shanker. Automatic generation of natural language summaries for java classes. In2013 21st International conference on program comprehension (ICPC), pages 23–32. IEEE, 2013
2013
-
[46]
Summarizing source code using a neural attention model
Srinivasan Iyer, Ioannis Konstas, Alvin Cheung, and Luke Zettlemoyer. Summarizing source code using a neural attention model. In Katrin Erk and Noah A. Smith, editors,Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2073–2083, Berlin, Germany, August 2016. Association for Computational ...
-
[47]
Recommendations for datasets for source code summarization
Alexander LeClair and Collin McMillan. Recommendations for datasets for source code summarization. InProceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 3931–3937, 2019
2019
-
[48]
Readsum: retrieval-augmented adaptive transformer for source code summarization.IEEE Access, 11:51155–51165, 2023
Yunseok Choi, Cheolwon Na, Hyojun Kim, and Jee-Hyong Lee. Readsum: retrieval-augmented adaptive transformer for source code summarization.IEEE Access, 11:51155–51165, 2023
2023
-
[49]
Bibek Poudel, Adam Cook, Sekou Traore, and Shelah Ameli. Documint: Docstring generation for python using small language models.arXiv preprint arXiv:2405.10243, 2024
Pith/arXiv arXiv 2024
-
[50]
Automatic code documentation generation using gpt-
Junaed Younus Khan and Gias Uddin. Automatic code documentation generation using gpt-
-
[51]
InProceedings of the 37th IEEE/ACM International Conference on Automated Software Engineering, pages 1–6, 2022
2022
-
[52]
ProConSuL: Project context for code summarization with LLMs
Vadim Lomshakov, Andrey Podivilov, Sergey Savin, Oleg Baryshnikov, Alena Lisevych, and Sergey Nikolenko. ProConSuL: Project context for code summarization with LLMs. In Franck Dernoncourt, Daniel Preo¸ tiuc-Pietro, and Anastasia Shimorina, editors,Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: Industry Track, pages...
-
[53]
Code summarization beyond function level
Vladimir Makharev and Vladimir Ivanov. Code summarization beyond function level. In2025 IEEE/ACM International Workshop on Large Language Models for Code (LLM4Code), pages 153–160. IEEE, 2025
2025
-
[54]
Large language models are few-shot summarizers: Multi-intent comment generation via in-context learning
Mingyang Geng, Shangwen Wang, Dezun Dong, Haotian Wang, Ge Li, Zhi Jin, Xiaoguang Mao, and Xiangke Liao. Large language models are few-shot summarizers: Multi-intent comment generation via in-context learning. InProceedings of the 46th IEEE/ACM International Conference on Software Engineering, pages 1–13, 2024
2024
-
[55]
Automatic semantic augmentation of language model prompts (for code summarization)
Toufique Ahmed, Kunal Suresh Pai, Premkumar Devanbu, and Earl Barr. Automatic semantic augmentation of language model prompts (for code summarization). InProceedings of the IEEE/ACM 46th international conference on software engineering, pages 1–13, 2024
2024
-
[56]
Code needs comments: Enhancing code llms with comment augmentation
Demin Song, Honglin Guo, Yunhua Zhou, Shuhao Xing, Yudong Wang, Zifan Song, Wenwei Zhang, Qipeng Guo, Hang Yan, Xipeng Qiu, et al. Code needs comments: Enhancing code llms with comment augmentation. InFindings of the Association for Computational Linguistics: ACL 2024, pages 13640–13656, 2024
2024
-
[57]
Rethinking-based code summarization with chain of comments
Liuwen Cao, Hongkui He, Hailin Huang, Jiexin Wang, and Yi Cai. Rethinking-based code summarization with chain of comments. In Owen Rambow, Leo Wanner, Marianna Apidianaki, Hend Al-Khalifa, Barbara Di Eugenio, and Steven Schockaert, editors,Proceedings of the 31st International Conference on Computational Linguistics, pages 3043–3056, Abu Dhabi, UAE, Janua...
2025
-
[58]
Summac: Re-visiting nli-based models for inconsistency detection in summarization.Transactions of the Association for Computational Linguistics, 10:163–177, 2022
Philippe Laban, Tobias Schnabel, Paul N Bennett, and Marti A Hearst. Summac: Re-visiting nli-based models for inconsistency detection in summarization.Transactions of the Association for Computational Linguistics, 10:163–177, 2022
2022
-
[59]
Fenice: Factuality evaluation of summarization based on natural language inference and claim extraction
Alessandro Scirè, Karim Ghonim, and Roberto Navigli. Fenice: Factuality evaluation of summarization based on natural language inference and claim extraction. InFindings of the Association for Computational Linguistics: ACL 2024, pages 14148–14161, 2024
2024
-
[60]
FIZZ: Factual in- consistency detection by zoom-in summary and zoom-out document
Joonho Yang, Seunghyun Yoon, ByeongJeong Kim, and Hwanhee Lee. FIZZ: Factual in- consistency detection by zoom-in summary and zoom-out document. In Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen, editors,Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 30–45, Miami, Florida, USA, November
2024
-
[61]
doi: 10.18653/v1/2024.emnlp-main.3
Association for Computational Linguistics. doi: 10.18653/v1/2024.emnlp-main.3. URL https://aclanthology.org/2024.emnlp-main.3/. 13
-
[62]
Suyoung Bae, CheolWon Na, Jaehoon Lee, Yumin Lee, YunSeok Choi, and Jee-Hyong Lee. Referee: Reference-free and fine-grained method for evaluating factual consistency in real-world code summarization.arXiv preprint arXiv:2604.10520, 2026
Pith/arXiv arXiv 2026
-
[63]
FActScore: Fine-grained atomic evaluation of factual precision in long form text generation
Sewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis, Wen-tau Yih, Pang Koh, Mohit Iyyer, Luke Zettlemoyer, and Hannaneh Hajishirzi. FActScore: Fine-grained atomic evaluation of factual precision in long form text generation. In Houda Bouamor, Juan Pino, and Kalika Bali, editors,Proceedings of the 2023 Conference on Empirical Methods in Natural Language Proc...
-
[64]
DevEval: A manually-annotated code generation benchmark aligned with real-world code repositories
Jia Li, Ge Li, Yunfei Zhao, Yongmin Li, Huanyu Liu, Hao Zhu, Lecheng Wang, Kaibo Liu, Zheng Fang, Lanshen Wang, Jiazheng Ding, Xuanming Zhang, Yuqi Zhu, Yihong Dong, Zhi Jin, Binhua Li, Fei Huang, Yongbin Li, Bin Gu, and Mengfei Yang. DevEval: A manually-annotated code generation benchmark aligned with real-world code repositories. In Lun-Wei Ku, Andre Ma...
-
[65]
u depends on / contains v
Robert Tarjan. Depth-first search and linear graph algorithms.SIAM journal on computing, 1 (2):146–160, 1972. 14 A Additional details about MemDocAgent A.1 Algorithms of dependency-aware traversal guiding Algorithm 1 describes dependency graph construction, and Algorithm 2 presents the topological traversal used for hierarchical generation. Algorithm 1Bui...
1972
-
[66]
Focus on the big picture, not implementation details
REPO : Repository-level Documentation: - Brief introduction and purpose of the overall system - Architecture overview with diagrams - High-level functionality of each sub-module including references to its documentation file - Link to other module documentation instead of duplicating information - Do not duplicate content covered in MODULE or COMPONENT do...
-
[67]
MODULE: Module-level Documentation: - Explanation of the module’s role within the system and its internal design, so a developer can understand *how* its components fit together before reading individual component details - Responsibility and boundaries of the module - List of core components with a one-line description each - Component interaction diagra...
-
[68]
COMPONENT: Component-level Documentation: - Providing enough detail to *reimplement* the function, method, or class correctly — covering inputs, outputs, behavior, edge cases, and constraints - Summary of what the component does and why it exists (not how it works) </DOCUMENTATION_STRUCTURE> <WORKFLOW>
-
[69]
You will first receive a sub-task, which includes the type of task (COMPONENT, MODULE, or REPO), the target component/module/repo to document, and other relevant information
-
[70]
Analyze the provided code components or module structure, explore the not given dependencies between the components if needed
-
[71]
For COMPONENT tasks, generate the documentation for the specific component, and save the documentation in memory with the name of ‘component_id’
-
[72]
For MODULE tasks, synthesize the documentations of sub-components and generate the module-level documentation, and save the documentation in memory with the name of ‘module_id’
-
[73]
For REPO tasks, synthesize the documentations of all modules and generate the repository-level documentation, and save the documentation in memory with the name of ‘repo_id’
-
[74]
For each task, you perform thought-action-observation loops to iteratively improve the documentation until it passes verification, then save the final documentation to memory and return. - At every turn, you MUST follow this structure: Thought:〈your reasoning about what to do next, what information you need〉 Action: Choose exactly one action from the list...
-
[75]
READ:If you think more information is needed to generate high-quality documentation of the target component, use this action to request relevant information. - During think step, you should analyze the current code and context, and explain what additional information might be needed (if any) - You have access to three types of information sources:
-
[76]
Sub-components or sub-modules (from memory): - If the target is a MODULE or REPO, you can request the documentation of its sub-components or sub-modules that have already been documented from memory. - This is the primary source of information for MODULE and REPO tasks, since the module/repo-level documentation should be synthesized based on the already g...
-
[77]
Internal Codebase Information (from local code repository): For Functions: - Code components called within the function body - Places where this function is called For Methods: - Code components called within the method body - Places where this method is called - The class this method belongs to For Classes: - Code components called in the __init__ method...
-
[78]
Only request it when understanding an external third-party API or library is essential for accurate documentation, and that information cannot be found within the target codebase
External Open Internet retrieval Information: - External Retrieval is extremely expensive. Only request it when understanding an external third-party API or library is essential for accurate documentation, and that information cannot be found within the target codebase. - Use the import statements in <IMPORT_INFORMATION_IN_THE_FILE> to identify candidates...
-
[79]
""Reads a file and returns its content as a list of lines
WRITE:If you think you have collected sufficient context, use this action and generate the documentation for the target task type. - General guidelines for high-quality documentation: - Make documentations actionable and specific: Focus on practical usage. - Use clear, concise language: Avoid jargon unless necessary, use active voice, and be direct and sp...
-
[80]
- Verification Process: - First read the target task information (source code and related information) as if you’re seeing it for the first time
VERIFY:After generating a documentation, use this action to self-evaluate the documentation quality along three criteria, each scored from 0.00 to 1.00 (two decimal places). - Verification Process: - First read the target task information (source code and related information) as if you’re seeing it for the first time. - Read the generated documentation an...
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.