{"total":16,"items":[{"citing_arxiv_id":"2607.01916","ref_index":26,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"ContextSniper: AntTrail's Token-Efficient Code Memory for Repository-Level Program Repair","primary_cat":"cs.AI","submitted_at":"2026-07-02T09:15:28+00:00","verdict":null,"verdict_confidence":null,"novelty_score":null,"formal_verification":null,"one_line_summary":null,"context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.31121","ref_index":25,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"The Past Is Prologue: A Plug-in Controller for Selective Updates in Sequentially Evolving LLM Memory","primary_cat":"cs.AI","submitted_at":"2026-06-30T04:33:18+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":5.0,"formal_verification":"none","one_line_summary":"Janus is a method-agnostic plug-in that uses a Memory Momentum Trigger and compact hybrid evaluation to selectively accept LLM memory updates, yielding +2.7 to +4.6 accuracy gains over base updaters on six datasets.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.28436","ref_index":2,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Dockerless: Environment-Free Program Verifier for Coding Agents","primary_cat":"cs.SE","submitted_at":"2026-06-26T06:14:55+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":7.0,"formal_verification":"none","one_line_summary":"Dockerless uses agentic repository exploration to verify patches without execution, enabling SFT and RL training of coding agents that reach 62.0/50.0/35.2% resolve rates on SWE-bench Verified/Multilingual/Pro while matching environment-based results.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.14061","ref_index":5,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"LLM Agents Can See Code Repositories","primary_cat":"cs.SE","submitted_at":"2026-06-12T03:14:40+00:00","verdict":"CONDITIONAL","verdict_confidence":"MODERATE","novelty_score":6.0,"formal_verification":"none","one_line_summary":"Adding visual dependency-graph images to a text-based coding agent cuts token consumption by up to 26% while keeping issue-resolution accuracy roughly unchanged.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.09577","ref_index":32,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Code Is More Than Text: Uncertainty Estimation for Code Generation","primary_cat":"cs.CL","submitted_at":"2026-06-08T14:52:43+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"Three code-specific uncertainty axes (lexical, algorithmic, functional) yield an ensemble that raises average AUROC from 0.696 to 0.776 across five code LLMs, with one single-pass signal matching multi-pass baselines at lower cost.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.07297","ref_index":4,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"SWE-Explore: Benchmarking How Coding Agents Explore Repositories","primary_cat":"cs.SE","submitted_at":"2026-06-05T14:08:27+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":7.0,"formal_verification":"none","one_line_summary":"SWE-Explore is a new benchmark evaluating repository exploration by coding agents on 848 issues across 203 repositories, using line-level ground truth from successful agent trajectories and showing agentic methods outperform classical retrieval on coverage and ranking.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.30105","ref_index":12,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"EvoRepair: Enhancing Vulnerability Repair Agents Through Experience-Based Self-Evolution","primary_cat":"cs.SE","submitted_at":"2026-05-28T15:46:58+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":7.0,"formal_verification":"none","one_line_summary":"EvoRepair is the first experience-based self-evolving agent framework for automated vulnerability repair, reporting 90.46% overall success on PATCHEVAL and SEC-bench benchmarks.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.15384","ref_index":6,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Is One Score Enough? Rethinking the Evaluation of Sequentially Evolving LLM Memory","primary_cat":"cs.LG","submitted_at":"2026-05-14T20:15:22+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"SeqMem-Eval reveals that high final accuracy in sequential LLM memory tasks often coexists with substantial forgetting and negative transfer, exposing stability-adaptability trade-offs hidden by standard aggregate metrics.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.10674","ref_index":7,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Step Rejection Fine-Tuning: A Practical Distillation Recipe","primary_cat":"cs.LG","submitted_at":"2026-05-11T14:55:20+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"Step Rejection Fine-Tuning masks loss on erroneous steps identified by a critic LLM in unresolved trajectories, raising SWE-bench Verified resolution rate by 3.7% to 32.2% versus 2.4% for trajectory-level rejection.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2604.08224","ref_index":19,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Externalization in LLM Agents: A Unified Review of Memory, Skills, Protocols and Harness Engineering","primary_cat":"cs.SE","submitted_at":"2026-04-09T13:19:41+00:00","verdict":"ACCEPT","verdict_confidence":"LOW","novelty_score":5.0,"formal_verification":"none","one_line_summary":"LLM agent progress depends on externalizing cognitive functions into memory, skills, protocols, and harness engineering that coordinates them reliably.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2604.04580","ref_index":11,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Beyond Fixed Tests: Repository-Level Issue Resolution as Coevolution of Code and Behavioral Constraints","primary_cat":"cs.SE","submitted_at":"2026-04-06T10:26:46+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"Agent-CoEvo is a multi-agent LLM framework that coevolves code patches and test patches to resolve repository-level issues, outperforming fixed-test baselines on SWE-bench Lite and SWT-bench Lite.","context_count":1,"top_context_role":"background","top_context_polarity":"background","context_text":"structured reasoning. Methods such as RepoGraph [38], knowledge-graph-augmented approaches [49], and intent-grounding systems like SpecRover [40] inject structural and semantic information into the repair process. Search-driven techniques including SWE-Search [5] and multi-agent debate frameworks [23] further guide patch exploration, while experience-based systems [11, 32] leverage memory to transfer prior repair knowledge across tasks. Recent studies also highlight the importance of inference-time compute scaling. Approaches such as Thinking Longer, Not Larger [29] and Trae Agent [15] show that deeper test-time reasoning and reflection substantially improve repair success. Complementary work like BugPilot [43] focuses on"},{"citing_arxiv_id":"2603.22048","ref_index":4,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Dynamic analysis enhances issue resolution","primary_cat":"cs.SE","submitted_at":"2026-03-23T14:48:54+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"Embedding dynamic analysis into an LLM repair agent yields a claimed 79.4% resolution rate on SWE-bench Verified while cutting token use by about 25%.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2603.12572","ref_index":9,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"LMEB: Long-horizon Memory Embedding Benchmark","primary_cat":"cs.CL","submitted_at":"2026-03-13T02:09:57+00:00","verdict":"CONDITIONAL","verdict_confidence":"MODERATE","novelty_score":6.0,"formal_verification":"none","one_line_summary":"LMEB is a new benchmark that evaluates embedding models on long-horizon memory retrieval and shows this skill is largely orthogonal to traditional passage-retrieval performance.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2602.01785","ref_index":21,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Seeing is Coding: On the Effectiveness of Vision Language Models in Code Understanding","primary_cat":"cs.CL","submitted_at":"2026-02-02T08:10:21+00:00","verdict":null,"verdict_confidence":null,"novelty_score":null,"formal_verification":null,"one_line_summary":null,"context_count":1,"top_context_role":"background","top_context_polarity":"background","context_text":"traditional digitization [5, 9, 45, 48, 57, 71, 78, 97] to end-to-end neural approaches. Early systems like TrOCR [53] and Nougat [12] demonstrated direct transcription without separate detection stages, while GOT-OCR2.0 [93] enhanced structure recovery for charts and tables. General-purpose MLLMs [10, 24, 31, 50, 51, 56, 62] have advanced high-resolution visual understanding, with specialized models for document comprehension [ 21, 40, 86] and GUI understanding [ 8, 101]. DeepSeek-OCR [94] introduced optical compression for documents, achieving up to 20×ratios. However, these works focus on natural documents or UI screenshots, where visual layouts are loosely structured. Code presents unique challenges with dense symbolic content and strict syntactic constraints [15, 79]. Our study systematically evaluates how MLLMs handle code-specific visual"},{"citing_arxiv_id":"2601.00376","ref_index":6,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"In Line with Context: Repository-Level Code Generation via Context Inlining","primary_cat":"cs.SE","submitted_at":"2026-01-01T15:56:24+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":7.0,"formal_verification":"none","one_line_summary":"InlineCoder reframes repository-level code generation as function-level coding by using a draft anchor to inline the target function into its call graph for upstream usage and downstream dependency context.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2509.14635","ref_index":5,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"SWE-QA: Can Language Models Answer Repository-level Code Questions?","primary_cat":"cs.CL","submitted_at":"2025-09-18T05:25:32+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":7.0,"formal_verification":"none","one_line_summary":"SWE-QA creates a new repository-level code QA benchmark with 576 pairs and an agentic LLM framework, showing promise but open challenges for models handling complex codebases.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null}],"limit":50,"offset":0}