{"id":"fe30cff4-3af7-4070-876c-9a3aac813985","arxiv_id":"2607.25975","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Scientists are developing personal, unshared 'landmarking' habits to separate human-meaningful code from agent context, and this fragmentation may complicate collaboration.","lead":"A study of how scientists use AI coding assistants finds individuals inventing private conventions—'landmarking strategies'—for marking code as human-readable versus agent context. A smart generalist might read it to see an early, concrete picture of how AI agents are quietly reshaping scientific teamwork and code maintainability.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Core generalization lacks a baseline: no evidence that landmarking is new, agent-caused, or more idiosyncratic than pre-agent practice.","rationale":"The reader flagged representativeness and attribution; I partially agree. My concern sharpens the attribution issue into a missing-baseline/causal-identification problem. The condition that must hold for the strongest claim is that landmarking strategies are a response to agentic tools and are becoming more idiosyncratic. The paper's 'What I'm seeing' section supplies only four cases and the author's narrative; no comparison or baseline is offered. Because the paper is positioned as a position piece and explicitly disclaims direct evidence for the collaboration prediction, the appropriate verdict remains CONDITIONAL rather than REJECT. I would not upgrade to unconditional accept without the survey-based comparison and a coding scheme for landmarking strategies. This is a request for empirical anchor, not a challenge to the concept's plausibility.","tokens_in":3705,"tokens_out":4400,"duration_ms":48173,"concrete_test":"Re-analyze the 2025 survey dataset (ref [7]) comparing agentic-tool users against chatbot-only/no-AI users on items about version-control practices, commit-message authorship, documentation habits, and whether code comments are written for humans or agents, matching on prior Git experience, programming experience, and field. If agentic-tool users do not show significantly more idiosyncratic/private landmarking behavior than the control group, the 'de-standardization caused by agents' claim is unsupported. As a complement, ask the four case participants to reconstruct their pre-agent version-control and documentation habits; if their current practices are unchanged, the attribution to agent use fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim ('the idiosyncrasies of scientific code may in fact be growing more particular'; 'conventions fragment') requires that the observed landmarking strategies be both caused by agentic tool use and increasing relative to a prior baseline. The evidence is four contemporaneous contextual-inquiry cases plus the author's own experience; there is no longitudinal data, no comparison group of non-agent or chatbot-only scientists, and no retrospective account from participants establishing that their practices differ from pre-agent habits. Many described behaviors—commit messages as lab-notebook entries, minimal diffs as decision records, spec documents, chat logs as memory—have recognizable antecedents in scientific programming before agents, and the paper itself notes many scientists do not use Git at all. Without controlling for pre-existing personal style or the researcher's presence, the observations could reflect ordinary variation in programming practice rather than agent-induced de-standardization. The >800-person survey is cited only for adoption rates, not for landmarking prevalence or change, so it cannot rescue the generalization. The collaboration claim is explicitly flagged as lacking direct evidence. This is an empirical support gap, not an internal contradiction; the central hypothesis is plausible but not established.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This position piece argues that as scientists adopt AI coding agents, the longstanding assumption that at least one human understands why code exists is breaking down. Based on an ongoing contextual inquiry (four completed cases), a survey of >800 scientific programmers, and the author's own analysis workflows, it introduces the concept of \"landmarking strategies\": personal, often unshared rules for marking which parts of a codebase are for human understanding and which are context for agents. Examples include typed commit messages as lab-notebook entries, minimal-merge diffs as decision records, spec documents committed instead of code review, and chat logs plus a single appended script for non-Git users. The paper argues that this quiet de-standardization may complicate collaboration and predicts that teams that explicitly delineate human-readable versus agent-context artifacts will be better able to maintain scientific code.","tokens_in":3908,"tokens_out":5106,"duration_ms":52254,"significance":"The topic is timely and the observations are vivid and thought-provoking. The paper identifies a plausible new phenomenon—the repurposing of version control and other infrastructure as personal landmarks in agent-mediated work—and connects it to existing concepts of cognitive and intent debt. The concrete prediction (teams that explicitly delineate human-readable vs agent context will fare better) is falsifiable and could be tested. However, the evidence base is thin for the central generalization: four cases, the author's own experience, and a survey used only for adoption statistics. If reframed as a hypothesis-generating position piece, the contribution is real; as a statement of what \"scientists are inventing,\" the current support is insufficient. The paper's explicit acknowledgment of lacking direct evidence for the collaboration claim is to its credit, but it makes the abstract's stronger wording harder to accept.","major_comments":[{"comment":"The central generalization—'scientists are inventing personal conventions' and 'quiet de-standardization'—rests on four completed contextual-inquiry cases. No coding scheme, saturation analysis, or systematic cross-case comparison is reported, and the cases are not described in enough detail. For a position piece this is acceptable as hypothesis generation, but the abstract states it as an observed fact. Please either report the cases more systematically or reframe the claim as 'four observed cases suggest…' and carry that caveat through the abstract.","section":"Abstract and §'What I'm seeing'"},{"comment":"The claim that idiosyncrasies are 'growing more particular' and are agent-caused requires a baseline. Several behaviors have recognizable pre-agent antecedents: commit messages as lab notebooks, minimal-diff merges as decision records, spec documents, and long scripts with comments. Without a comparison group of non-agent/chatbot-only programmers or retrospective participant data, the observations could reflect pre-existing personal style. The >800-person survey [7] is cited only for adoption rates (roughly three quarters via ChatGPT), not for landmarking prevalence or change, so it does not rescue the temporal or causal claim.","section":"§'What I'm seeing' — baseline/causality"},{"comment":"The collaboration prediction is load-bearing for the abstract's 'could complicate collaboration', yet the paper explicitly states 'I do not have direct evidence to corroborate this yet.' The expectation is plausible, but presenting it as an implication ('For these reasons...') conflates a hypothesis with a finding. Either gather direct evidence (e.g., paired collaborator interviews or codebase-handoff studies) or mark this as a speculative design implication in both the abstract and body.","section":"§Collaboration"},{"comment":"The author's own Git-paralysis example is presented as a supporting data point. It is an anecdote from a single individual who is also the researcher, and the generalization to 'I would bet this is not uncommon' is unsupported. If kept, it should be labeled as motivation/hypothesis, not as a fifth observation, and the basis for the 'not uncommon' bet should be explicit.","section":"§'What I'm seeing' — personal narrative"}],"minor_comments":[{"comment":"The paper never states how the four cases were selected, what disciplines or tools were represented, or how field notes/interviews were analyzed. A short 'method' note or a table of case characteristics would help readers judge transferability.","section":"Methods/positioning"},{"comment":"Reference [7] lists the first author as 'O'Brien, G.', which appears inconsistent with the author's name on this paper (Elle O'Brien). If this is the author's own survey, the name should be consistent; if it is a different author, the citation may be a typo.","section":"References"},{"comment":"There are typographical errors, e.g., 'suﬀicient' (p. 1) and 'doesn ’tmatter' (p. 2). Also the phrase 'For many scientists ( terminology note:...' uses an unneeded space after the opening parenthesis.","section":"Typos"},{"comment":"The term 'scientist' is deliberately stretched to include research software engineers, but the paper sometimes slides between 'scientists I've met' and 'all of these scientists.' Consider using 'scientific programmers' consistently or flag the scope each time.","section":"Terminology"}],"recommendation":"major_revision","confidential_remarks":"This is an arXiv preprint-style position paper. For a CS.HC venue that welcomes position pieces, the topic and examples are valuable, but the central empirical claim needs substantial softening or additional evidence. The author is transparent about the missing collaboration evidence, which is good; the main fix is aligning the abstract and framing with the actual evidence base. The self-citation [7] with a mismatched initial should be corrected. No concerns about the authors' integrity; the issue is evidence-proportionate wording."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The landmarking concept is the real contribution here. Naming the practice of triaging code artifacts into 'human-readable' versus 'agent context' gives researchers a handle on something that is happening but was previously diffuse. The four vignettes—commit messages as lab-notebook entries, minimal-diff-as-choice, spec-as-reviewable-artifact, unread agent comments left for future agents—are concrete and recognizable. The paper is also honest about its status as a position piece: the author explicitly says she has four cases so far and that she lacks direct evidence for the collaboration breakdown. That transparency earns credit.\n\nThe soft spot is exactly what the stress-test note flags: no baseline. The claim that idiosyncrasies 'may in fact be growing more particular' requires showing that these practices are new, agent-caused, and increasing relative to pre-agent habits. We don't get that. The four cases are contemporaneous, the author's own workflow is one of them, and several behaviors have clear antecedents in how scientists have always managed exploratory code. The >800-person survey is cited only for adoption rates, not for landmarking prevalence or change. So the central generalization is plausible but unproven.\n\nThat said, I don't think this is a load-bearing flaw. The paper frames itself as an ongoing contextual inquiry and repeatedly hedges the ambitious claims with 'I expect' and 'I don't have a full theory yet.' For a position piece, it is doing what position pieces should do: naming a phenomenon, illustrating it, and sketching a research agenda. The collaboration prediction is explicitly flagged as a prediction. The main remedy is resubmission-with-revision: add a coding scheme, a sampling/saturation discussion, and a plan for measuring change over time or comparing agent vs. non-agent groups. Also, the author should consider whether 'scientists are inventing' is too strong given that some observed strategies may be repurposed personal habits rather than new inventions.\n\nThe writing is clear, the citations are appropriate, and the self-citation to the survey is minor and does not drive the central claim. This paper deserves a serious referee. It is exactly the kind of intervention that can shape future work in human-AI collaboration in scientific programming, as long as the empirical limits are treated as open problems rather than solved questions.\n\nRecommendation: send to peer review. A conditional accept is right, with the condition being a more disciplined framing of the generalization and an explicit future-work plan.","headline":"A genuinely new conceptual label ('landmarking strategies') with vivid case observations, but the empirical generalization outruns the four cases; still worth refereeing as a position piece.","tokens_in":4392,"tokens_out":1593,"would_cite":true,"duration_ms":19285,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Scientists using coding agents are inventing personal 'landmarking strategies' that mark which parts of a codebase are for humans versus agents, quietly fragmenting collaboration.","keywords":["agent-written code","landmarking strategies","human-readable artifacts","scientific programming","version control conventions","cognitive debt","collaboration","contextual inquiry"],"falsifier":"A larger observational study that finds no systematic difference in the idiosyncrasy of version-control and documentation practices between agent-using and non-agent-using scientific teams would directly undercut the claim that agents are driving de-standardization.","tokens_in":3560,"feed_emoji":"🤖","tokens_out":3465,"duration_ms":31964,"temperature":0.7,"pith_summary":"The paper argues that the longstanding assumption of scientific code—that at least one person understands why it exists—is breaking down as coding agents generate more code than any one person can review. Drawing on four in-depth contextual-inquiry cases, a survey of over 800 scientific programmers, and the author's own analysis workflow, it describes how scientists are developing idiosyncratic 'landmarking strategies': personal rules for triaging artifacts into those meant for human understanding and those left as context for agents. These strategies often repurpose shared infrastructure like version control in unconventional ways, making commit messages and diffs unreliable as shared documents. The author contends this quiet de-standardization will complicate collaboration across teams with heterogeneous practices, and predicts that teams that explicitly delineate what is human-readable versus agent context will be better able to develop, document, and maintain scientific codebases.","feed_headline":"Agents make scientific code more personal, not less","feed_subtitle":"Scientists are inventing private rules for what humans should read in code—fragmenting team collaboration.","key_machinery":"The central object is the 'landmarking strategy'—a scientist's private, often unarticulated rule for triaging artifacts in a codebase into those meant for human review and those left as context for agents. It carries the argument because the paper claims these strategies are multiplying idiosyncratically and are not interoperable across teams, which is the mechanism that threatens collaboration.","core_discovery":"On the paper's terms, the discovery is that scientists are not merely accepting agent output wholesale; they are inventing personal conventions for what counts as a human-understandable landmark in the codebase. Concretely: one engineer treats commit messages as lab-notebook entries typed only by the scientist; a postdoc treats only merged diffs as the 'choice' to understand and keeps agent specifications transient; another writes spec documents to markdown and reviews specs rather than code; a non-Git user keeps all code in one growing script, leaving agent-written comments he never reads but that he intends as context for future agents. The paper generalizes from these cases to a claim abo","pith_inferences":["A testable extension: compare commit-message style, branch structure, and comment density in repositories before and after a scientist starts using a coding agent; the fragmentation hypothesis predicts measurable increases in idiosyncratic conventions.","The mechanism resembles how personal knowledge-management practices fragment across a lab; a lightweight shared 'landmark schema' (for example, a README pointer or a standardized comment marker) could serve as a coordination artifact without forcing a single workflow.","The paper implies that code review culture, already weak in science, will face new pressure: agents can generate too much code to review, so landmarking defines what little does get reviewed and might become the new locus of oversight.","An implicit consequence: tools that surface repository context (like issue trackers) for non-Git users could be adapted to encode landmark conventions, making intent legible without requiring everyone to adopt Git."],"forward_implications":["Commit messages and diffs can no longer be assumed to be reviewable human units; collaborators must ask where the scientist's 'why' is actually recorded.","Newcomers to a codebase—new students, replicators, or reviewers—will need explicit orientation to a team's landmark conventions or risk misreading what is trustworthy.","Heterogeneous tool use (agents versus simple chatbots, large versus small token budgets) creates incompatible context across collaborators, even within one team.","Teams that explicitly negotiate and document what is human-readable versus agent context will be better able to develop, document, and maintain code.","Version-control systems may need to evolve new conventions that assume both human and agent readers."],"fun_headline_variants":["Agents push scientists to add personal landmarks to code","Code splits: humans and agents need different signposts","Personal code conventions emerge as agents join research","Scientists invent private rules for reading agent-written code"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The central generalization rests on the author's four completed contextual-inquiry cases being representative of scientists broadly, and on the observed behaviors being caused by agentic tools rather than reflecting pre-existing personal habits.","fun_headline_variants_meta":{"raw":{"variants":["Agents push scientists to add personal landmarks to code","Code splits: humans and agents need different signposts","Personal code conventions emerge as agents join research","Scientists invent private rules for reading agent-written code"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000312,"raw_usage":{"total_tokens":1569,"prompt_tokens":660,"completion_tokens":909,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":404,"completion_tokens_details":{"reasoning_tokens":849}},"tokens_in":404,"tokens_out":909,"duration_ms":7323,"temperature":1.0,"reasoning_tokens":849,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T00:55:50.499641+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A larger observational study that finds no systematic difference in the idiosyncrasy of version-control and documentation practices between agent-using and non-agent-using scientific teams would directly undercut the claim that agents are driving de-standardization.","supporting_citations":[],"review_version":1}