{"id":"07db50a9-6aef-4c91-97a3-961c5570ec02","arxiv_id":"2606.23127","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"AFTER benchmark shows single refinement improves LLM agent performance by 3.7-6.7 points and multi-model procedural skills reach 73.1% cross-model accuracy on 382 tasks.","lead":"The paper introduces the AFTER benchmark with 382 enterprise tasks to test how procedural memory helps LLM agents reuse skills across jobs and models. A smart generalist might read it to see concrete numbers on whether memory systems make AI agents more reliable in real work settings.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"Benchmark representativeness for real enterprise workflows remains the central unverified assumption","rationale":"The reader's weakest_assumption correctly isolates the gap between benchmark-internal results and the industrial applicability asserted in the strongest_claim. Full-text details on benchmark construction do not add external validation data that would mitigate this risk, so the provisional UNVERDICTED stance is appropriate.","tokens_in":1675,"tokens_out":329,"duration_ms":14568,"concrete_test":"Sample 50-100 anonymized production task traces from an enterprise agent deployment; compute distributional overlap (task length, role-specific terminology density, error-mode frequency, average branching factor) against AFTER; if key statistics diverge by >25%, re-evaluate the main refinement and cross-model experiments on the real traces and check whether the 3.7-6.7 point and 73.1% figures hold.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline claim that procedural memory 'delivers consistent gains in industrial workflows' (with 3.7-6.7 point improvements and 73.1% cross-model accuracy) depends on AFTER's 382 tasks and 22 skills being representative of production settings. The paper positions the benchmark as spanning six roles with controlled transfer settings, yet provides no external anchoring (e.g., comparison to deployed logs, expert validation of task distributions, or coverage of long-horizon/noisy real-world conditions). If AFTER systematically under-samples certain workflow characteristics common in industry, the measured local improvement, cross-task/role transfer, and multi-model generalization results do not necessarily extrapolate.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper introduces AFTER, a benchmark of 382 realistic enterprise tasks spanning six professional roles and 22 procedural skills, to evaluate procedural memory in LLM agents under controlled settings for local improvement, cross-task/role transfer, and cross-model generalization. It claims that a single refinement round yields 3.7-6.7 point aggregate gains and that skills evolved from diverse multi-model traces reach 73.1% cross-model accuracy (outperforming single-model sources), while some skills generalize broadly and others specialize to roles.","tokens_in":1812,"tokens_out":385,"duration_ms":10566,"significance":"If the benchmark holds as representative, the results supply actionable guidance on skill evolution and transfer for production agent platforms, including the value of multi-model traces and the distinction between generalizable versus role-specific skills.","major_comments":[{"comment":"Abstract: the headline claim that procedural memory 'delivers consistent gains in industrial workflows' is load-bearing on AFTER's 382 tasks and 22 skills being representative of real enterprise settings, yet the manuscript supplies no external anchoring (comparison to deployed logs, expert validation of task distributions, or coverage of long-horizon/noisy conditions).","section":"Abstract"},{"comment":"Abstract (performance claims): the reported numeric gains (3.7-6.7 points, 73.1% cross-model accuracy) are presented without details on task selection criteria, scoring rubrics, statistical significance testing, or controls for prompt variation, leaving the central empirical results on unexamined experimental design choices.","section":"Abstract"}],"minor_comments":[{"comment":"Abstract: the phrase 'controlled evaluation settings for local improvement, cross-task transfer...' is used without a forward reference to the specific protocol definitions or tables that implement them.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the detailed feedback on the abstract and empirical presentation. We address each major comment below, acknowledging limitations where the manuscript falls short of ideal standards for benchmark validation and experimental transparency. We propose targeted revisions to strengthen the claims without overstating the work's scope.","responses":[{"response":"We agree that the phrasing 'industrial workflows' implies broader representativeness than the benchmark construction can support. The 382 tasks were synthesized from role descriptions and procedural patterns drawn from public enterprise documentation and expert consultation, but the manuscript indeed provides no direct comparison to proprietary deployed logs or formal expert validation of task distributions. We cannot supply such anchoring without access to confidential production data. We will revise the abstract to qualify the claim as applying 'in the controlled evaluation settings of the AFTER benchmark' and add an explicit limitations paragraph discussing the synthetic nature of the tasks and the absence of long-horizon or noisy real-world conditions. This addresses the concern without misrepresenting the contribution.","revision_made":"yes","referee_comment":"[Abstract] Abstract: the headline claim that procedural memory 'delivers consistent gains in industrial workflows' is load-bearing on AFTER's 382 tasks and 22 skills being representative of real enterprise settings, yet the manuscript supplies no external anchoring (comparison to deployed logs, expert validation of task distributions, or coverage of long-horizon/noisy conditions)."},{"response":"The full manuscript contains dedicated sections on benchmark construction (including task selection criteria and coverage of the 22 skills), the evaluation protocol (scoring rubrics with human-verified rubrics), and experimental controls (including prompt templates and model backbones). However, the referee is correct that these details are not summarized in the abstract and that statistical significance testing and explicit controls for prompt variation are not highlighted in the reported results. We will expand the abstract with a brief methods clause and add statistical significance results (paired t-tests with p-values) plus prompt-variation ablation tables to the main results section. These changes make the experimental design choices more transparent without altering the reported numbers.","revision_made":"yes","referee_comment":"[Abstract] Abstract (performance claims): the reported numeric gains (3.7-6.7 points, 73.1% cross-model accuracy) are presented without details on task selection criteria, scoring rubrics, statistical significance testing, or controls for prompt variation, leaving the central empirical results on unexamined experimental design choices."}],"tokens_in":1291,"tokens_out":489,"duration_ms":10034,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The useful part is the AFTER benchmark: 382 tasks, six roles, 22 skills, with explicit settings for local refinement, cross-task transfer, cross-role transfer, and cross-model generalization. The multi-model trace result at 73.1% and the observation that some skills stay general while others specialize are concrete data points that earlier agent benchmarks did not report in this combination.\n\nThe paper does a clean job of running the same skills through single-model versus mixed-model traces and measuring the difference. That setup lets you see where diversity helps without fitting extra parameters.\n\nThe soft spot is exactly the one the stress test flags. Nothing in the abstract anchors the 382 tasks to real enterprise logs, expert review, or coverage of noisy long-horizon cases. Without that, the 3.7-6.7 point gains and the cross-model number stay local to the benchmark. The abstract also skips any mention of scoring rubrics, inter-annotator agreement, or prompt-variation controls, so the numeric claims are hard to read as robust.\n\nThis is for people who build or evaluate production agent platforms and want transfer numbers to compare against. A reader already working on memory mechanisms or agent benchmarks will find the experimental grid worth looking at.\n\nIt should go to peer review. The benchmark design and the multi-model comparison are new enough that referees can check the missing details and decide how far the numbers travel.","headline":"AFTER gives a new controlled benchmark for procedural memory transfer across tasks/roles/models, but the industrial gains rest on unanchored task selection.","tokens_in":2317,"tokens_out":362,"would_cite":false,"duration_ms":11889,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Procedural memory improves LLM agent performance on 382 enterprise tasks by 3.7-6.7 points after one refinement round, with multi-model skills reaching 73.1% cross-model accuracy.","keywords":["procedural memory","LLM agents","enterprise workflows","skill transfer","benchmark evaluation","cross-model generalization","agent refinement"],"falsifier":"Running the same refinement and skill-evolution procedures on a fresh collection of enterprise tasks outside the AFTER benchmark and observing no performance gains or transfer benefits.","tokens_in":2599,"feed_emoji":"🤖","tokens_out":680,"duration_ms":12725,"temperature":0.7,"pith_summary":"The paper presents the AFTER benchmark of 382 realistic enterprise tasks across six professional roles and 22 procedural skills to measure how procedural memory transfers in LLM agents. Experiments demonstrate that a single round of skill refinement raises aggregate performance by 3.7-6.7 points in industrial workflows. Skills derived from diverse multi-model execution traces reach 73.1% accuracy when tested on different models, beating all single-model sources. Some skills transfer broadly across tasks and models, while others specialize to particular roles and lose effectiveness outside those contexts. The results supply concrete guidance for constructing, testing, and operating procedural memory in production agent platforms.","feed_headline":"Procedural memory lifts LLM agent scores 3.7-6.7 points","feed_subtitle":"One refinement round plus multi-model skills reach 73.1% accuracy across models on 382 workplace tasks.","key_machinery":"The AFTER benchmark, which tests procedural memory transfer across tasks, roles, and model backbones through controlled settings for local improvement, cross-task transfer, cross-role transfer, and cross-model generalization.","core_discovery":"Procedural memory delivers consistent gains in industrial workflows: a single refinement round improves aggregate performance by 3.7-6.7 points, while skills evolved from diverse multi-model execution traces achieve 73.1% cross-model test accuracy, outperforming all single-model trace sources. Some skills generalize broadly across tasks and models, whereas others become specialized to role-specific workflows and lose effectiveness under transfer.","pith_inferences":["If the benchmark tasks capture typical enterprise patterns, organizations could maintain a shared library of refined skills updated from multiple model sources.","Specialized skills that lose transfer value suggest a need for role-aware skill routing mechanisms in agent systems.","Cross-model generalization at 73.1% indicates procedural memory could help reduce lock-in to any particular LLM provider."],"forward_implications":["A single refinement round on procedural memory yields measurable performance lifts across multiple industrial workflows.","Skills built from multi-model execution traces generalize better across different LLM backbones than skills from any single model.","Broadly generalizing skills can be reused across tasks and roles, while role-specific skills require separate maintenance.","Production agent platforms can use cross-model accuracy as a selection criterion when evolving reusable skills."],"fun_headline_variants":["LLM agents gain 3.7-6.7 points via procedural memory after refinement","Multi-model traces achieve 73.1% cross-model accuracy in agents","Procedural skills show broad or role-specific transfer in enterprise tasks","382 workplace tasks benchmark procedural memory control and adaptation"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The 382 tasks and 22 skills in AFTER are representative enough of real enterprise workflows that the measured transfer and generalization effects will hold outside the benchmark.","fun_headline_variants_meta":{"raw":{"variants":["LLM agents gain 3.7-6.7 points via procedural memory after refinement","Multi-model traces achieve 73.1% cross-model accuracy in agents","Procedural skills show broad or role-specific transfer in enterprise tasks","382 workplace tasks benchmark procedural memory control and adaptation"]},"model":"grok-4.3","cost_usd":0.007161,"raw_usage":{"total_tokens":3284,"prompt_tokens":624,"num_sources_used":0,"completion_tokens":74,"cost_in_usd_ticks":71612000,"prompt_tokens_details":{"text_tokens":624,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":2586,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":624,"tokens_out":74,"duration_ms":12823,"temperature":1.0,"reasoning_tokens":2586,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-26T08:38:15.703917+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Running the same refinement and skill-evolution procedures on a fresh collection of enterprise tasks outside the AFTER benchmark and observing no performance gains or transfer benefits.","supporting_citations":[],"review_version":1}