{"id":"f6e4765a-11c0-4f29-8af5-ddd0bd8c003c","arxiv_id":"2504.20903","paper_version":4,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":7,"one_line_summary":"In a stylized NK/NKC simulation, joint performance is highest when a high-performing human searches first and an optimizing AI refines second, not when AI acts first.","lead":"This paper simulates human-AI collaboration with an NK-style model where humans weight recent decisions more heavily and AI weights all decisions equally. It argues that in sequenced tasks, letting a high-performing human go first and AI refine afterward yields better results than the usual AI-first design.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"H→AI advantage in §4.4 may be an artifact of conditioning on ex post high-performing human while not conditioning the AI-first arm on equally good upstream AI.","rationale":"The reader identifies the linear, non-NK payoff and the seeding of the H-to-AI advantage as the weakest assumption. My concern extends that observation to the formal comparison itself: the H→AI arm is conditioned on ex post high H performance, while the AI→H arm is not symmetrically conditioned. This asymmetry is more immediately load-bearing than the NK labeling because it attacks the internal logic of the central claim before any generalization to real organizations is attempted. The same linearity that the reader flags makes the conditioning especially potent: with a mean-based payoff and a moving-average threshold update, a high-1s seed almost mechanically produces a high-1s downstream string. I do not claim the result is false; rather, the manuscript has not yet shown that the result is a property of task sequencing rather than of favorable selection. The proposed symmetric-conditioning test would settle this directly. Because the reader's verdict is already CONDITIONAL, my concern does not change the verdict but sharpens the condition: the authors should either adopt symmetric conditioning or explicitly frame the claim as applying only after a human is known to be high-performing. This is also separate from the missing code/data and the absent experimental validation, which remain secondary reproducibility issues.","tokens_in":20775,"tokens_out":7073,"duration_ms":83299,"concrete_test":"Re-run the simulation with symmetric conditioning: from the identical generated runs, compute Avg. APO[AI | high-H] for H→AI and Avg. APO[H | high-AI] for AI→H, where high-AI is defined by the same top-decile rule on the first-stage AI payoff; also report unconditional means for both sequences. If the H→AI advantage shrinks or reverses under symmetric conditioning, the Section 4.4 generalization is an artifact of comparing a selected favorable case against an unselected baseline, and the 'AI-first is worse' prescription needs to be restated as 'after a proven high-performing first agent, the second agent should be an optimizer.' If the advantage persists, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central comparison in Section 4.4 is asymmetrically conditioned. The H→AI side is stated as 'Avg. APO [AI| high-quality H]', i.e., downstream AI payoff averaged only over runs in which the first-stage H string already scored high. The AI→H side is not conditioned on an equally favorable first-stage AI realization; it is the maximum over H|AI configurations. Since Section 3.4 defines PO_H and PO_AI as the mean of realized binary states and the AI update rule in Section 3.3 is a uniform moving-average threshold, a 'high-quality' H is, by construction, a string with a high density of 1s, and uniform averaging propagates that density: a seed with mean above 0.5 produces a downstream string with mean above 0.5. The comparison is therefore close to selecting a favorable initial condition on one side and an unbiased/random initial condition on the other. What the simulation shows may be that conditioning on ex post success makes the first agent look good, not that H→AI sequencing is intrinsically superior. Managers do not know ex ante which human runs will be high-performing; the design prescription requires comparing unconditional expected APO under both sequences, or comparing H→AI after a high-H realization with AI→H after an equally high first-stage AI realization.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper develops an agent-based simulation model of sequential human–AI decision-making. The authors distinguish a recency-weighted 'satisficing' human adaptation rule from a uniformly weighted 'optimizing' AI adaptation rule, and vary task scope (N), within-task complexity (K), and cross-agent interdependence (C) across modular and sequenced task structures. The main reported findings are: (i) in modular tasks, AI and human search are substitutes on aggregate, with complementarities peaking at moderately broad AI search and low human task complexity; (ii) in AI-to-H sequenced tasks, complementarity attenuates as C and AI search breadth increase; (iii) in H-to-AI sequenced tasks, joint performance is maximized when a high-performing human initiates search and AI subsequently optimizes, a result framed as contradicting the dominant 'AI-first' design principle; and (iv) a memoryless 'probabilistically delusional' AI can outperform rule-based AI when rescuing low-performing upstream human search. The paper claims a generalizable contingency logic for task division, supported by simulations of a model labeled NK/NKC.","tokens_in":21087,"tokens_out":5084,"duration_ms":55584,"significance":"If the central results were established, the paper would make a useful contribution to the management literature on human–AI collaboration: it formalizes a parsimonious distinction between two memory regimes, draws attention to task sequencing as a design variable, and offers a counterintuitive hypothesis that 'AI follows' can outperform 'AI first.' The proposed role of memoryless random search as an escape mechanism is also an interesting idea. However, the current evidentiary value is limited: the model is not actually an NK/NKC fitness-landscape model, no code or data are provided, no significance tests or confidence intervals appear anywhere, and the headline H-to-AI advantage rests on an asymmetrically conditioned comparison. These issues are central rather than peripheral, so the paper's claims are currently not supported to the standard expected of a simulation-based theoretical contribution.","major_comments":[{"comment":"The model is not an NK/NKC model. Payoff is defined as the simple average of realized binary states (PO_AI = (1/|N_AI|) Σ x_i^AI and PO_H similarly), while K appears only as a window length in a moving-average threshold rule and C as a seed length. There is no random fitness contribution, no epistatic payoff structure, and hence no rugged landscape with local peaks in the Kauffman sense. Consequently, the extensive interpretations in §4.1–4.3 in terms of 'rugged search landscapes,' 'local optima,' and 'local traps' are not implied by the formal model. This is load-bearing because the propositions and managerial implications are expressed in that language. The authors should either introduce a genuine NK payoff function or reframe the paper explicitly as a model of threshold dynamics in a bit-string space, removing the landscape claims.","section":"§3.3–3.4, Eqs. for x_{i+1}^H, x_{i+1}^AI, PO_AI, PO_H"},{"comment":"The central comparison is asymmetrically conditioned. The H-to-AI arm conditions on ex post high-performing human strings, whereas the AI-to-H arm is not conditioned on equally favorable upstream AI strings; it is taken as a maximum over H given AI. Because PO_H and PO_AI are means of bits, and because the AI's uniform averaging rule propagates the mean of the seed sequence, a high-density H seed will mechanically produce a high downstream APO. The reported H-to-AI advantage may therefore reflect selection of favorable initial conditions on one side only, rather than an intrinsic property of the sequence. The authors should compare unconditional expected APO under both sequences, or condition both arms on equally high-performing upstream agents (e.g., the same quantile of first-stage payoff), and report the full distributions. This issue directly affects the paper's main design prescription and must be addressed before the claim is credible.","section":"§4.4, 'Avg. APO [AI| high-quality H] > max APO [H|AI]'"},{"comment":"The paper claims to 'statistically generalize' the H-to-AI advantage, but the manuscript contains no significance tests, confidence intervals, or standard errors for any of the reported pairwise comparisons or heatmap differences. The 'blank cells denote failure of convergence' in Figures 9 and 10 is also unexplained: the reader is not told what convergence criterion failed, how often, or whether the missing cells affect the reported qualitative patterns. Since Proposition 3a and the concluding preference ordering depend on these results, the absence of inferential statistics and convergence diagnostics is a major gap.","section":"§4.4, Fig. 6 and Robustness Checks §5.2–5.3"},{"comment":"Several headline 'mechanisms' are direct consequences of the update equations rather than emergent simulation findings. For example, the statement that uniform memory 'amplifies inherited trajectories' follows immediately from the definition of x_{i+1}^AI as an unweighted average over a window with a 0.5 threshold: if the seed mean exceeds 0.5, the average will tend to stay above 0.5. Similarly, the claim that recency weighting 'corrects locally but is volatile at scale' mirrors the linear recency weights in the H equation. The authors should clarify which results are analytically derivable from the definitions and which genuinely require simulation; otherwise the explanatory contribution is overstated.","section":"§3.3 and §4.3 (mechanisms)"}],"minor_comments":[{"comment":"The reference list heading is misspelled as 'REFRERNCES'; this should be corrected.","section":"References heading"},{"comment":"Notation is inconsistent: the manuscript uses N_AI/N_H in some places and NAI/NH or N^{AI} in others; please standardize the notation for all parameters (N, K, C, PO, APO).","section":"Throughout"},{"comment":"The initialization of the first K states is not specified. The moving-average rules in §3.3 require a starting window, but the text does not state whether it is drawn from Bernoulli(0.5), set to a fixed vector, or initialized differently, nor how this affects results.","section":"§3.2–3.3"},{"comment":"Figure 1, Figure 5, Figure 6, and Figure 7 are referenced but not included in the manuscript text; all figures need to be embedded, and Figure 7 ('AI wastage') requires a definition of 'local peaks discovered' that is currently missing.","section":"Figures"},{"comment":"The abstract promises 'experimental validation (a separate study),' but no citation or description of that study is provided anywhere in the manuscript; either add the reference or remove the claim.","section":"Abstract and §6.4"},{"comment":"The simulation parameter ranges are not fully reported: the text states 1,000 runs but does not give the grid of N, K, and C values, the number of replications per cell, or a parameter table; this should be added for reproducibility.","section":"§3.4"}],"recommendation":"major_revision","confidential_remarks":"The central finding (H-to-AI superiority) appears to be an artifact of conditioning on ex post successful H strings without conditioning the AI-to-H arm on equally successful upstream AI strings. If the authors cannot fix this by re-running matched comparisons, the headline claim will not survive. I also have concerns about the disconnect between the NK/NKC framing and the actual model mechanics, and about the absence of code/data and inferential statistics. The paper's idea is potentially interesting, but the current evidence base is thin. I would recommend requesting a careful revision with the conditioning issue addressed head-on, rather than accepting it in its present form."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nWhat you should know: this paper has a real idea buried in it—that a high-performing human should go first and a uniform-memory AI should refine, rather than the default AI-first sequence—but in its current form the evidence for that idea is much weaker than the paper claims. The model is not an NK landscape, the central comparison is asymmetrically conditioned, and nothing is shipped to check it.\n\nThe genuinely new piece is the memory-regime contrast: H uses recency-weighted adaptation, AI uses uniform weighting. That is a clean, defensible difference, and the paper states it clearly. The \"AI wastage\" concept in AI→H sequences is also a useful way to frame why AI-first can underperform. The prose is organized and propositions are stated explicitly.\n\nNow the soft spots, in rough order of severity.\n\nFirst, the model is labeled NK/NKC but the payoff is the average of binary decision states; K is just a moving-average window length, and C is a seed length. There are no random fitness contributions and no rugged landscape, yet the discussion talks repeatedly about local peaks and escaping basins. That is not a cosmetic naming issue—it changes what the results can mean.\n\nSecond, the headline comparison in Section 4.4 is loaded. \"Avg. APO [AI | high-quality H] > max APO [H|AI]\" conditions the H→AI side on ex post high-performing upstream H, while the AI→H side is not conditioned on equally good upstream AI. And because a high-performing H is just a string with many 1s, and AI's uniform average propagates that density, the advantage may largely be an artifact of starting with a favorable density. The paper needs unconditional comparisons, or symmetric conditioning on equally good first-stage realizations.\n\nThird, the abstract promises an experimental validation \"separate study\" that never appears in the manuscript. No code, no data, no significance tests or confidence intervals, and several sensitivity cells are blank with \"failure of convergence.\" For a simulation paper, that is not enough.\n\nWho is it for: people thinking about task sequencing in human-AI collaboration will find the framing useful, but only as a hypothesis, not as a result. I wouldn't cite it yet. I would send it to review, though—the question is important, the update rules are simple enough to verify, and a revision that fixes the landscape label, conditions symmetrically, and ships the code could make it a solid paper.\n\nRecommendation: engage with it, but push hard for a major revision.","headline":"A real idea about human-first sequencing is buried under an NK mislabel, an asymmetric comparison, and no shared code or data; the paper is worth engaging as a hypothesis, not as an established result.","tokens_in":21599,"tokens_out":2420,"would_cite":false,"duration_ms":25454,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"In sequenced tasks, joint performance peaks when a high-performing human searches first and an optimizing AI refines, beating AI-first sequencing.","keywords":["human-AI collaboration","task sequencing","NK model","satisficing","AI-first design","agent-based simulation","memory regimes"],"falsifier":"Re-run the same H-to-AI versus AI-to-H protocol with true NK payoffs, assigning a random fitness contribution to each bit given its K neighbors, and check whether Avg. APO[AI | high-quality H] > max APO[H | AI] still holds; if the advantage depends on counting 1-bits rather than on landscape structure, it should weaken or reverse.","tokens_in":20500,"feed_emoji":"🤖","tokens_out":8663,"duration_ms":71036,"temperature":0.7,"pith_summary":"This paper claims that in sequenced decision tasks the best division of labor is usually human-first, not AI-first: joint performance peaks when a high-performing human generates the initial solution and an optimizing AI then refines and extends it. The claim matters because it contradicts the dominant design prescription that AI should generate options and humans should select among them. The authors build an agent-based simulation in which the only difference between the two agents is how they remember past decisions - humans weight recent outcomes more heavily, AI weights all outcomes equally - and show that this single difference, combined with task structure (modular versus sequenced, and coupling strength C), produces substitution, complementarity, and rescue regimes. If the simulation is a fair model of organizational search, then effective human-AI collaboration is a matter of task architecture, not of AI capability alone.","feed_headline":"AI should follow a high-performing human, model shows","feed_subtitle":"A bit-string simulation finds task architecture, not AI capability, decides when AI adds value.","key_machinery":"The load-bearing mechanism is a pair of threshold-update rules on binary decision strings. A human satisficing agent sets the next bit to 1 when a recency-weighted moving average of the previous $K_H$ bits reaches 0.5; an AI optimizing agent uses the same threshold on an equal-weighted moving average of the previous $K_{AI}$ bits. $N$ is the number of bits, $K$ is the window length that the paper calls task complexity, $C$ is the number of the first agent's realized bits that seed the second agent's sequence, and the payoff is the fraction of 1-bits realized, averaged over runs. The H-to-AI advantage is carried by the asymmetry $|N_{AI}| > |N_H|$: once the AI exhausts the human's short string, it keeps applying its equal-weight rule to its own outputs, so it compounds whatever fraction of 1s it inherited.","core_discovery":"On the paper's own terms, the central discovery is a formal ordering of task sequences: in interdependent sequenced tasks, the average joint payoff of an H-to-AI sequence with a high-performing upstream human exceeds the maximum joint payoff of any AI-to-H sequence, written as Avg. APO [AI | high-quality H] > max APO [H | AI]. The reason is compounding: a high-quality human's bit string contains many 1-states, and AI's uniform-memory averaging propagates and extends that favorable starting condition; when AI leads, much of its broad search is wasted because a following human can use only a small subset of it (the paper's 'AI wastage'), and under high coupling the human's recency weighting becomes a bias rather than a calibration device. A second discovery concerns low-performing humans: a memory-less, stochastically exploring AI beats a rule-based optimizing AI as a downstream partner, because randomness at scale escapes the local trap that inherited-memory optimization would reinforce.","pith_inferences":["Because the payoff field counts 1-bits rather than measuring true NK fitness, the 'high-performing human' is defined by a high fraction of 1s; a natural extension is to replace the payoff with an NK fitness function and test whether the H-to-AI advantage survives.","The model implies a managerial portfolio rule outside its explicit claims: identify and seed high-quality human judgment before deploying optimization AI, and use broad random AI search only when the human input is weak, rather than always defaulting to AI-generated option menus.","Since $K$ is a temporal moving-average window rather than a structural coupling, the core result may transfer to other sequential pipelines such as prompt engineering or human-in-the-loop fine-tuning, where the same averaging-and-inheritance dynamics apply.","An experimental test of the ordering would be straightforward: compare expert-drafted initial solutions refined by an LLM against LLM-generated options selected by experts, measuring output quality across tasks of varying interdependence."],"forward_implications":["In modular tasks, optimal joint performance occurs when AI searches moderately broadly relative to the human and the human's task has low complexity; joint payoff falls when the AI search space is too narrow or too broad.","In AI-to-H sequences, optimization-satisficing complementarity attenuates as task interdependence $C$ and AI's relative search breadth increase, because human recency weighting turns into bias under strong coupling.","In H-to-AI sequences with a high-performing upstream human, downstream AI optimization maximizes joint payoff across the whole parameter space, outperforming any AI-to-H configuration.","When the upstream human is low-performing, memory-less stochastic AI outperforms rule-based optimizing AI as the downstream agent, acting as an escape mechanism from local traps.","Across all configurations, the paper's preference order is: high-quality human first then AI, then AI first then high-quality human, then random AI rescuing a low-quality human."],"supporting_citations":[{"why":"Supplies the NK adaptive-search tradition of payoff landscapes and local peaks that the model builds on.","marker":"(Levinthal, 1997)"},{"why":"Supplies the binary-variable NK implementation used to represent decision states.","marker":"(Rivkin & Siggelkow, 2003)"},{"why":"Supplies the NKC extension with the cross-agent interdependence parameter C used for sequenced tasks.","marker":"(Ganco, Kapoor, & Lee, 2020)"},{"why":"Supplies the behavioral satisficing and path-dependence foundations of adaptive search.","marker":"(Levinthal & March, 1981)"},{"why":"Defines AI as unboundedly rational and articulates the AI-first design principle the paper contests.","marker":"(Csaszar, 2025)"},{"why":"Provides the 'probabilistically delusional' AI characterization used for memory-less random search.","marker":"(Esanu, 2024)"},{"why":"Supplies the mapping of binary decision states to good/bad payoffs and the bounded-rationality modeling apparatus.","marker":"(Puranam et al., 2015)"}],"fun_headline_variants":["AI best as follower of high-performing humans in sequenced tasks","Task structure, not AI capability, decides when AI adds value","AI follows high performer, or rescues low performer stochastically","Joint payoff peaks when AI follows a strong human lead","AI-first designs lose to human-led sequences, model shows"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The model's payoff is not an NK fitness landscape: performance is just the average fraction of 1-bits on each agent's string, and K is only a window length in a threshold rule, so the conclusions may not carry over to tasks with non-additive, context-dependent payoffs.","fun_headline_variants_meta":{"raw":{"variants":["AI best as follower of high-performing humans in sequenced tasks","Task structure, not AI capability, decides when AI adds value","AI follows high performer, or rescues low performer stochastically","Joint payoff peaks when AI follows a strong human lead","AI-first designs lose to human-led sequences, model shows"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000745,"raw_usage":{"total_tokens":3351,"prompt_tokens":1007,"completion_tokens":2344,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":623,"completion_tokens_details":{"reasoning_tokens":2261}},"tokens_in":623,"tokens_out":2344,"duration_ms":18189,"temperature":1.0,"reasoning_tokens":2261,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T05:16:54.460163+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the same H-to-AI versus AI-to-H protocol with true NK payoffs, assigning a random fitness contribution to each bit given its K neighbors, and check whether Avg. APO[AI | high-quality H] > max APO[H | AI] still holds; if the advantage depends on counting 1-bits rather than on landscape structure, it should weaken or reverse.","supporting_citations":[],"review_version":1}