{"id":"ea918ff1-2634-4398-b00a-369c5af616d3","arxiv_id":"2506.22893","paper_version":1,"verdict":"UNVERDICTED","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"Enterprise AI should be reorganized into a user-centric market of specialized agents guided by six tenets rather than built around general-purpose assistants.","lead":"This paper argues that enterprise AI should shift from requiring users to adapt to models toward AI that adapts to users, delivered through task-specific agents on market-based platforms. It proposes six design tenets for such user-centric agentic systems and discusses where current large language models fall short for strategic decisions.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Tenet 4's incentive-compatibility claim is asserted, not derived: with AA8 allowing capability misrepresentation and Tenet 6 conceding imperfect verifiability, no mechanism guarantees truthful disclosure or quality-diverse exit.","rationale":"The reader's weakest assumption identifies exactly the same load-bearing premise: the market mechanism must align incentives so that users get agency, privacy, and safety while underperforming agents exit. My concern sharpens this: the paper's own assumptions and caveats create a specific internal tension between AA8 (agents may communicate capabilities truthfully or not), PA4 (planner maximizes own utility), and Tenet 6's admission of overclaiming and imperfect verifiability. Section 3.2.4 asserts that disclosure and feedback render rewards incentive-compatible, but no formal argument or example shows how. The cited mechanism-design literature does not automatically transfer because the paper acknowledges the agentic setting differs from ad-bidding in important ways (direct user-agent communication, private user information, safety obligations, entry/exit frictions).\n\nThis is not a reason to reject the paper outright. It is a position paper proposing a research agenda, and it explicitly lists open questions in the conclusion. The verdict of UNVERDICTED remains appropriate: the paper has no testable claim to accept or reject, and the missing mechanism is exactly the kind of gap that keeps it unverified. My concrete test would convert the existence claim into a tractable mechanism-design question: if a truthful equilibrium cannot exist even in a minimal one-shot model, the central recommendation would need significant revision; if it can, the authors would have a starting point for a real contribution. Either way, the verdict does not change because the paper as written does not supply the needed analysis.","tokens_in":11900,"tokens_out":2731,"duration_ms":32865,"concrete_test":"Formalize the minimal game suggested by Section 2 and Section 4: each agent has a private capability type θ ∈ {H, L}; the agent sends a public message m about its capability; the planner sets a reward r(m, s, f) based on m, a noisy performance signal s, and user feedback f; users choose an agent based on m and receive a payoff that depends on θ. Check whether there exists any reward scheme r such that truthful reporting (m = θ) is a Bayesian Nash equilibrium, and in which the planner's own utility maximization (PA4) is consistent with users' utility maximization (UA5). If no such r exists even in this one-shot setting, the paper's assertion that disclosure plus feedback 'renders incentive compatible reward' fails in the simplest model.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's strongest claim is that the shift to User-Centric AI is 'attainable through task-specific Agents and their efficient and effective organization,' specifically via a 'Walled-Garden Platform with agent autonomy, governed by a market mechanism.' The load-bearing premise is that such a market can align incentives so that agents disclose capabilities truthfully, users get agency and safety, and underperforming agents exit. This premise is never modeled.\n\nSection 2 sets up the necessary ingredients but not the mechanism: UA5 users maximize utility, AA6 agents receive rewards set by the platform, AA8 agents 'communicate their capabilities publicly, whether truthful or not,' PA4 the planner maximizes its own utility, and PA6 the planner is responsible for safety. Section 3.2.4 claims that public disclosure plus user feedback 'renders incentive compatible reward to the agent creator / developer to improve the agent.' But Section 3.2.6 itself concedes the 'proclivity of agents (developers) to overclaim their capability, along with planner's imperfect verifiability of claims.' That is a direct tension: if agents can misreport and the planner cannot verify, the reward scheme described in Section 4—'agents are rewarded based on their own claim of performance, planner's observance of their performance and users' feedback'—does not by itself ensure truthfulness. Standard mechanism design requires explicit payment rules, allocation rules, or verification structures that satisfy incentive-compatibility and individual-rationality constraints; the paper cites Akerlof, Hart, and Myerson but does not instantiate any such structure for the agentic setting, and it even notes that ad-bidding 'differs significantly from an agentic platform.'\n\nThe central recommendation therefore rests on an unproven existence claim: that some market mechanism can simultaneously deliver truthful disclosure, quality-price diversity, low entry/exit barriers, user agency, privacy, and safety.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This position paper argues that enterprises should move from an 'AI-Centric User' paradigm, where users adapt to inflexible AI, to a 'User-Centric AI' paradigm in which AI and agents adapt to users, their tasks, workflows, and decision-making contexts. The authors define three entities (users, agents, platforms) with explicit assumptions UA1-UA6, AA1-AA8, and PA1-PA6, propose six tenets (process orientation, forward thinking, locally privacy-preserving learning, a market-mechanism platform, risk-reward and quality-price diversity, and low entry/exit barriers), and illustrate the ideas with a four-step data-analytics workflow. The paper's main proposal is a walled-garden platform with autonomous agents governed by a market mechanism, in which agents are rewarded based on self-reported performance, planner observation, and user feedback, and underperforming agents exit.","tokens_in":12276,"tokens_out":4098,"duration_ms":41598,"significance":"If the framework were made precise, it could contribute a useful conceptual foundation for agentic AI in enterprises, especially in connecting user agency, privacy, and agent governance to market-mechanism thinking from economics. The paper deserves credit for making its assumptions explicit, for identifying capability overclaiming and imperfect verifiability as central problems, and for presenting a concrete workflow as a running example. At the same time, the contribution is currently a framing and a set of postulates rather than a validated design: no mechanism, equilibrium concept, or incentive-compatibility condition is specified, and the paper explicitly leaves the main implementation questions open in its conclusion. The significance is therefore conditional on later formalization and empirical evaluation.","major_comments":[{"comment":"The central claim that a market mechanism makes the agent platform incentive-compatible is asserted rather than derived. In §3.2.4 the paper states that public disclosure plus user feedback 'renders incentive compatible reward to the agent creator / developer to improve the agent,' and §4 proposes that agents 'are rewarded based on their own claim of performance, planner's observance of their performance and users' feedback.' However, AA8 explicitly permits agents to communicate capabilities 'whether truthful or not,' and §3.2.6 concedes the 'proclivity of agents (developers) to overclaim their capability, along with planner's imperfect verifiability of claims.' No payment rule, allocation rule, or verification structure is given that would make truthful disclosure and quality-diverse exit an equilibrium. Please either provide a concrete mechanism sketch with an incentive-compatibility argument, or explicitly re-scope this claim as an open conjecture rather than a technical result.","section":"§3.2.4, §3.2.6, §4"},{"comment":"The six tenets are largely entailed by the assumptions in Section 2 rather than independently established. For example, Tenet 5 (risk-reward and quality-price diversity) follows directly from UA3-UA5 combined with AA6, and Tenet 3 follows from UA6 together with AA6-AA8; the rationales in §3.2 mostly re-cite these assumptions. For a framework paper this is acceptable if the assumptions are presented as normative postulates, but the current text presents the tenets as conclusions with rationales. Please state explicitly that the tenets are postulates derived from the premises and that their adequacy is to be tested by future empirical and simulation work, as the conclusion already suggests.","section":"§2, §3"},{"comment":"Section 4 argues that LLMs fall short on strategic decision-making and that User-Centric AI agents would fill the identified gaps, but the paper does not provide evidence that process-oriented, forward-thinking, market-governed agents would actually address the 'severe gaps' in goals, judgment, subjectivity, private knowledge, and environment. The examples A-C are illustrative, not demonstrative. As written, the claim that the shift is 'attainable' (Introduction and §4) is stronger than what the paper supports; please soften the wording to a research agenda or add design-level support for how the proposed tenets close the gaps.","section":"§4"}],"minor_comments":[{"comment":"The phrase 'and to complement (I) and (III)' appears to be a typo; it should likely read 'and to complement (I) and (II).'","section":"§3.1"},{"comment":"Figure 1 is referenced heavily through points I-VIII, but the figure itself is not reproduced in the provided version; ensure the published version includes the figure and that all labels (I-VIII) are legible and match the caption.","section":"§4, Figure 1"},{"comment":"The article header and footer contain template leftovers that conflict with the current preprint: the copyright line says 2018, the ACM reference format says 2018, and the received/revised dates are 2007/2009, while the arXiv submission is dated June 2025. These should be corrected or removed.","section":"Article metadata"},{"comment":"The key terms 'Walled-Garden Platform' and 'market mechanism' are used as central concepts but are not formally defined; please add concise definitions so the proposal is less ambiguous.","section":"§4"},{"comment":"In the model-selection example, the sentence 'In other cases, the risk may be worthwhile if the senior finds something useful for the future from such a model' is vague; clarify who evaluates the risk and which agent or user acts on it.","section":"§3.2.5"}],"recommendation":"major_revision","confidential_remarks":"This is a position paper whose central proposal depends on an incentive-compatibility claim that is asserted but not formalized. The missing mechanism design is likely addressable in revision, so I do not recommend rejection, but the paper needs substantial strengthening to move from a vision statement to a technical contribution. The template/copyright metadata also appears erroneous and should be fixed before any publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is a position paper, not a research result, and it is honest about that. Its contribution is a coherent synthesis: six tenets for user-centric enterprise agents—process orientation, forward thinking, privacy-preserving local learning, a market platform, risk/quality diversity, and low entry/exit barriers—built on explicit assumptions about users, agents, and a platform planner. The running data-analytics workflow keeps the ideas concrete, and the conclusion lists implementation questions as open research, which is refreshing. What the paper does well: it names a real gap—enterprise agents are mostly designed AI-centric—and gives a shared vocabulary to argue about alternatives. The assumptions are stated cleanly, and the examples illustrate genuine heterogeneity in user agency and task contexts. The soft spot is exactly the one the stress-test flags. Tenet 4 says that public capability disclosure plus user feedback renders incentive-compatible rewards to agent developers, but no mechanism, equilibrium concept, payment rule, or verification structure is specified. The internal tension is real: AA8 allows agents to report capabilities truthfully or not, and Tenet 6 concedes the planner's imperfect verifiability of claims. So the central recommendation—a walled-garden market aligning incentives for truthful disclosure, quality diversity, and low-barrier entry—rests on an unproven existence claim. That is a load-bearing gap, but it is a gap typical of a research agenda, not a fatal empirical error. A related minor issue: the Section 2 assumptions entail the tenets almost by construction—UA5 and PA4 give the market, UA3/UA4 give diversity. So the framework is internally consistent but not independently justified. For a position paper that's acceptable, but the authors should either soften the incentive-compatibility claim or add a stylized model. The citation pattern is fine; the self-citation to prior IUI work is legitimate and relevant. There is no empirical or formal result to verify, so the UNVERDICTED label fits. This paper is a useful agenda-setting document for the enterprise AI community, not a solved result. I would cite it as an example of enterprise-agent framing. A serious editor should send it to peer review if the venue accepts position papers—it deserves referee time and a push to either formalize the mechanism or explicitly mark it as a hypothesis. It should not be desk-rejected.","headline":"A clear position paper that deserves peer review as an agenda-setting synthesis, despite an asserted market mechanism that is never actually modeled.","tokens_in":779,"tokens_out":1623,"would_cite":true,"duration_ms":36837,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that enterprise AI should shift from an 'AI-Centric User' model, where people adapt to inflexible tools, to 'User-Centric AI,' delivered by task-specific agents organized on a market-style platform.","keywords":["User-Centric AI","AI-Centric User","Agentic Enterprise","Enterprise Decision-Making","Multi-Agent Platform","Market Mechanism","Human Agency","Walled-Garden Platform"],"falsifier":"A pilot or simulation would settle it: run the proposed four-step workflow with specialized agents rewarded from their own claims, the planner's observation, and user feedback, then check whether agents with inflated capability claims lose reward over time. If overclaiming agents survive or if expert users withhold data despite the privacy controls, the market-mechanism organization is not incentive-compatible as proposed.","tokens_in":11676,"feed_emoji":"🤖","tokens_out":8189,"duration_ms":84583,"temperature":0.7,"pith_summary":"This paper is a position statement about how enterprises should deploy AI. Its central claim is that the current 'AI-Centric User' paradigm, in which people must craft prompts, learn model quirks, and settle for general-purpose outputs, is the wrong direction, and that enterprises should move to 'User-Centric AI,' where AI adapts to each user's tasks, workflows, and decision-making context. The authors argue this shift is achievable through task-specific AI agents organized on a platform that behaves like a market, with specialized agents, user feedback, and rewards that push weak agents out. They offer six tenets: process orientation, forward thinking, locally privacy-preserving learning, a market-mechanism platform, risk-reward and quality-price diversity, and low entry and exit barriers, illustrated on a four-step enterprise workflow from data preparation to presentation. The stakes are concrete: enterprise AI adoption is described as immature, and the paper claims that organizing AI as an incentive-compatible agent market could move repeated enterprise decision-making toward reliable automation.","feed_headline":"Stop forcing users to adapt to AI; give them agent markets","feed_subtitle":"Task-specific agents in a walled-garden market could automate the decisions users now wrestle from general-purpose LLMs.","key_machinery":"The load-bearing objects are the six tenets and the Walled-Garden Platform with agent autonomy, governed by a market mechanism. The platform is the organizational device: users communicate directly with specialized agents, a planner supervises and manages agents through rewards rather than direct control, and agents are rewarded on the basis of their own claimed performance, the planner's observation, and user feedback, so that underperformers can be exited and new agents can enter cheaply. The market mechanism, borrowed by analogy from ad-bidding and mechanism design, is what is supposed to align the self-interest of agents, users, and the platform. The paper also introduces three primitives: User Agency, Agent Foresight, and User Feedback and Agent Learning, along with a running four-stage Workflow (data preparation, model selection, results evaluation, presentation) whose discrete tasks make it possible to assign a specialized agent to each step. The workflow and the primitives do the argumentative work by turning the abstract slogan 'user-centric AI' into a concrete design space with identifiable points where users and agents can cooperate or conflict.","core_discovery":"On the paper's own terms, the discovery is a re-framing: the missing ingredient in enterprise AI is not more capable models but a user-centric organization of AI delivery. The paper asserts that general-purpose LLMs and current agentic frameworks keep the human in the role of adapting to the machine, and that this explains why GenAI rollouts remain immature and strategic decision-making remains largely unautomated. The proposed alternative is a Walled-Garden Platform with direct user-agent communication, supervised by a planner, in which agents are rewarded from three sources: their own performance claims, the planner's observation, and user feedback. Incentives, not direct control, govern agent survival. On top of this market mechanism, the paper stacks six tenets specifying what user-centric AI must do: emphasize process over outcome, anticipate users' next needs, learn from user feedback with local privacy control, offer risk-reward and quality-price diversity, and keep entry and exit barriers low. The paper also distinguishes tactical decisions, where automation is already proven in trading, revenue management, recommendations, and ad-bidding, from strategic decisions, which require user goals, judgment, private knowledge, and environment, information no LLM has, and argues that user-centric agents, not bigger models or better prompting, are the path to automating those decisions.","pith_inferences":["Beyond the paper's claims: the market-mechanism analogy leaves a concrete design problem open, namely what auction, pricing, or reputation rules make truthful capability disclosure an equilibrium; the paper does not specify them, so the natural next step would be simulated agent markets that test whether truthful disclosure survives.","Beyond the paper's claims: if process-orientation is taken seriously, evaluation of enterprise AI would shift from answer correctness to process quality, such as whether an agent preserves user agency, interjects for clarification judiciously, and improves the user's own skill over time.","Beyond the paper's claims: the walled-garden architecture suggests that the first viable agent markets will be enterprise-internal, with cross-enterprise or open agent markets only appearing after data ownership and leakage controls mature.","Beyond the paper's claims: the distinction between tactical and strategic decisions implies a phased adoption path, automating tactical workflows first, as the paper says is already happening, then extending the same market organization to atomic pieces of strategic decisions before attempting whole strategic problems."],"forward_implications":["If the platform vision is right, enterprise users would stop engineering prompts and instead assemble their own workflows from task-specific agents that already know the surrounding steps.","Rewarding agents from claims plus planner observation plus user feedback would make agent quality a competitive outcome: low-performing agents lose reward and exit, while better agents take their place.","Locally privacy-preserving learning would let expert users feed their private judgment into agents without giving up the knowledge that makes them valuable, increasing the pool of high-quality training signal.","Risk-reward and quality-price diversity would let enterprises match agent behavior to the user's risk appetite, such as choosing an exploratory new model versus a conservative proven one for the same task.","If strategic decisions become automatable with user-in-the-loop agents, enterprises could push more decision-making toward automation without sacrificing human accountability."],"supporting_citations":[{"why":"Supplies the definitions of Agents and Workflow that the paper builds its premise on.","marker":"[27]"},{"why":"Supplies the mechanism-design tradition that Tenet 4 draws on for market platforms.","marker":"[23]"},{"why":"Motivates why quality uncertainty in agent markets matters, the lemons problem behind the market mechanism.","marker":"[2]"},{"why":"Supports treating the market mechanism as an incentive scheme for the planner.","marker":"[11]"},{"why":"Supplies the psychology-of-agency basis for the User Agency primitive.","marker":"[4]"},{"why":"Frames the human-centered AI tradition the paper extends toward enterprise agentic systems.","marker":"[34]"},{"why":"Provides evidence that specialist instruction tuning beats generalists on specific NLP tasks, motivating task-specific agents.","marker":"[33]"},{"why":"Supplies the promotion-versus-prevention distinction used to differentiate user agency in decision-making.","marker":"[15]"},{"why":"Supports the agency-plus-automation design principle behind process orientation.","marker":"[14]"}],"fun_headline_variants":["Forget bigger LLMs, give agents a market to serve users","User-centric AI: let agents compete in a walled garden","Enterprise AI needs agent markets, not more prompting","Stop adapting to AI; let agents adapt to you","Agentic enterprise: incentives over control for AI"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The argument rests on the assumption that a walled-garden market of self-interested agents can be designed so that agents disclose their capabilities truthfully, users choose well, weak agents exit, and enterprise data stays protected, all at once, and the paper asserts this by analogy to ad-bidding and mechanism design without specifying the rules that would guarantee it.","fun_headline_variants_meta":{"raw":{"variants":["Forget bigger LLMs, give agents a market to serve users","User-centric AI: let agents compete in a walled garden","Enterprise AI needs agent markets, not more prompting","Stop adapting to AI; let agents adapt to you","Agentic enterprise: incentives over control for AI"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000179,"raw_usage":{"total_tokens":1306,"prompt_tokens":957,"completion_tokens":349,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":573,"completion_tokens_details":{"reasoning_tokens":270}},"tokens_in":573,"tokens_out":349,"duration_ms":4020,"temperature":1.0,"reasoning_tokens":270,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T21:55:04.830249+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A pilot or simulation would settle it: run the proposed four-step workflow with specialized agents rewarded from their own claims, the planner's observation, and user feedback, then check whether agents with inflated capability claims lose reward over time. If overclaiming agents survive or if expert users withhold data despite the privacy controls, the market-mechanism organization is not incentive-compatible as proposed.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the definitions of Agents and Workflow that the paper builds its premise on."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the mechanism-design tradition that Tenet 4 draws on for market platforms."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Motivates why quality uncertainty in agent markets matters, the lemons problem behind the market mechanism."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supports treating the market mechanism as an incentive scheme for the planner."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Frames the human-centered AI tradition the paper extends toward enterprise agentic systems."},{"cited_title":"Specialist or Generalist? Instruction Tuning for Specific NLP Tasks","cited_arxiv_id":"2310.15326","evidence_quote":"Provides evidence that specialist instruction tuning beats generalists on specific NLP tasks, motivating task-specific agents."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the promotion-versus-prevention distinction used to differentiate user agency in decision-making."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supports the agency-plus-automation design principle behind process orientation."}],"review_version":1}