{"id":"34415a72-5dfe-421d-bc4a-02e5089d6785","arxiv_id":"2507.10644","paper_version":4,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"The paper organizes Multi-Agent Systems, the Semantic Web, and LLM-based agents into one narrative in which the location of semantic effort migrated from platform, to data, to model.","lead":"This survey unifies three decades of research on autonomous web agents under a single story: where meaning comes from moved from agent platforms, to annotated web data, to the large language model itself. A generalist should read it for a structured map of why earlier agent-web efforts stalled and what the current wave of protocols and regulations is trying to solve.","discovery_kind":"unification","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The predictive claim is undercut by the paper's own seven 'generation-invariant' challenges: if most open problems persist across all three eras, they cannot follow from the current model-side locus of semantic effort.","rationale":"The paper is a genuinely useful synthesis with transparent methodology, released bibliometric artifacts, and falsifiable 2028 markers; those are real strengths. The central claim, however, is the predictive migration thesis, and the most load-bearing weakness is internal: the paper simultaneously asserts that current open problems follow from the model-side locus (abstract, §5) and that seven generation-invariant challenges persist across all three eras (§6, §6.5). The pairing matrix in Figure 5 shows only Challenge 5 directly paired with Lesson 4; the other six challenges are paired with other lessons and are documented by the paper's own text as having existed in Gen I and Gen II (for example, FIPA's platform-bound IDs, the Semantic Web's missing revenue model, repeated centralised registries). Consequently the existence of these open problems cannot be a consequence of the Gen III locus; only their technological manifestation is. The reader's bibliometric concern is real but secondary: even corrected bibliometrics would not establish the causal direction. The proposed test—a systematic tabulation of whether each challenge was recognised in each era—would settle the scope of the predictive claim. If all seven are present in all eras, the paper should restate the claim as one about the form of persistent problems, which is exactly the CONDITIONAL verdict the reader already reached. No change to the verdict is needed.","tokens_in":911,"tokens_out":998,"duration_ms":105154,"concrete_test":"Systematically tabulate, from the paper's own primary sources and the released bibliography, whether each of the seven generation-invariant challenges (C1–C7) is explicitly recognised as an open problem in each of the three eras (FIPA/MAS, Semantic Web, LLM agents). If all seven challenges appear in all three eras, the central claim must be restated as 'the form of the challenges follows the locus,' not 'the open problems follow from the locus'; the paper's §6.1–6.4 passages already suggest this outcome, so the tabulation settles the scope of the predictive claim.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract and §5 assert that the Gen II→Gen III shift is predictive because 'each generation's failure modes and current open problems follow from where that generation located its semantic effort.' But §6 defines Challenges 1–7 as 'generation-invariant problems' persisting across FIPA, Semantic Web, and LLM eras, and §6.5's pairing matrix (Figure 5) shows only Challenge 5 (non-verifiable tool semantics) traced directly to Lesson 4 (the migration); Challenges 1–4, 6, and 7 are attributed to other lessons or to invariant structural features such as identity persistence, discovery, economic substrate, reputation, consent, and liability. The paper's own evidence confirms this: FIPA-era agents had platform-bound IDs (Challenge 1), the Semantic Web failed for lack of a revenue model (Challenge 3), and UDDI, SPARQL endpoints, and the MCP Registry all recapitulate centralised discovery (Challenge 2). Therefore the existence of these open problems does not follow from the model-side locus; only their technological manifestation does. This internal tension is load-bearing because the paper's novelty rests on the predictive force of the migration thesis; if the locus predicts only the surface form of persistent problems, the headline claim is overstated and the 'predictive' language should be tempered accordingly.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript is a narrative survey of the Web of Agents (WoA) spanning roughly 1990--2026, organized around a central thesis: the locus of 'semantic effort' has migrated chronologically from platform-side coordination (Generation I, FIPA-era MAS), through data-side annotation (Generation II, Semantic Web), to model-side interpretation (Generation III, LLM-based agents). The authors call the Gen II to Gen III transition 'semantics-in-data to semantics-in-models' and claim this shift is predictive, in that each generation's failure modes and current open problems follow from where it located semantic effort. The paper contributes a four-dimensional comparative framework (semantic foundation, communication paradigm, locus of intelligence, discovery mechanism), applies it to sixteen systems, covers the November 2024--August 2026 institutional convergence (AAIF, A2A v1.0, MCP specification, payment-network protocols, EU AI Act, NIST CAISI, International AI Safety Report 2026), and derives seven named lessons paired with seven generation-invariant challenges. It also presents bibliometric evidence from OpenAlex with disclosed coverage caveats and releases query results, raw counts, and scripts in a Zenodo supplement.","tokens_in":41794,"tokens_out":3374,"duration_ms":43316,"significance":"If read as a retrospective organizing narrative, the paper is a valuable synthesis: it connects three research communities that are usually surveyed separately, applies a uniform analytical lens across them, and documents a recent institutional layer that few surveys cover. The transparency of the methodology is a genuine strength: the authors explicitly state that this is not a PRISMA-style systematic review, describe the seed-and-snowball and keyword procedures, report the selection funnel, and release the bibliometric data and scripts. The seven lessons and seven challenges, and especially the paired matrix in Figure 5, provide a usable framework for structuring future work, and the falsifiable markers stated in Lesson 4 are a commendable attempt to make the narrative empirically testable. However, the paper's headline claim that the shift is 'predictive' is not supported by the paper's own evidence, because most of the listed challenges are generation-invariant and are attributed to features other than the locus of semantic effort.","major_comments":[{"comment":"The central claim that 'each generation's failure modes and current open problems follow from where that generation located its semantic effort' is not supported by the paper's own pairing structure. In Figure 5, only Challenge 5 (Non-Verifiable Tool Semantics) is directly paired with Lesson 4 (Semantic-Effort Migration); Challenges 1, 2, 3, 4, 6, and 7 are attributed to other lessons or to generation-invariant structural features such as identity persistence, discovery, economic substrate, reputation, consent, and liability. The paper explicitly calls these 'generation-invariant problems' that persist across FIPA, Semantic Web, and LLM eras, which means their existence cannot follow from the current model-side locus; at most, their technological manifestation can. The 'predictive' language should be tempered to describe an organizing or retrospective thesis, or the paper should provide a per-challenge derivation showing how each problem is a structural consequence of the locus-based trade-offs.","section":"§5 and §6.5 (Figure 5)"},{"comment":"The bibliometric evidence is partially circular with respect to the migration thesis. The primary OpenAlex queries are chosen per generation ('FIPA ACL' OR 'agent communication language' for Gen I, 'semantic web service' for Gen II, 'agentic AI' for Gen III), and the generation boundaries are themselves defined by the assumed locus of semantic effort in Table 3. Consequently, Figure 2 largely traces the labels selected for each era. The disclosed caveats (pre-2002 coverage gap, Gen III undercounting) are welcome, but they are not corrected for; as presented, the figure cannot independently confirm the migration pattern. I recommend either adding robustness checks that use the same query terms across all years or explicitly demoting Figure 2 from 'empirical confirmation' to an illustrative visualization of known trends.","section":"§1.2 and Figure 2"},{"comment":"The 'directional pattern' claim rests on assumptions that are also used to define the generations. Generation I (1995--2005) and Generation II (2001--2012) overlap, and the four comparative dimensions are selected to surface the migration rather than being tested against alternative decompositions. This does not invalidate the framework, but it means the migration is, at present, an interpretive lens rather than a measured empirical law. The manuscript would be strengthened by an explicit statement that the platform-data-model ordering is a historiographic claim, not a consequence of the data, and by acknowledging that the 2028 falsifiable markers in Lesson 4 are the first genuine out-of-sample test of the thesis.","section":"§2, Table 3, and §5"}],"minor_comments":[{"comment":"The acronym 'OW ASP' appears twice and should be 'OWASP'.","section":"§6.3"},{"comment":"The MCP specification date is given as 'version 2025-11-25' but the text elsewhere calls it 'November 2025 specification'; please use one notation consistently.","section":"§4.2"},{"comment":"For SWE-agent and Devin, the Discovery cell reads 'N/A (single-repo)', which could be misread as 'not applicable' in the sense of 'no data'; consider writing 'Not applicable (operates within a fixed repository)' to avoid ambiguity.","section":"Table 4"},{"comment":"The statement 'the loss it incurs becomes the dominant unresolved problem of the next phase' uses 'dominant' without supporting evidence; consider softening to 'a major unresolved problem' unless the dominance is demonstrated, e.g., via challenge prevalence or adoption-failure analyses.","section":"§5"},{"comment":"Reference [84] (Collabnix blog) is a low-authority source for a claim about Kubernetes orchestration of agentic AI; consider replacing it with a peer-reviewed or official reference, or qualify the claim as an industry practice description.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper is transparent about its narrative methodology and releases its data and scripts, which is commendable. The main revision needed is not to add new evidence but to recalibrate the strength of the claims: the phrase 'predictive' is used in the abstract, introduction, and Section 5, yet the paper's own challenges matrix shows that most open problems are generation-invariant and not derived from the locus of semantic effort. If the authors reframe the contribution as a retrospective synthesis with a plausible directional thesis and explicitly position the 2028 markers as the test of its predictive value, the paper would be publishable as a survey/position paper in an appropriate venue. For a journal with high standards on empirical claims, the current wording would risk being seen as overclaiming."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper is worth reading: it gives the agent-interoperability field a single, clear narrative arc and backs it with real artifacts. The three-generation platform→data→model framing is a useful organizing device, the four-dimensional classification of sixteen systems is applied consistently, and the November 2024–August 2026 institutional layer (AAIF, A2A v1.0, MCP 2025 spec, payment protocols, EU AI Act, NIST, AISR 2026) is genuinely new relative to the seven concurrent surveys they compare against. They also ship the OpenAlex queries, raw counts, and scripts, and they are transparent about the pre-2002 coverage gap, the partial 2026 data point, and the Gen III query under-counting. That is real evidence and it deserves credit.\n\nThe soft spot is the load-bearing 'predictive' claim. The stress-test note is right: §6 defines Challenges 1–7 as generation-invariant problems persisting across FIPA, Semantic Web, and LLM eras, and Figure 5 shows only Challenge 5 traced directly to Lesson 4 (the migration). If most open problems—identity, discovery, economics, reputation, consent, liability—persist across all three generations, they cannot follow from where semantic effort was located. What the locus predicts is the technological surface form, not the existence of the problems. The authors slide between 'follow from' and 'manifest as,' and the strong reading in the abstract and §5 is not supported by their own pairing matrix. The bibliometric evidence is also partly circular: generations are defined by their assumed locus, and the chosen queries ('FIPA ACL', 'semantic web service', 'agentic AI') trace those labels. That makes Figure 2 a useful descriptive plot, not an independent confirmation.\n\nThe 2028 falsifiable markers in Lesson 4 are a good start and show the authors understand how to make a narrative checkable. But the markers are future commitments, not current evidence. The prediction of 'semantics-in-verified-contracts' is plausible but speculative, and some 2025–2026 institutional events are press-release sourced and anticipated rather than observed.\n\nNet: this is a solid, well-structured survey with a strong synthesis and honest methodology, but the central thesis needs recalibration. The fix is simple—replace 'predictive' with 'explanatory' or 'characterizing,' and clarify that the locus shapes the manifestation of persistent challenges rather than generating them. I would send this to peer review with a request for revision.\n\nFor you: cite it if you need a compact three-generation frame or a map of the 2024–2026 institutional layer. It is a serious piece of work, not a sloppy one.","headline":"A genuinely useful three-generation survey whose central 'predictive' thesis is overstated; the generation-invariant challenges the authors themselves define undercut the claim that open problems follow from the locus of semantic effort.","tokens_in":42355,"tokens_out":1152,"would_cite":true,"duration_ms":16265,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Three decades of agent-web research form a single directional migration of semantic effort—from platform, to data, to learned model—and that migration predicts each era's failures and its open problems.","keywords":["Web of Agents","Agentic AI","Multi-Agent Systems","Large Language Models","Semantic Web","Agent Interaction Protocols","Model Context Protocol","Agent Governance"],"falsifier":"One concrete observation that would settle whether the central claim is right: by 2028, check whether at least half of publicly registered MCP servers require signed tool manifests and whether reported deployment failures become dominated by contract-enforcement gaps rather than semantic ambiguity. Failing either marker would contradict the paper's prediction of a migration to semantics-in-verified-contracts; additionally, if a Generation III open problem is found that demonstrably stems from data-annotation costs rather than model-side non-verifiability, the predictive link between locus and failure modes would be broken.","tokens_in":41336,"feed_emoji":"🤖","tokens_out":10751,"duration_ms":114175,"temperature":0.7,"pith_summary":"The paper is a narrative survey built around one organizing claim: across three decades, agent-interoperability research has moved the location of meaning in a fixed order, from agent platforms (the FIPA-era multi-agent systems of the 1990s), to externally annotated data (the Semantic Web of the 2000s), to the learned weights of large language models (the current agentic-AI era). The authors call the middle-to-latest step the semantics-in-data $\\rightarrow$ semantics-in-models shift and argue it is predictive rather than merely descriptive: each generation's adoption failures and its unresolved open problems follow from where that generation put its semantic effort. If the claim holds, today's pain points—non-verifiable tool descriptions, ephemeral agent identity, centralised discovery, consent propagation across delegation chains, and a liability vacuum—are structural costs of the current model-side choice, not incidental engineering gaps. The survey supports the thesis with a four-dimensional comparative framework applied to sixteen systems, a bibliometric time series of publication share across the three generations, and coverage of the 2024–2026 institutional wave of protocols, payment-network standards, and regulation. A sympathetic reader would care because the framework turns a scattered history into a testable prediction: the next migration should restore verifiability, in a phase the paper calls semantics-in-verified-contracts.","feed_headline":"30 years of agent research follow one path: platform to data to model","feed_subtitle":"A survey argues each era's failures follow from where meaning lived—and predicts the next phase restores verifiability.","key_machinery":"The load-bearing machinery is the four-dimensional comparative framework (semantic foundation, communication paradigm, locus of intelligence, discovery mechanism), used to classify sixteen representative systems from 1995 to 2026. Its four dimensions are chosen to expose whether a generation's architecture is internally aligned; the paper's Lesson 1 states that adoption at open-web scale requires mutual compatibility of all four dimensions, and a mismatch on any one is enough to prevent it. The same framework also carries the central identity of the paper, the semantic-effort migration: platform $\\rightarrow$ data $\\rightarrow$ model, made concrete in a three-lane chronology and a publication-share series. It does the work of turning 'where the meaning lives' into a measurable, comparable variable, and it grounds the prediction that the next migration will run toward semantics-in-verified-contracts.","core_discovery":"On the paper's own terms, the discovery is that the Web of Agents has not been three separate research programmes (multi-agent systems, the Semantic Web, LLM agents) but one continuous project whose defining variable is the locus of semantic effort. Generation I (FIPA-era MAS, roughly 1995–2005) encoded meaning in the platform: performative speech acts, belief-desire-intention logic, and platform-side coordination services. Generation II (the Semantic Web, roughly 2001–2012) moved meaning into external data structures—RDF/OWL knowledge graphs—so that comparatively simple agents could reason over annotated graphs. Generation III (LLM-based agents, 2020s onward) moved meaning into the model: data is left largely unannotated and a large language model interprets natural-language tool descriptions at runtime. The authors claim this trajectory is directional and predictive: the loss each migration incurred becomes the dominant unresolved problem of the next phase, so the current era's inability to verify what a model thinks a tool does is the direct price of semantics-in-models. They further predict that the next phase, semantics-in-verified-contracts, will trade some flexibility back for verifiability through signed tool manifests, cryptographically verifiable agent cards, and runtime pre/post-condition checks.","pith_inferences":["If the migration is genuinely directional, the same logic suggests that improvements in model capability will push agents toward interpreting human-facing interfaces rather than structured APIs, deepening the verifiability bottleneck and making runtime contract enforcement a central research area rather than a peripheral one.","The framework is directly transferable as an evaluation tool: any new agent protocol can be scored on the four dimensions, and its open-web prospects judged by whether it repeats a known mismatch (for example, decentralised discovery without persistent identity).","The bibliometric evidence should be read cautiously: 'agentic AI' only became common terminology in late 2023, so part of the Generation III publication peak reflects naming, and the true migration pattern depends on the triangulation queries; a reader should weight the qualitative chronology more heavily than the raw curve.","A testable extension of the paper's own falsification markers: track whether reported deployment failures shift from 'the model misunderstood the tool' to 'the model violated an enforced contract'; that shift is the observable signature of the predicted fourth phase."],"forward_implications":["If the migration thesis is right, the current generation's open problems are structural, not incidental: non-verifiable tool semantics, ephemeral identity, centralised discovery, cross-boundary consent asymmetry, and liability gaps will not disappear just from better protocol design.","Protocols that align on three dimensions but fail on a fourth are unlikely to survive at open-web scale; the prediction applies to current and future agent-interoperability protocols.","Centralised discovery will keep recapitulating the same failure: the current registries (the MCP Registry, agent-card hosting) will face the same fate as the FIPA Directory Facilitator and UDDI unless discovery becomes structurally decentralised.","The first credible economic substrate for agent commerce—provided by incumbent payment networks entering the space—makes the other pillars (security, trust, governance) negotiable for the first time in three decades.","The next phase, semantics-in-verified-contracts, is already visible in early markers; if it consolidates by 2028, model flexibility will be constrained by machine-checkable contracts without losing the large language model's adaptive reach."],"supporting_citations":[{"why":"It supplies the Generation II vision of meaning placed in machine-readable data so that agents can reason over the open web.","marker":"[13]"},{"why":"It defines the FIPA ACL performative semantics, the platform-side semantic foundation of Generation I.","marker":"[32]"},{"why":"It documents the FIPA-era failure to reach open-web scale through the architectural mismatch between agent platforms and the web's RESTful substrate.","marker":"[19]"},{"why":"It records the Semantic Web's acknowledged unrealised state, grounding the incentive-failure account of semantics-in-data.","marker":"[49]"},{"why":"It defines the Model Context Protocol, the Generation III agent–tool interface with implicit, model-interpreted semantics.","marker":"[10]"},{"why":"It defines the Agent-to-Agent protocol and Agent Card discovery, forming the Generation III agent–agent channel.","marker":"[11]"},{"why":"It supplies the MCP November 2025 specification (Tasks, OAuth 2.1, Registry), a marker of the predicted verified-contract phase.","marker":"[89]"},{"why":"It supplies the A2A v1.0 specification with signed Agent Cards, the other early marker of the predicted fourth phase.","marker":"[96]"}],"fun_headline_variants":["Agent research: one 30-year arc, from platform smarts to model smarts","Semantics moved from data to models—what that predicts next","Web of Agents: a 30-year migration of meaning, now in models","From knowledge graphs to LLMs: the same web, new locus of smarts","Three eras, one variable: where meaning lives determines the failures"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the hand-selected bibliometric queries—'FIPA ACL' or 'agent communication language', 'semantic web service', and 'agentic AI'—and the qualitative reading of the historical record genuinely track where each era located its semantic effort; if those proxies are biased, the empirical confirmation of the migration pattern weakens.","fun_headline_variants_meta":{"raw":{"variants":["Agent research: one 30-year arc, from platform smarts to model smarts","Semantics moved from data to models—what that predicts next","Web of Agents: a 30-year migration of meaning, now in models","From knowledge graphs to LLMs: the same web, new locus of smarts","Three eras, one variable: where meaning lives determines the failures"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000247,"raw_usage":{"total_tokens":1656,"prompt_tokens":1174,"completion_tokens":482,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":790,"completion_tokens_details":{"reasoning_tokens":384}},"tokens_in":790,"tokens_out":482,"duration_ms":5379,"temperature":1.0,"reasoning_tokens":384,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T17:29:47.861042+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"One concrete observation that would settle whether the central claim is right: by 2028, check whether at least half of publicly registered MCP servers require signed tool manifests and whether reported deployment failures become dominated by contract-enforcement gaps rather than semantic ambiguity. Failing either marker would contradict the paper's prediction of a migration to semantics-in-verified-contracts; additionally, if a Generation III open problem is found that demonstrably stems from data-annotation costs rather than model-side non-verifiability, the predictive link between locus and failure modes would be broken.","supporting_citations":[],"review_version":1}