{"id":"6a7deeda-9164-4aa8-9a28-84906d2132e0","arxiv_id":"2507.19593","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"This systematic review of 44 agent-compatible hypergame papers finds that practical applications simplify the theory, graph-based and hierarchical models dominate deceptive reasoning, and no formal hypergame language for agents exists.","lead":"This paper reviews 44 studies that apply hypergame theory, a framework for modeling how agents see different games, to multi-agent systems. It finds that most practical implementations simplify the theory and that a shared formal language for describing hypergames is still missing.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The agent-compatibility filter in §3.1 is correlated with model type, so the §4.1 prevalence claims (graph-based dominance, HNF rarity) may partly be selection artifacts; independent corpus reconstruction is needed before the roadmap is treated as established.","rationale":"The reader correctly identified corpus representativeness as the weakest assumption. I sharpen that concern: the risk is not merely that the single-query search misses papers, but that the agent-compatibility inclusion criterion is correlated with the very model types whose prevalence the survey reports. HNF-based methods are typically formulated as analytical decision aids, which the filter tends to exclude, while graph-based and hierarchical methods are more naturally used in agent controllers. This makes the headline prevalence claims selection-sensitive. The paper is otherwise well-structured: the formal introduction is useful, the classification tables are detailed, and the publication-trend analysis is plausibly correct. The numerical inconsistencies (49 vs. 44 papers, and minor funnel-count discrepancies between text and Figure 1) are real but secondary. Because the central claims are conditional on a reproducible corpus that is not yet demonstrated, the existing CONDITIONAL verdict is appropriate. I do not see evidence of fatal error, so I would not move the verdict to REJECT; an independent replication could either confirm the roadmap or force a refinement of the prevalence claims.","tokens_in":24172,"tokens_out":6766,"duration_ms":78915,"concrete_test":"Run an independent corpus reconstruction: (1) query Scopus, Web of Science, and DBLP for 'hypergame*' AND ('agent*' OR 'multi-agent' OR 'MAS'), supplementing the Google Scholar results with backward and forward citation chaining from the 44 included papers and the 6 identified surveys; (2) have two independent annotators apply the stated agent-compatibility criterion to the candidate set, reporting inter-annotator agreement; (3) recompute the key distributions in Figures 4, 9, 10, and 11, including model-type counts, domain-model associations, and HNF share. If HNF share rises above roughly 20% or graph-based dominance disappears outside cybersecurity, the §4.1 roadmap would need re-grounding; if the distributions replicate, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central descriptive claims in §4.1—prevalence of hierarchical and graph-based models in deceptive reasoning, simplification in practical applications, limited adoption of HNF, and the absence of formal hypergame languages—are generalizations over a corpus assembled by a single Google Scholar query for \"hypergame theory\" and then filtered by manual relevance and agent-compatibility judgments (§3.1). The load-bearing step is the agent-compatibility filter, which removed 69 of 113 candidate papers because they used hypergames \"solely from an analyst perspective.\" This is not a neutral sampling step: HNF-based models are naturally analyst-oriented decision aids, whereas graph-based and hierarchical formalisms are more readily integrated into reactive agent controllers, e.g., through LTL-style strategy synthesis in the Kulkarni and Fu line of work. If the filter disproportionately removes HNF and analyst-oriented papers while retaining graph and hierarchical papers, then the observed \"limited HNF adoption\" and \"graph-based dominance\" are partly consequences of the inclusion rule rather than properties of the literature. The same concern applies to the claimed absence of formal hypergame languages: only 44 pre-selected papers were examined, and the paper itself acknowledges HML and HAT in §4.2, so absence in this corpus is weak evidence. The 49-versus-44 corpus count discrepancy between the abstract and the full text adds a reproducibility concern, but it is not the central weakness. The argument is not internally inconsistent, but its support is method-sensitive.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript presents a systematic review of hypergame theory applications in multi-agent systems (MAS). It provides a formal introduction to hypergame theory, hierarchical hypergames, and the Hypergame Normal Form (HNF), then defines agent-compatibility criteria and applies them to build a corpus of 44 selected papers (49 in the arXiv metadata abstract). The authors classify each paper by domain, sub-domain, hypergame formalism, integration fidelity, and computational task, and report trends: hierarchical and graph-based models dominate deceptive-reasoning applications, practical deployments simplify the theoretical frameworks, HNF adoption is limited, and no agent-based hypergame modeling language has emerged. The paper concludes with a roadmap for hypergame-based MAS research, including formal languages and human-agent misalignment.","tokens_in":24469,"tokens_out":9584,"duration_ms":101962,"significance":"The survey fills a genuine gap: no prior systematic review examines hypergame theory from an MAS/agent perspective. The detailed classification tables (Tables 3 and 4) and the author-level co-authorship analysis are useful resources, and the explicit selection funnel in Section 3.1 is a strength for reproducibility. If the corpus is representative, the observed tendencies and gaps provide actionable guidance for researchers. However, because the central claims are prevalence claims over a manually filtered corpus from a single Google Scholar query, the value of the roadmap depends on the robustness of the corpus construction; the current manuscript does not yet establish that robustness.","major_comments":[{"comment":"The agent-compatibility filter is not neutral with respect to model type, and this threatens the central prevalence claims. Section 3.1 excludes 69 of 113 papers because they used hypergames 'solely from an analyst perspective,' but HNF and hierarchical hypergames were originally developed as post-hoc analytical frameworks (Vane and Lehner, 2000; Wang et al., 1988), whereas the graph-based LTL-synthesis line of work (Kulkarni and Fu, 2019–2024) is designed for reactive agent controllers. Consequently, the observations in §4.1.1 that 'graph-based models in particular stand out as the most popular implementation' and that 'HNF has seen limited use in agentic contexts' may be partly consequences of the inclusion rule rather than properties of the literature. The authors should provide a sensitivity analysis: re-run the prevalence statistics in Figures 4, 9, 10, and 12 on the full 113 relevant papers (including the 69 analyst-perspective papers), or at least report the model-type distribution within the excluded set, and ideally reconstruct an independent corpus using additional search terms such as 'hypergame normal form,' 'hierarchical hypergame,' and 'hypergames on graphs.'","section":"§3.1 and §4.1.1"},{"comment":"The funnel numbers are internally inconsistent, which undermines the reproducibility of the systematic review. The text reports 320 unique results, 30 duplicates removed, 17 non-English papers excluded, 154 papers excluded by relevance filtering, and then 'From the remaining 113 works' 69 papers excluded by the agent-compatibility filter, giving 44. Figure 1, however, reports 29 duplicates, 16 non-English, 153 irrelevant, 119 relevant, 44 agentic, 69 analytical, and 6 surveys; 119 + 198 = 317, not 320. The relationship between the '113 works' in the text and the '119 relevant' in the figure is unexplained, and the handling of the 6 survey papers (retained for comparison but not in the core set) should be made explicit in the funnel. Because the corpus is the evidence base for all prevalence claims, these arithmetic and procedural inconsistencies must be corrected.","section":"§3.1 and Figure 1"},{"comment":"The claim that 'an agent-based hypergame modeling language or simulation platform has not emerged' is presented as a structural gap, but the evidence base is the 44-paper filtered corpus, and the paper itself names HML (Brumley, 2003), HAT (Gibson, 2013), and SPA (Vane, cited in Kovach et al., 2015) in the same section. Absence from this pre-selected, agent-compatible corpus is weak evidence of absence in the broader literature. A targeted search for hypergame languages and tools—for example, querying 'hypergame markup language,' 'HYPANT,' 'hypergame analysis tool,' and 'hypergame normal form' in addition to 'hypergame theory'—should be conducted before concluding that no formal language exists. At minimum, the paper should distinguish 'no language in our corpus' from 'no language in the literature.'","section":"§4.2"}],"minor_comments":[{"comment":"The arXiv metadata abstract says 49 selected studies while the full-text abstract, Section 3.1, and the conclusion all say 44; this discrepancy must be resolved.","section":"Abstract and §3.1"},{"comment":"In the hypergame definition, the relation R_{ij} is defined as a subset of A_i × A_j, but preferences are over joint outcomes, not over an opponent's action set; the rock-paper-scissors example uses ordered pairs of outcomes, so the definition should read R_{ij} ⊆ A_i × A_i.","section":"§2.2"},{"comment":"Section 3.3 reports 7 papers referring to HNF and 2 hybrid papers, while Section 4.1.1 states that 'only 9 works adopt HNF-based solutions'; the paper should clarify whether the 9 includes the hybrid category and make the terminology consistent across Figures 4 and 9.","section":"§3.3 and §4.1.1"},{"comment":"Section 3.4.2 says experimental and practical works each number 11, but Section 4.1.1 refers to 'the 14 works we classified as practical' in connection with Figure 11; these counts should be reconciled.","section":"§3.4.2 and §4.1.1"},{"comment":"The text and the figure give slightly different counts for duplicates (30 vs. 29), non-English papers (17 vs. 16), and irrelevant papers (154 vs. 153); the numbers should be made to match exactly or the rounding should be explained.","section":"§3.1 and Figure 1"},{"comment":"There is a typo in the phrase 'a standard game-theoretic without misaligned perceptions'—the word 'model' or 'game' appears to be missing—and '0 th' should be formatted consistently as '0th'.","section":"§2.3"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a potentially useful survey, but its central claims rest on a corpus constructed by a single Google Scholar query and a manually applied agent-compatibility filter that is correlated with model type. The authors should be asked to provide sensitivity analyses on the excluded papers or an independent corpus, and to fix the internal count inconsistencies. The 49/44 discrepancy in the abstract is a visible reproducibility issue that should be corrected before publication. The two self-citations in the corpus are not a concern, but they should be disclosed transparently if the paper is revised."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Frankly, this is a solid piece of survey work. It's the first review I know that looks at hypergame theory specifically through an agent-compatible MAS lens, and it does two things well: it gives a clear formal recap of Bennett/Wang-style hierarchical hypergames and Vane's HNF, with a running rock-paper-scissors example that actually helps; and it builds a detailed classification—domain, formalization, fidelity, task, dynamics—across 44 papers. The tables are dense but usable. The author-frequency and co-authorship analysis is a nice extra and supports the claim that the field is slowly broadening beyond its original circles.\n\nThe trends the paper extracts—practical deployments simplify the full theoretical machinery, cybersecurity dominates, HNF remains niche, formal agent-oriented languages are missing—are plausible and mostly line up with the literature I know. So the review deserves credit for organizing a scattered subfield.\n\nThe soft spot is corpus construction. Everything comes from one Google Scholar query for \"hypergame theory,\" followed by manual relevance and agent-compatibility screening, with no released screening protocol. The stress-test note about the agent-compatibility filter is fair: removing 69 of 113 papers for using hypergames \"solely from an analyst perspective\" is not a neutral step. HNF and HEU are naturally analyst-facing frameworks, so that filter plausibly trims HNF-heavy work and inflates the apparent dominance of graph-based and hierarchical models. The paper's claim that HNF is 'too rigid for modern agent-based modeling' is an interpretation that goes beyond what the corpus can support. Similarly, the 'lack of formal languages' claim is okay only when stated carefully—they themselves cite HML and HAT, so the real gap is 'no agent-oriented language,' which they do say later.\n\nThe 49-in-the-abstract vs 44-in-the-text count is a genuine inconsistency, probably from an earlier version, but it needs fixing. It doesn't sink the trends, though; the full text is consistently 44, and the tables are built on 44.\n\nWho gets value: grad students entering hypergame theory from MAS, and researchers working on deceptive reasoning, theory of mind, or cyber deception who want a map of where hypergames have and haven't been applied. It deserves a real referee. I'd send it to review with a request for a second database, a sensitivity analysis around the agent-compatibility judgments, and a cleaned-up abstract. The core is worth engaging with.","headline":"A useful first systematic map of agent-compatible hypergame applications, with plausible but method-sensitive trend claims; the 44/49 count and single-engine manually filtered corpus should be fixed before the roadmap is taken as established.","tokens_in":24949,"tokens_out":4301,"would_cite":true,"duration_ms":50274,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Surveying 44 agent-compatible applications, this review argues hypergame theory has settled into a cybersecurity-heavy, graph-based mainstream: practical models simplify the formalism, and no formal hypergame language exists.","keywords":["hypergame theory","multi-agent systems","perceptual games","misaligned perceptions","nested beliefs","theory of mind","deception modeling","systematic review"],"falsifier":"Run the same screening funnel on a multi-database search (Scopus, Web of Science, IEEE Xplore, dblp) using 'hypergame' and variants such as 'perceptual game,' 'subjective game,' and 'misperception,' with forward and backward citation chasing: if this surfaces many agent-compatible applications the single-query corpus missed—especially HNF-based models outside cybersecurity, or an existing agent-oriented hypergame language—then the survey's prevalence and gap claims would be overturned.","tokens_in":23987,"feed_emoji":"🧠","tokens_out":13981,"duration_ms":134940,"temperature":0.7,"pith_summary":"This survey sets out to establish how hypergame theory—game theory in which each agent plays its own subjective 'perceptual game' rather than a shared game—has been turned into a working tool for multi-agent systems. Screening 320 papers down to 44 agent-compatible applications across cybersecurity, robotics, social simulation, communications, and general game theory, it classifies each by domain, model type, integration fidelity, and computational task. Its central finding is that the field shows clear prevailing tendencies: hierarchical and graph-based hypergames dominate deceptive reasoning, practical deployments tend to flatten or simplify the theoretical frameworks, the hypergame normal form has barely spread beyond cybersecurity, and no formal hypergame description language exists. It also finds that human-agent and agent-agent misalignment remain largely unexplored territory. If this assessment is right, it gives researchers a concrete roadmap for moving hypergames from post-hoc conflict analysis into dynamic agent architectures.","feed_headline":"44 studies: hypergame theory thrives in cyber, stalls elsewhere","feed_subtitle":"Graph-based and hierarchical hypergames dominate; the normal form stays rare outside cyber defense.","key_machinery":"The central object is the hypergame, a tuple $H=(N,\\{G_i\\})$ in which each player $i$ acts inside its own perceptual game $G_i=(N_i, A_i, R_i)$—its subjective list of who is playing, what actions exist, and how outcomes are ranked. Two formalisms carry the theory and the survey's taxonomy: hierarchical multi-level hypergames, which nest viewpoints through perceptual functions $f_i: \\Gamma_i \\to \\Gamma_{ij}$ so that a third-level hypergame encodes what $i$ believes $j$ believes about $k$'s game, analyzed with the Hypergame Nash Equilibrium (HNE), a strategy profile that is a Nash equilibrium in every player's subjective game; and the hypergame normal form (HNF), a decision-theoretic matrix in which a row player assigns belief-context probabilities to candidate opponent mixed strategies and evaluates hyperstrategies by hypergame expected utility (HEU), which mixes expected utility with a fear-of-being-outguessed parameter. The survey also distinguishes underperceived, overperceived, and misperceived game components, and caps nested reasoning at level three on the empirical ground that human strategic reasoning seldom goes deeper.","core_discovery":"On the paper's own terms, the discovery is an empirical map of a small field: after a keyword search, deduplication, language filtering, and manual relevance and agent-compatibility screening, 44 papers remain in which hypergames actually shape an agent's reasoning, decisions, or learning. The survey shows that multi-level hypergames dominate (35 papers), with graph-based models the single most common formalization (14 papers), concentrated in cybersecurity attack-defense and defensive deception; only nine works use HNF-based solutions, all inside cybersecurity; reasoning is the dominant task (33 papers); and while 35 of the 44 works are 'complete' integrations of a hypergame model, theoretical papers still outnumber experimental and practical ones combined, and practical systems disproportionately resort to flattened, perceptual, or HEU-based simplifications. From this distribution the paper argues that hypergame theory is being adopted selectively—deeply where deception and nested beliefs are mission-critical, shallowly elsewhere—and that the missing infrastructure (a formal language and dynamic belief-update support) is what keeps it from broader deployment in multi-agent systems.","pith_inferences":["A 'minimal viable hypergame' pattern—one perceptual game per agent plus a belief-update rule—could be distilled from the flattened and perceptual models in the corpus, giving practitioners a standard starting point the survey itself does not extract.","The graph-based dominance suggests a concrete unification test: whether hypergames on graphs with temporal-logic objectives can be re-expressed in HNF belief-context terms, which would give the normal form the scalability it currently lacks.","The single-keyword corpus leaves a checkable opening: an expanded search using terms such as 'subjective game,' 'misperception game,' and 'theory of mind' would show whether HNF's confinement to cybersecurity is a real property of the literature or an artifact of the query.","The human-agent misalignment gap points toward an untried use: hypergame models of AI systems that misperceive human goals would give alignment failures a game-theoretic vocabulary that purely probabilistic frameworks only approximate."],"forward_implications":["Cybersecurity is the proving ground: 24 of the 44 agent-compatible papers sit there, and it is the only domain in which HNF-based models appear.","Practical applications in the corpus benefit most from simplified formalisms—perceptual games, flattened L-th order models, and HEU-based heuristics—rather than from full analytical frameworks.","The absence of a formal hypergame language is a structural gap: existing tools such as HML and HAT are static or single-formalism, and future languages should draw on epistemic logic, GDL-III, and theory-of-mind formalisms.","Graph-based and HEU-based models are the scalable formulations that integrate with learning algorithms, making them the natural carriers for hypergame-based learning.","Planning is the least-used task (4 papers), which the survey attributes to the lack of expressive agent-compatible formalisms, marking planning support as a clear next step."],"supporting_citations":[{"why":"Foundational definition of hypergames as models of conflict built from each player's perceptual game; the object the survey tracks.","marker":"(Bennett, 1980)"},{"why":"The 'Fall of France' hypergame study that anchors the notion of strategic surprise from perceptual asymmetry, used as the running motivation.","marker":"(Bennett and Dando, 1979)"},{"why":"Formalizes hierarchical hypergames and the hypergame Nash equilibrium; basis of the multi-level (MLH) category that dominates the corpus.","marker":"(Wang et al., 1988)"},{"why":"Defines the hypergame normal form and its hyperstrategies; basis of the HNF category whose limited adoption the survey flags.","marker":"(Vane and Lehner, 2000)"},{"why":"Adds hypergame expected utility and the uncertainty parameter g; basis of the HEU-based models in the taxonomy.","marker":"(Vane, 2000)"},{"why":"HYPANT and the hypergame markup language HML; evidence for the survey's claim that early formal languages are static and limited.","marker":"(Brumley, 2003)"},{"why":"The HAT tool with XML-based hypergame representation; further evidence for the missing agent-oriented hypergame language.","marker":"(Gibson, 2013)"},{"why":"The prior broad review of hypergame theory for conflict, misperception, and deception; the comparison baseline for this survey's MAS-specific angle.","marker":"(Kovach et al., 2015)"},{"why":"Cognitive hierarchy model; supplies the empirical claim that human strategic reasoning rarely exceeds three levels, justifying the survey's level-3 cap.","marker":"(Camerer et al., 2004)"}],"fun_headline_variants":["Hypergames: deep in cyber, shallow elsewhere","Survey: hypergame theory thrives in cyber, stalls in other MAS","44 studies show hypergames only fully embraced in cybersecurity","Hypergame theory's selective adoption: cyber deep, rest shallow"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"Everything rests on the assumption that one Google Scholar keyword search for the exact phrase 'hypergame theory,' plus the authors' manual judgments of relevance and 'agent-compatibility,' captured the whole relevant literature—if studies using other terminology or outside that index were missed, the trends and gaps the survey reports could be artifacts of the search rather than facts about the field.","fun_headline_variants_meta":{"raw":{"variants":["Hypergames: deep in cyber, shallow elsewhere","Survey: hypergame theory thrives in cyber, stalls in other MAS","44 studies show hypergames only fully embraced in cybersecurity","Hypergame theory's selective adoption: cyber deep, rest shallow"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000237,"raw_usage":{"total_tokens":1560,"prompt_tokens":1049,"completion_tokens":511,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":665,"completion_tokens_details":{"reasoning_tokens":444}},"tokens_in":665,"tokens_out":511,"duration_ms":5154,"temperature":1.0,"reasoning_tokens":444,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T14:13:50.491710+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same screening funnel on a multi-database search (Scopus, Web of Science, IEEE Xplore, dblp) using 'hypergame' and variants such as 'perceptual game,' 'subjective game,' and 'misperception,' with forward and backward citation chasing: if this surfaces many agent-compatible applications the single-query corpus missed—especially HNF-based models outside cybersecurity, or an existing agent-oriented hypergame language—then the survey's prevalence and gap claims would be overturned.","supporting_citations":[{"cited_title":": Hypergames: Developing a model of conflict","cited_arxiv_id":null,"evidence_quote":"Foundational definition of hypergames as models of conflict built from each player's perceptual game; the object the survey tracks."},{"cited_title":", Hipel , K.W","cited_arxiv_id":null,"evidence_quote":"Formalizes hierarchical hypergames and the hypergame Nash equilibrium; basis of the multi-level (MLH) category that dominates the corpus."},{"cited_title":", Gibson , A.S","cited_arxiv_id":null,"evidence_quote":"The prior broad review of hypergame theory for conflict, misperception, and deception; the comparison baseline for this survey's MAS-specific angle."}],"review_version":1}