{"id":"8e0844b1-1d5e-4aa4-b108-471d12b5c29d","arxiv_id":"2501.05435","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A systematic review of 158 Neuro-Symbolic AI papers finds research concentrated in learning and inference, with explainability, trustworthiness, and Meta-Cognition as underrepresented gaps.","lead":"This paper is a systematic review of Neuro-Symbolic AI research published between 2020 and 2024, based on a PRISMA-style search of five databases. It finds that most work focuses on learning and inference, while explainability, trustworthiness, and especially Meta-Cognition are underrepresented.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Code-availability filter is applied unevenly and directly produces the 5% Meta-Cognition figure; the distributional claim is not robust to this inclusion criterion.","rationale":"The reader's weakest assumption identified representativeness of the PRISMA process, including the codebase filter, as the key vulnerability. I agree with that diagnosis but sharpen it: the filter is not merely unvalidated; it is applied inconsistently by the paper's own account, and that inconsistency directly manufactures the 5% Meta-Cognition figure. This is the most load-bearing concern because the strongest claim is the distributional one, and the most spectacular single number, Meta-Cognition at 5%, becomes zero under the paper's own inclusion rule when the waiver is removed. The concern is a correctness risk, not a disagreement with community consensus: if the filter is restored, the Meta-Cognition count changes dramatically; if the filter is removed, the other category percentages likely shift because code availability correlates with research area (systems and ML papers ship code more than theory and explainability work). The paper does have genuine value: the proposed Meta-Cognition definition is a modest conceptual contribution, and the compiled list of 158 papers with code links is a useful resource for the community. However, there is no formal verification and no external validation of the sampling; the 167/158 count discrepancy and missing PRISMA flow diagram add further uncertainty. My reading does not move the verdict away from the reader's CONDITIONAL; it reinforces the conditions: the authors should publish the full paper list and recompute percentages under alternative, consistent inclusion rules. A non-finding would not be honest here because the central empirical claim is genuinely load-bearing and currently rests on an inclusion criterion that the paper itself applies asymmetrically.","tokens_in":25044,"tokens_out":3219,"duration_ms":30730,"concrete_test":"Obtain the list of 392 candidate papers (or an independently drawn random sample of at least 100) and classify each into the five categories using the paper's definitions. Then compute the category proportions under three rules: (a) all 392 candidates, (b) only papers with a public codebase, and (c) papers with a public codebase but with Meta-Cognition waived as in the paper. If the proportions under (a) and (b) differ by more than 10 percentage points for any category, or if Meta-Cognition's proportion under (b) is 0%, the reported 63/44/35/28/5 distribution is an artifact of the code-availability filter and the central underrepresentation claim is not established. This test also forces a single auditable list that resolves the 167-vs-158 count discrepancy.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central distributional claim (63% learning/inference, 44% knowledge representation, 35% logic/reasoning, 28% explainability/trustworthiness, 5% meta-cognition) rests entirely on the 158-paper sample. That sample is created by a code-availability filter that the paper applies unevenly. Section 2.3 and Section 3 state that 225 of 392 candidate papers were excluded because no public codebase could be found, 'except for entries on Meta-Cognition as no code-bases could be found' for that category. Because the filter is waived for exactly the category the paper declares most underrepresented, the 5% figure is not comparable with the other percentages. If Meta-Cognition papers had been held to the same code criterion, the count would drop from 8 to 0 (the paper states no codebases could be found), making the headline 'gap' an artifact of an inconsistent inclusion rule. Conversely, if code availability is not a meaningful quality filter, then excluding 58% of candidates for lack of code is unjustified and may differentially remove explainability and trustworthiness work, which often does not ship code. Either way, the percentages are not estimates of the research landscape; they are percentages of an inconsistently filtered subset. The paper provides no PRISMA flow diagram, no inter-rater reliability check, and no validation that the code-availability filter preserves the category distribution, so the load-bearing assumption of an unbiased and comparable sample is unsupported.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents a PRISMA-based systematic review of Neuro-Symbolic AI between 2020 and 2024. The authors define a five-area taxonomy (Knowledge Representation, Learning and Inference, Explainability and Trustworthiness, Logic and Reasoning, and Meta-Cognition), screen an initial pool of 1,428 papers down to 158 (or 167, inconsistently reported) included papers, and report the distribution of research across the five areas. They find concentration in Learning and Inference (63%), Knowledge Representation (44%), and Logic and Reasoning (35%), with less work in Explainability and Trustworthiness (28%) and very little in Meta-Cognition (5%). The paper concludes that Meta-Cognition is an underrepresented area and recommends interdisciplinary research to address the gaps.","tokens_in":25343,"tokens_out":5509,"duration_ms":46561,"significance":"If the distributional counts were reliable, this review would be a useful quantitative map of Neuro-Symbolic AI research activity and a clear case for increased attention to Meta-Cognition. The paper also provides a substantial annotated bibliography with links to many repositories, which is a valuable community resource. However, the central quantitative claims are currently undermined by inconsistencies in the sample definition and by an inclusion criterion applied unevenly across categories, so the review's main contribution is not yet established. The paper does offer one genuinely useful contribution: a concrete, well-scoped definition of Meta-Cognition for Neuro-Symbolic AI, along with an explicit summary of open questions in that area.","major_comments":[{"comment":"The number of included papers is internally inconsistent: the Abstract reports 167 papers meeting inclusion criteria, while Section 3 states that after full-text review a further 9 papers were removed, leaving 158 included papers. All reported percentages are computed with denominator 158 (e.g., 70/158 = 44%, 99/158 = 63%), so the abstract should state 158. This is not a cosmetic discrepancy: the entire distributional summary is keyed to a sample size that is never reported consistently.","section":"Abstract and Section 3"},{"comment":"The public-codebase inclusion criterion is applied unevenly. Section 2.3 states that 225 of 392 candidates were excluded because no public codebase could be found, 'except for entries on Meta-Cognition as no code-bases could be found' for that category. Consequently, all 8 Meta-Cognition papers were admitted without the code requirement, while all other categories were subject to it. If the code filter is a quality or reproducibility gate, Meta-Cognition should be held to the same standard (which would reduce its count to 0, not 8); if it is not a meaningful filter, then excluding 58% of candidates on this basis is unjustified and may differentially remove work in areas such as Explainability and Trustworthiness that often does not ship code. Either way, the reported 5% figure is not comparable with the other category percentages, and the gap claim is an artifact of the inclusion rule.","section":"Section 2.3 and Section 3"},{"comment":"The paper claims to follow PRISMA but does not supply a PRISMA flow diagram, the exact search strings used for the five databases, or the detailed screening protocol (title/abstract inclusion rules, deduplication method, and the criteria used to judge 'relevance'). Without these, the candidate set of 392 papers is not reproducible, and the distributional percentages cannot be verified or compared with other systematic reviews. The authors should add the full search strategy and a PRISMA-style flow diagram as supplementary material.","section":"Section 2.3"},{"comment":"The code-availability claims are not supported by the reference list. Many entries do not link to the cited paper's implementation; examples include [32] (URL points to 'symbolic-execution-papers'), [47] (URL points to a chat transcript), [70] and [81] (URLs point to unrelated paper lists), [87] (URL points to 'Autonomous-Agents'), and [96]–[97] (URLs point to general LLM paper collections). If these URLs were used to satisfy the 'public codebase' inclusion criterion, the screening is not credible; if they were not used, the paper should state the actual repository for each included paper.","section":"References and Section 3"}],"minor_comments":[{"comment":"The keyword list contains a doubled comma in 'Logic and Reasoning,,'; the typo should be corrected.","section":"Keywords"},{"comment":"The table header uses 'Ar𝜒iV' instead of 'arXiv', and the row labeled 'Total (after screening)' appears to contain raw per-database search counts rather than screened totals; relabeling would avoid confusion.","section":"Table 1"},{"comment":"The caption uses 'Meta-Level Cognition' while the text and taxonomy use 'Meta-Cognition'; the terminology should be unified.","section":"Figure 2 caption"},{"comment":"The percentage breakdown after deduplication is internally inconsistent: 45% (n=641) duplicates, 28% (n=395) removed at title/abstract, and 28% (n=392) held cannot all be correct percentages of the same total of 1,428, since 395 and 392 are different numbers that round to the same reported 28%.","section":"Section 3"},{"comment":"The phrase 'to improve machine generalization' is awkwardly written as 'the improve machine generalization'; this should be corrected for readability.","section":"Section 4.2"}],"recommendation":"major_revision","confidential_remarks":"The paper appears to be a workshop paper (CEUR Vol-3819) and a master's scholarly paper. The central quantitative claims need to be corrected and re-validated before the review can be relied upon. The most productive path would be to re-run the screening with a consistent code-availability policy (or drop the criterion entirely), provide the PRISMA flow and search strings, and re-verify the repository links for all 158 included papers. If the authors can do that, the review could be a useful resource. The current version, however, does not support the headline gap analysis."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper does three useful things: it compiles a large and genuinely useful reference list (177 entries), it proposes Meta-Cognition as a fifth taxonomy category for neuro-symbolic AI, and it gives a plausible qualitative picture of where the field is concentrating. The meta-cognition definition is a re-framing of established cognitive-science concepts rather than a new discovery, but it is a reasonable addition to the taxonomy and the authors are explicit about borrowing it. The discussion sections read like an honest attempt to organize a fast-moving literature, and the citations to surveys and foundational books are appropriate.\n\nThe soft spot is the sampling. The stress-test note is correct: the code-availability filter is explicitly waived for Meta-Cognition, and the paper itself says no codebases could be found for any entry in that category. So the 5% figure is not comparable with the other percentages. If the same criterion were applied, that count would be 0, and the headline gap becomes an artifact of an inconsistent inclusion rule. There is also a numerical inconsistency: the abstract says 167 papers met inclusion criteria, while Section 3 says 167 were gathered, 9 were removed, leaving 158. That may be a reporting slip, but in a systematic review you need the numbers to line up. The PRISMA flow is incomplete: no flow diagram, no search strings, no inter-rater reliability check, and no validation that the code filter preserves category distributions. The taxonomy is self-defined, which is unavoidable but should be disclosed as a potential source of bias.\n\nNone of this destroys the qualitative conclusion: learning and inference are clearly the dominant focus, and meta-cognition is under-explored. But the specific percentages should be treated as approximate, not as reliable estimates of the research landscape. The authors have done the hard work of reading and categorizing a large corpus; what is missing is disciplined reporting and a consistent inclusion rule.\n\nWho is this for? Anyone looking for a quick map of neuro-symbolic AI and a bibliography to start from. It deserves a serious referee, but not as-is. The right outcome is major revision: fix the count, publish the search protocol and included-paper list, re-run the analysis with a consistent criterion or drop code-availability as a filter, and add a PRISMA flow diagram. I would not cite the distributional claims in their current form, but I would cite the paper for its taxonomy and reference list if those were cleaned up.","headline":"Useful map of neuro-symbolic AI with a sensible taxonomy, but the distributional claims are not robust because the code-availability filter is applied unevenly, and the 5% meta-cognition figure is partly an artifact.","tokens_in":25841,"tokens_out":1651,"would_cite":false,"duration_ms":17163,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Review maps where Neuro-Symbolic AI work is concentrated and where it is missing.","keywords":["Neuro-Symbolic AI","systematic review","PRISMA","meta-cognition","explainability","knowledge representation","learning and inference","logic and reasoning"],"falsifier":"A direct test would be to repeat the identical search and screening on the same five databases and keyword combinations, but include the 225 papers excluded for lacking a public codebase, and compare the category percentages to the reported 63, 44, 35, 28, and 5 percent. If meta-cognition or explainability percentages change materially once the code filter is removed, the claimed gap is an artifact of the filter rather than a property of the literature.","tokens_in":24820,"feed_emoji":"🧠","tokens_out":1610,"duration_ms":17234,"temperature":0.7,"pith_summary":"This review aims to map the post-2020 boom in Neuro-Symbolic AI by systematically collecting and classifying research into five areas: knowledge representation, learning and inference, explainability and trustworthiness, logic and reasoning, and meta-cognition. After screening 1,428 candidate papers down to 158 studied in detail, the authors report that most work sits in learning and inference (63 percent), knowledge representation (44 percent), and logic and reasoning (35 percent), while explainability and trustworthiness gets less attention (28 percent) and meta-cognition almost none (5 percent). The central claim is that this distribution reveals a real gap: a field that pairs neural learning with symbolic reasoning has not yet built the self-monitoring and trust-building mechanisms needed for reliable, adaptable deployment. A sympathetic reader would care because the paper turns a scattered literature into a quantitative map of where effort is going and where interdisciplinary openings are being left open.","feed_headline":"Neuro-symbolic AI avoids thinking about its own thinking","feed_subtitle":"A 158-paper review finds learning dominates while self-monitoring makes up just 5 percent of the field.","key_machinery":"The load-bearing machinery is a five-category taxonomy of Neuro-Symbolic AI, synthesized from six surveys and four books, used in a PRISMA-style literature screening pipeline. The taxonomy defines knowledge representation, learning and inference, explainability and trustworthiness, logic and reasoning, and a newly proposed meta-cognition category, and the pipeline counts how many of the 158 included papers fall into each category and each pairwise intersection. That counting is what carries the argument: the reported percentages and the claim of underrepresentation are direct outputs of applying the taxonomy to the screened corpus.","core_discovery":"The paper's central claim is that Neuro-Symbolic AI research from 2020 to 2024 is unevenly distributed across five foundational areas, with learning and inference (63 percent), knowledge representation (44 percent), and logic and reasoning (35 percent) dominating, while explainability and trustworthiness (28 percent) and especially meta-cognition (5 percent) are underrepresented. The authors further claim that the few intersection points, such as the single system, AlphaGeometry, that touches all four main areas, show how rare broad integration is, and that the near absence of work combining explainability with the other areas signals a concrete opportunity for interdisciplinary research. They also offer a definition of meta-cognition within Neuro-Symbolic AI as the system's capacity to monitor, evaluate, and adjust its own reasoning and learning processes, and argue that this missing layer is what would let future systems act with the lazy-when-possible, focused-when-required character of human cognition.","pith_inferences":["One testable extension would be to run the same five-category taxonomy on the full 1,428-document pool without the code-availability filter to see whether the 5 percent meta-cognition figure is an artifact of the filter, since the review itself waived that filter for meta-cognition.","A neighboring question the paper leaves implicit is whether the code-availability criterion is a proxy for a younger, more systems-oriented slice of the field; that would change how the percentages should be read as evidence about the whole research community.","The taxonomy's meta-cognition category could be operationalized by checking whether specific mechanisms, such as reward-model self-critique, reflection loops, or architecture-level monitors, appear in the literature, which would give a more direct measure than keyword counting.","If the gap is real, a practical consequence is that benchmark suites for neuro-symbolic systems should include meta-cognitive tasks, such as detecting and correcting one's own reasoning errors, to measure progress in the area the review identifies as most neglected."],"forward_implications":["If the distributional map is right, funding and research effort in Neuro-Symbolic AI will likely keep flowing into learning, inference, and knowledge representation while explainability remains a secondary concern.","The low 5 percent figure for meta-cognition implies that self-monitoring and self-regulating architectures are a nearly open frontier, with room for new frameworks rather than incremental improvements.","The sparse intersections involving explainability and trustworthiness suggest that adding explanation mechanisms to existing neuro-symbolic systems is a low-competition, high-need direction, especially for real-world deployment.","If meta-cognition is genuinely absent, a field aiming at reliable autonomy cannot get there by scaling up current learning and reasoning work alone; a separate control layer will be needed.","The single project sitting at the intersection of all four main areas, AlphaGeometry, indicates that cross-area integration is possible but currently exceptional rather than routine."],"supporting_citations":[{"why":"One of the six survey papers whose review of Neuro-Symbolic AI taxonomies supplies the basis for the paper's own five-category taxonomy.","marker":"[20]"},{"why":"A survey of neural-symbolic learning systems that supports the learning and inference category and its framing.","marker":"[21]"},{"why":"A survey of cognitive Neuro-Symbolic AI that contributes to the taxonomy and the cognitive framing, including links to meta-cognition.","marker":"[22]"},{"why":"A survey of statistical relational to neurosymbolic AI that grounds the logic and reasoning and explainability categories.","marker":"[23]"},{"why":"A systematic review of Neuro-Symbolic methods for trustworthy AI that directly grounds the explainability and trustworthiness category.","marker":"[24]"},{"why":"A survey of applications of neurosymbolic AI that helps define the scope and the search terminology for the review.","marker":"[25]"},{"why":"AlphaGeometry, identified as the only project lying at the intersection of all four main research areas, is the load-bearing example for the claim that broad integration is rare.","marker":"[29]"},{"why":"The Garcez and Lamb definition of Neuro-Symbolic AI that the paper adopts as the working definition for inclusion and classification.","marker":"[18]"},{"why":"Kahneman's System 1 and System 2 framing, which the paper uses to motivate the need for a meta-cognitive control layer above fast and slow processing.","marker":"[19]"}],"fun_headline_variants":["Neuro-symbolic AI: 5% self-reflection, 63% learning","Meta-cognition: the gaping hole in neuro-symbolic AI","Self-monitoring: just 5% of neuro-symbolic research","Neuro-symbolic AI: smart but self-unaware"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The reported percentages and gap analysis depend on the assumption that the PRISMA-style search, the five self-defined categories, and the filter requiring a public codebase produced a representative sample of the Neuro-Symbolic AI literature, an assumption the paper does not validate with a flow diagram, inter-rater reliability checks, or a test of whether the code filter distorts the category counts.","fun_headline_variants_meta":{"raw":{"variants":["Neuro-symbolic AI: 5% self-reflection, 63% learning","Meta-cognition: the gaping hole in neuro-symbolic AI","Self-monitoring: just 5% of neuro-symbolic research","Neuro-symbolic AI: smart but self-unaware"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000592,"raw_usage":{"total_tokens":2809,"prompt_tokens":1012,"completion_tokens":1797,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":628,"completion_tokens_details":{"reasoning_tokens":1719}},"tokens_in":628,"tokens_out":1797,"duration_ms":15115,"temperature":1.0,"reasoning_tokens":1719,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T21:13:39.317132+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A direct test would be to repeat the identical search and screening on the same five databases and keyword combinations, but include the 225 papers excluded for lacking a public codebase, and compare the category percentages to the reported 63, 44, 35, 28, and 5 percent. If meta-cognition or explainability percentages change materially once the code filter is removed, the claimed gap is an artifact of the filter rather than a property of the literature.","supporting_citations":[{"cited_title":"Survey on Applications of Neurosymbolic Artificial Intelligence","cited_arxiv_id":"2209.12618","evidence_quote":"A survey of applications of neurosymbolic AI that helps define the scope and the search terminology for the review."}],"review_version":1}