{"id":"a3320020-406d-4a7c-96e0-a58f384af7c9","arxiv_id":"2505.24758","paper_version":4,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"A broad survey that scores about fifty graph database systems on 46 features and summarizes recent developments in models, query languages, and storage.","lead":"This survey compares dozens of graph database systems on 46 features, including data models, query languages, storage layouts, and operational capabilities. It is a practical reference for choosing a graph database and a map of which features the field currently provides or neglects.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 46-feature comparison is scored against a July 2023 cutoff (Table 15) while the paper is published in February 2026 and claims to survey recent advancements; this staleness undercuts the accuracy of the central comparison.","rationale":"I read the paper as making a survey claim whose concrete deliverable is the 46-feature comparison and the resulting aggregate indices. The authors are transparent about their methodology and even provide a GitHub repository, which is genuine independent support for reproducibility. However, the scoring rubric in Table 15 anchors the evaluation to July 2023, while the manuscript is dated February 2026 and repeatedly frames itself as surveying recent advancements. The reader's stated weakest assumption was that public documentation accurately captures system capabilities; that is adjacent but not identical to my concern. I see the time-anchor as more load-bearing because even if all public documentation is accurate, scores for features like Active development, Trendiness, and GQL support can be outdated by more than two years, and the aggregate indices in Section 3.3 inherit that staleness. This does not destroy the survey's value as a structured snapshot, but it does condition the central recency claim. Since the reader already returned CONDITIONAL and cited the July 2023 window in the rationale, my concern does not move the verdict; it sharpens the condition under which the central claim would be accepted. A concrete re-scoring audit as of late 2025 would settle whether the staleness is material or merely a cosmetic date mismatch.","tokens_in":45735,"tokens_out":4916,"duration_ms":54320,"concrete_test":"Pull the raw feature scores from the linked GitHub repository and re-evaluate each time-sensitive feature (Active development, Trendiness, Documentation up-to-date, SaaS offering, Query Language/standard support) as of a fixed recent date, e.g., 2025-12-01, using release histories, vendor documentation, and Google Trends. Then recompute the aggregate indices and the Table 13/14 percentages. If more than 20% of systems change any time-sensitive score, or any system-level Product, Database, or Data index moves by more than 0.05, the Section 3.3 conclusions and the 'recent survey' claim are not supported as written.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that the survey is comprehensive and recent, and that its 46-feature comparison (Section 3.2) and aggregate indices (Section 3.3) accurately reflect the current capabilities of the included systems. The scoring rubric in Table 15 fixes the evaluation window to July 2023: 'Active development' is scored by updates 'within three months before July 2023', 'Trendiness' by Google Trends for 'July 2022 - July 2023', and other features such as SaaS offering, Live backups, Query Language, and GQL support are not given a collection date at all. The paper is dated 23 Feb 2026, with many references accessed 17-Dec-2025. In a field where GQL was standardized in 2024, SQL/PGQ engines such as Oracle 23ai, DuckDB, and PostgreSQL were evolving through 2024-2025, and several systems changed maintenance status, a comparison scored against mid-2023 information cannot support the abstract's claim of 'recent advancements' and a 'detailed analysis of recent advancements' without re-validation. The authors disclose the cutoff, but disclosure does not cure the mismatch: the aggregate percentages in Section 3.3 (e.g., 60.78% active development, 56.86% open source) are presented as current ecosystem facts, not as a historical snapshot. This is the load-bearing weakness because the entire value of the survey rests on the feature matrix and indices being a reliable, current guide.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This survey paper reviews the graph database landscape. It introduces graph data models (property graphs, RDF, and hybrids), surveys graph query languages (Cypher, GQL, SQL/PGQ, Gremlin, SPARQL, GraphQL, and academic proposals), and describes storage architectures ranging from unstructured BLOB stores to succinct and shared-memory designs. The core contribution is a feature-based comparison of 46 features across 51 graph database systems (Section 3.2), aggregated into Product, Database, and Data dimension indices, followed by an adoption analysis in Section 3.3 and a related-work survey in Section 4. The paper claims to provide a thorough and recent survey that can guide developers and researchers in choosing graph database technologies.","tokens_in":46027,"tokens_out":5754,"duration_ms":65348,"significance":"If the feature matrix and aggregate indices were current and methodologically sound, this would be a useful reference for practitioners and researchers: the feature set is broad, the scoring rubric is explicit, and the GitHub repository provides reproducibility. The paper also covers several systems rarely compared in prior surveys, including research prototypes and industry systems such as ByteGraph, G-Tran, LiveGraph, and ZipG. However, the central value depends on the feature scores accurately reflecting current system capabilities, and that claim is currently undermined by the July 2023 scoring cutoff, undocumented system-selection criteria, and heterogeneous denominators in the aggregate statistics. These issues are fixable, but they affect the paper's main contribution as written.","major_comments":[{"comment":"The feature scoring is anchored to July 2023, while the paper is dated February 2026 and claims to analyze 'recent advancements' and to provide a 'comprehensive' and current survey. In Table 15, 'Active development' is scored by updates 'within three months before July 2023', and 'Trendiness' is measured over 'July 2022 - July 2023'. The abstract and Section 3.3 present the resulting percentages (e.g., 60.78% active development, 56.86% open source) as current ecosystem facts. This mismatch is load-bearing because the field changed materially after the cutoff: GQL was published in 2024, several systems cited in Section 2.2 (Oracle 23ai, DuckPGQ, PostgreSQL, Google Spanner Graph) added SQL/PGQ or GQL support, and maintenance statuses have shifted. The authors disclose the cutoff, but disclosure does not cure the inconsistency with the paper's stated recency. I request either (a) re-collecting the feature scores as of a stated 2025/2026 date and updating Tables 1-14, or (b) explicitly reframing the entire feature comparison and all aggregate claims as a historical snapshot 'as of July 2023' and revising the abstract and Section 3.3 accordingly.","section":"Section 3.2, Table 15; Abstract; Section 3.3"},{"comment":"The survey does not state inclusion or exclusion criteria for the systems it scores, despite the Introduction claiming an 'exhaustive list' of systems. More concretely, some scored entries are explicitly not graph databases: KatanaGraph is described as 'not a graph database in itself' (Section 3.2.1), and ZipG is described as lacking 'traits of a graph database such as transaction support, replication, graphical user interfaces or graph query languages' (Section 3.2.1). These systems nevertheless appear in the feature tables and appear to contribute to the aggregate percentages in Tables 13-14. Including non-DBMS systems in the denominator biases the reported feature-adoption rates downward and weakens the comparison's validity. The paper should either apply a clear inclusion criterion (e.g., a graph database management system or graph store with a public implementation), exclude such systems from the aggregate statistics, or analyze them in a separate category.","section":"Section 3.2, Sections 3.2.1-3.2.4, Tables 9-10"},{"comment":"The percentages in Tables 13-14 are computed over different denominators without reporting sample sizes. For example, 60.78% corresponds to 31/51, 55.17% to 32/58, and 51.22% to 21/41; similar discrepancies appear across rows. Section 3.3 uses these percentages to make 'majority of systems' and 'features of less interest' claims, but without knowing the effective N per feature or how missing values were handled, the comparisons are not reliable. Please report the denominator for each row and specify how 'no information' or 'not applicable' cases are treated in the aggregate indices.","section":"Tables 13-14, Section 3.3"},{"comment":"The paper acknowledges that 'the evaluation is based on information available to public and on scientific papers, if present,' and that proprietary databases may have more features than publicly reported. This is a serious limitation for the interpretation in Section 3.3, where low scores are attributed to features being 'not essential to provide, or more difficult to implement by most of the community.' For proprietary systems, an absent feature score can be a false negative caused by lack of public documentation, not by absence of the feature. The paper should label the scores as 'reported capabilities' and soften the interpretive language, or provide a sensitivity analysis for the subset of open-source systems where the source code can independently verify the scores.","section":"Section 3.2 and Section 3.3"}],"minor_comments":[{"comment":"The Query Language row contains a typo: 'Socre 1 if' should read 'Score 1 if'. Also, given that Section 2.2 discusses the ISO GQL standard released in 2024, the definition of a 'widely accepted query language' should be updated to mention GQL and SQL/PGQ rather than only 'Cypher/Gremlin or any other widely accepted query language'.","section":"Table 15"},{"comment":"The text 'Reactive Multi-database (64.0%)' appears to conflate two separate features; 'Reactive programming' belongs to the Database dimension (Table 13) and 'Multi-database' belongs to the Data dimension (Table 14). Please list them separately.","section":"Section 3.3, paragraph on features at most partially supported"},{"comment":"There are minor wording and spelling issues: 'accross' should be 'across' in Cluster Re-balancing, and 'the last 6 month before July 2023' should be 'the last 6 months before July 2023' in Active development.","section":"Table 15"},{"comment":"The text 'showing some perfomance benefits' contains a typo: 'perfomance' should be 'performance'.","section":"Section 3.2.3, MillenniumDB description"}],"recommendation":"major_revision","confidential_remarks":"The paper has clear value as a broad survey, but the feature comparison's July 2023 cutoff combined with a February 2026 publication date is a central credibility issue. I would not reject the paper outright because the issue is fixable (re-run the scoring or reframe the claims). I would ask the authors to address the system-selection and denominator issues as part of the same revision, since these affect the validity of the aggregate statistics in Section 3.3."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe paper is a serious, broad survey of graph databases with a 46-feature comparison across 51 systems. The organization by storage architecture and the feature taxonomy (Product/Database/Data) is genuinely useful as a reference. The authors also do a solid job covering recent standards (GQL, SQL/PGQ) and include newer engines like Kuzu and MillenniumDB that earlier surveys miss. The related-work section is thorough and fair.\n\nThe soft spot is real: the feature scores are anchored to a July 2023 cutoff (Table 15 explicitly says 'updated within three months before July 2023' for active development, and trendiness is measured July 2022–July 2023), while the paper is v3 dated February 2026 and many references were accessed in December 2025. The aggregate percentages in Section 3.3 (60.78% active development, etc.) are presented as current ecosystem facts, not as a historical snapshot. That mismatch undercuts the abstract's claim to 'recent advancements.' The authors disclose the cutoff, but disclosure doesn't fix the presentation: either update the scores or reframe the survey as explicitly a 2023 snapshot.\n\nA smaller but real error: the Ultipa section attributes the GQL standard to LDBC ('the developing international GQL standard of the LDBC'). GQL is an ISO/IEC standard; LDBC is a benchmark council. This will mislead readers who don't already know.\n\nThe reliance on public documentation without independent testing is a known limitation, and the authors acknowledge it. That's a minor concern for a survey of this scope.\n\nThe paper has no derivations or fitted predictions, so there's no circularity concern. The self-citations are to prior storage-compression work and are not load-bearing.\n\nWho is this for? Practitioners and researchers who want a structured comparison of graph databases and a catalog of features to consider when choosing a system. It's not a research contribution in the sense of a new algorithm or result, but it's a competent and broad survey.\n\nMy verdict: it deserves peer review, but it needs revision. The clearest fix is to either bring the feature scores up to a stated 2025/2026 evaluation date (a lot of work, but doable with a GitHub repo and commit hash) or to explicitly rebrand the survey as a July 2023 snapshot. I'd also fix the GQL/LDBC attribution and a few typos. If the authors do that, the survey is a worthwhile reference.","headline":"Broad and useful survey, but the feature matrix is a July 2023 snapshot presented as current, which needs fixing before it can serve as a reference.","tokens_in":46564,"tokens_out":3144,"would_cite":true,"duration_ms":33035,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A 46-feature comparison of over 50 graph database systems shows which capabilities are common and which are rare.","keywords":["graph databases","property graph model","RDF","graph query languages","GQL","SQL/PGQ","storage architectures","feature comparison"],"falsifier":"Choose any system the matrix scores as absent on a binary feature such as data encryption, granular locking, or automatic updates, then inspect its current official documentation or a running current release. If the feature is demonstrably present in a supported configuration, the survey's score is a false negative and that system's aggregate index is understated; a systematic re-check across all such false negatives would determine whether the reported adoption percentages shift materially.","tokens_in":1554,"feed_emoji":"🕸️","tokens_out":1499,"duration_ms":77735,"temperature":0.7,"pith_summary":"This paper tries to establish a current, feature-level picture of the graph database ecosystem: which systems exist, how they store graphs, what query languages they speak, and which capabilities are actually common. It evaluates dozens of systems against forty-six features grouped into Product, Database, and Data dimensions, scoring each feature on a 0-to-1 scale using public documentation and scientific papers. The point is to give developers and researchers a structured way to match a use case to a technology, and to show where the ecosystem as a whole is strong or weak. A sympathetic reader would take away that core capabilities like transaction support, Linux compatibility, REST APIs, and secondary indexes are now standard, while features like automatic updates, data versioning, granular locking, multiple isolation levels, and data encryption remain rare.","feed_headline":"Graph databases agree on basics, skimp on security","feed_subtitle":"A 46-feature audit of 50-plus systems finds transactions everywhere but encryption, versioning, and granular locking rare.","key_machinery":"The carrying device is a 46-feature evaluation framework divided into three dimensions: Product (adoption and deployment), Database (convenience, distribution, software development), and Data (access, consistency, control, security). Each feature is scored on a discrete 0-to-1 scale (absent, limited, partial, significant, or full support) using public documentation and scientific papers, and each dimension is summarized as a simple average per system. The resulting feature matrices and aggregate indices are what let the paper compare systems and identify which features are widely supported and which are not.","core_discovery":"The paper's central claim is that a feature-based comparison compiled from publicly available information can accurately represent the current graph database landscape and guide system choice. Beyond listing systems, it organizes the field by data model (property graphs, RDF, multi-model, and alternative models), by query language (Cypher, Gremlin, SPARQL, GQL/SQL/PGQ, GraphQL, and academic languages), and by storage architecture (unstructured, linear, non-linear, relational, and advanced compressed or shared-memory designs). Its evaluation yields aggregate scores per system in three dimensions and cross-system adoption rates per feature, showing, for example, that transaction support is fully present in 84% of systems while data encryption is fully present in only 26%.","pith_inferences":["The scores are a snapshot tied to documentation available at review time, so a system's aggregate index could shift if vendors publish previously internal capabilities; the matrix is best treated as an updatable baseline rather than a permanent ranking.","Because the survey deliberately excludes empirical benchmarks, its feature comparison could be combined with standardized workload measurements to separate claims of supporting a feature from demonstrated performance at that feature.","The standardization of GQL and SQL/PGQ is likely to compress the diversity of proprietary query languages over the next few years, and the survey's language taxonomy could serve as a baseline for tracking that convergence.","The finding that security and versioning features are rare suggests an opening for vendors and open-source projects to differentiate, with a testable expectation that encryption, granular locking, and data versioning become more common defaults in the next wave of graph databases."],"forward_implications":["Decision-makers can use the 46-feature matrix as a checklist to shortlist systems by the capabilities their use case actually requires.","The survey identifies features the ecosystem treats as table stakes, such as transactions, Linux support, REST APIs, and secondary indexes, versus differentiators like data versioning, granular locking, multiple isolation levels, and data encryption.","The recent GQL and SQL/PGQ standards are presented as a convergence point, with Oracle Database 23ai cited as the first commercial SQL/PGQ system and several engines already claiming partial GQL coverage.","The storage-architecture taxonomy, spanning BLOB-based, linear, non-linear, relational, succinct, and shared-memory designs, gives system designers a map of the graph-native storage design space.","The four distributed transaction challenges identified in the survey frame why newer systems adopt decentralized or RDMA-based designs rather than traditional centralized coordination."],"supporting_citations":[{"why":"Foundational survey of graph database models that frames the paper's review of property graphs, RDF, and query languages.","marker":"[116]"},{"why":"Taxonomy of graph database data organization and system designs that the paper builds on for its storage-architecture classification.","marker":"[208]"},{"why":"Earlier feature-based evaluation of contemporary graph databases that this survey updates and broadens to 46 features.","marker":"[197]"},{"why":"Survey of RDF stores and SPARQL engines that underpins the paper's coverage of the RDF data model and triple-store systems.","marker":"[18]"},{"why":"Foundations of modern graph query languages that supplies the classification of query features such as path queries and navigational patterns.","marker":"[31]"},{"why":"Source of the four distributed graph transaction challenges the paper lists and a surveyed system whose scores enter the comparison.","marker":"[110]"}],"fun_headline_variants":["Graph DB survey: transactions common, encryption rare","Survey: graph DBs lag on encryption and locking","50+ graph systems audited: security features sparse","Graph databases: 84% do transactions, 26% do encryption"],"cache_read_input_tokens":48640,"weakest_assumption_plain":"The survey's scores rest on the assumption that public documentation, vendor websites, and scientific papers accurately reflect each system's real capabilities, which the authors themselves note may undercount proprietary systems like TAO and ByteGraph whose features can be kept as corporate secrets.","fun_headline_variants_meta":{"raw":{"variants":["Graph DB survey: transactions common, encryption rare","Survey: graph DBs lag on encryption and locking","50+ graph systems audited: security features sparse","Graph databases: 84% do transactions, 26% do encryption"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000564,"raw_usage":{"total_tokens":2663,"prompt_tokens":918,"completion_tokens":1745,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":534,"completion_tokens_details":{"reasoning_tokens":1679}},"tokens_in":534,"tokens_out":1745,"duration_ms":13771,"temperature":1.0,"reasoning_tokens":1679,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T12:13:39.762965+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Choose any system the matrix scores as absent on a binary feature such as data encryption, granular locking, or automatic updates, then inspect its current official documentation or a running current release. If the feature is demonstrably present in a supported configuration, the survey's score is a false negative and that system's aggregate index is understated; a systematic re-check across all such false negatives would determine whether the reported adoption percentages shift materially.","supporting_citations":[{"cited_title":"Evaluation of contemporary graph databases","cited_arxiv_id":null,"evidence_quote":"Earlier feature-based evaluation of contemporary graph databases that this survey updates and broadens to 46 features."},{"cited_title":"A survey of RDF stores & SPARQL engines for querying knowledge graphs","cited_arxiv_id":null,"evidence_quote":"Survey of RDF stores and SPARQL engines that underpins the paper's coverage of the RDF data model and triple-store systems."}],"review_version":1}