{"id":"f064d10a-54fe-47bf-ba33-c52c591695f1","arxiv_id":"2505.09798","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":3.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"The paper builds an ontology-based knowledge graph for North Macedonian procurement and a nearest-neighbor value predictor that achieves near-zero predictive skill (R2=0.056).","lead":"This study converts 896 high-value North Macedonian public procurement contracts into an RDF knowledge graph, then tests a text-embedding and FAISS nearest-neighbor model for estimating contract values. It shows how standard semantic web tools and a simple retrieval model can make government contracting data queryable and analytics-ready.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Predictive evaluation in §VII lacks a stated train/test split, and even the reported R² ≈ 0.056 is near zero; the ML-driven risk-assessment claim is the least secure part of the central claim.","rationale":"The reader's weakest assumption identifies the same load-bearing concern: the predictive evaluation has no documented out-of-sample protocol, so the reported ML metrics may not reflect true generalization. I read the paper in good faith: the ontology, RML-based RDF transformation, SHACL validation, and SPARQL analytics are described concretely and are consistent with the stated data-engineering contribution. Those parts are not the contested ground. The contested ground is the predictive modeling, which the abstract and conclusion elevate to a central feature ('insights into procurement trends and risk assessment'). Table III is the sole support for that claim, and its evaluation protocol is absent. The paper itself contains no appended limitation statement or caveat acknowledging this gap. The concern is therefore not about disagreement with a field consensus; it is about an internally unsupported empirical claim. A proper leave-one-out or temporal-split check would settle whether the reported metrics survive out-of-sample evaluation. If they do not, the abstract and conclusion must be revised to remove or substantially qualify the ML and risk-assessment language. The ontology/RDF pipeline appears to be a legitimate but modest contribution, so a conditional acceptance requiring this check and more measured claims is appropriate rather than outright rejection. Thus the reader's CONDITIONAL verdict stands unchanged.","tokens_in":4333,"tokens_out":2392,"duration_ms":27659,"concrete_test":"Rerun the Table III evaluation with strict leave-one-out cross-validation on all 896 contracts: for each contract, remove that contract (and any contracts with identical subject text) from the FAISS index, embed the subject description, retrieve the top-9 nearest neighbors, and take the median of their known values as the prediction. Also run a temporal split where training contracts are restricted to those dated before each test contract's award date. Report RMSE, MAE, R², and confidence intervals for both protocols. If R² stays near 0.056 and the RMSE/MAE improvement over the median baseline shrinks to near zero, the ML claim in §VII is unsupported; if the metrics change materially, the original evaluation was in-sample.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim includes machine-learning-driven predictive modeling for procurement trends and risk assessment, but the only evidence is Table III, comparing a median baseline against a FAISS kNN estimator. Section VII-A describes encoding all 'previously known procurement contracts' into the FAISS index and querying 'a new procurement contract', yet it never states whether the test contract's own vector was excluded from the index during evaluation. With only 896 contracts, self-retrieval would materially inflate the reported metrics, making the RMSE/MAE improvements over the baseline artifacts of an in-sample lookup rather than evidence of generalization. Furthermore, even taken at face value, R² = 0.056 means the model explains roughly 6% of the variance in contract values; the accompanying statement that 'the model captures key patterns in procurement data' overstates a result whose error reduction over a constant median baseline is only about 5%. Because the abstract and conclusion present predictive modeling as a central extension enabling 'risk assessment' and 'advanced analytics', this unsupported evaluation is load-bearing: if the metrics do not reflect true out-of-sample performance, the ML-driven part of the central claim collapses, leaving only the modest ontology/RDF/SPARQL pipeline, which appears sound but is not what the abstract emphasizes.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents an applied pipeline that aggregates North Macedonian high-value public procurement records (896 contracts, 2009–2021), maps them from CSV/XLSX to RDF using an OWL ontology and RML mappings, validates the resulting graph with SHACL, and runs descriptive SPARQL queries to report summary statistics and top institutions, suppliers, and contract values. It then proposes a contract-amount prediction approach based on multilingual sentence embeddings (multilingual-e5-large-instruct) and a FAISS k-nearest-neighbor search, reporting RMSE, MAE, and R2 against a median baseline. The abstract and conclusion frame the work as a framework for semantic procurement analytics that includes machine-learning-based forecasting and risk assessment.","tokens_in":4693,"tokens_out":7795,"duration_ms":75235,"significance":"If judged only as a data-engineering case study, the paper is a plausible and potentially useful demonstration: the RML/SHACL pipeline is standards-based, the SPARQL queries are clearly described, and the summary statistics address transparency-relevant questions. The predictive component, however, is a central element of the stated contribution and is not established: the evaluation lacks a specified train/test split, uses a small dataset, and reports R2=0.056, which is near zero; no risk-assessment or anomaly-detection analysis is actually carried out. Strengths of the paper are its use of public data and the relatively simple, reproducible structure of the ontology and queries; no code, ontology file, or SPARQL scripts are provided, which limits verification. As it stands, the paper can be accepted only as a descriptive case study after the ML claims are substantially revised or removed.","major_comments":[{"comment":"The evaluation protocol for the predictive model is underspecified. §VII-A states that 'all previously known procurement contracts from our historical dataset are pre-encoded and stored in a FAISS index' and that a query is posed for 'a new procurement contract', but the paper never states how a test contract is separated from the indexed contracts, nor how many folds or repetitions were used. With only 896 contracts, retrieving the test contract's own embedding as its nearest neighbor would materially inflate similarity and deflate error. The authors must specify a true out-of-sample scheme (e.g., leave-one-out cross-validation with the query excluded from the index) and report the resulting performance; without this, the RMSE/MAE improvements in Table III cannot be interpreted as evidence of generalization.","section":"§VII-A and VII-B"},{"comment":"Even if the evaluation were out-of-sample, the reported result does not support the claim that 'the model captures key patterns in procurement data'. R2=0.056 means the model explains only about 5.6% of the variance in contract values, and the reductions relative to the median baseline are small (RMSE 42.21→39.92, MAE 13.23→12.77). The paper does not define how R2 is computed for a median baseline, nor does it provide confidence intervals, standard deviations, or a significance test. The near-zero R2 and modest error reduction should be described as a weak baseline comparison, not as evidence of predictive skill.","section":"Table III and end of §VII-B"},{"comment":"The abstract's claim that the system offers 'insights into procurement trends and risk assessment' and the conclusion's mention of 'anomaly detection' are not supported by the analysis presented. No risk model, risk indicator, anomaly-detection experiment, or threshold-based outlier analysis is defined in §VII. Predicting the monetary value of a contract via nearest neighbors is not equivalent to risk assessment. Either a concrete risk/anomaly analysis must be added, or the claims in the abstract and conclusion must be narrowed to contract-amount estimation and descriptive trend analysis.","section":"Abstract and §VIII"},{"comment":"The dataset is said to contain 'high-value contracts exceeding 1,000,000 euros', but all monetary values reported in the paper are in MKD (e.g., 241,083,174,450 MKD in Table II) without a stated exchange rate or conversion. This unit ambiguity matters for the ML target as well: it is unclear whether the RMSE/MAE values in Table III are in millions of MKD and how the EUR threshold was applied. The paper should specify the currency conversion and clearly state that all conclusions apply only to the high-value stratum, which is a selected subset of procurement activity.","section":"§II-B, §VI, Table II"}],"minor_comments":[{"comment":"The entry '7,2205' for 'Average contracts per institution' appears to be a typo; 896/127 ≈ 7.06, and the decimal separator should be consistent (elsewhere dots are used).","section":"Table II"},{"comment":"References [12] (Jones, term specificity) and [13] (Breiman, random forests) are not used appropriately: [12] is not cited in the text, and [13] is cited in §VII-B although no random-forest model is employed.","section":"References"},{"comment":"For reproducibility, the authors should provide a link to the ontology, the RML mapping files, the SHACL shapes, the SPARQL queries, and the prediction code, as well as the exact snapshot date of the open-data portal.","section":"Reproducibility"},{"comment":"The text says 'Unique Resource Identifiers' but the standard term is 'Uniform Resource Identifiers' (or IRIs); the paper should also describe the URI scheme used for institutions, suppliers, and contracts.","section":"§V-A"},{"comment":"Figures 3 and 4 are referenced but not accompanied by the actual plots in the submitted manuscript; if they are available, they should be embedded, and the axes and units should be labeled.","section":"Figures 3 and 4"},{"comment":"The embedding model 'multilingual-e5-large-instruct' has no citation or version identifier; the paper should cite the model card or a corresponding publication, and report embedding dimensionality and the choice of k=9 nearest neighbors without sensitivity analysis.","section":"§VII-A"}],"recommendation":"major_revision","confidential_remarks":"To the editor: this is a modest applied case study. The ontology is extremely simple, and the evaluation of the ML part is not yet at the level expected for a claim of 'risk assessment'. If the journal is interested in practical Semantic Web deployments, the revised version may fit, but the novelty and generality should be stated more honestly. The reference list contains citations that do not match the methods used, which suggests the manuscript should be checked carefully before resubmission."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Nothing here is formally new: RML, SHACL, SPARQL, sentence embeddings, and FAISS are all standard. The contribution is the concrete pipeline applied to North Macedonian high-value procurement contracts, and that part is decent. The ontology is minimal (three classes, two object properties, three datatype properties), but it is enough to make the data queryable, and the descriptive stats extracted via SPARQL look internally consistent. If you want a worked example of turning a small public dataset into RDF and running a few aggregation queries, this is fine. The data source is public, so most of this is reproducible.\n\nThe soft spot is Section VII. The authors build a FAISS index over \"all previously known procurement contracts\" and then query with \"a new procurement contract,\" but they never state that the test contract was excluded from the index. With only 896 contracts and top-9 neighbors, self-retrieval would put the query's own value into the median and artificially inflate the results. On top of that, the reported R² is 0.056—essentially zero predictive power. The RMSE/MAE improvement over a constant median baseline is about 5%, and the claim that the model \"captures key patterns\" overstates that. The abstract's mention of \"risk assessment\" is not backed by any analysis in the paper. These are not minor omissions; they hit the most emphasized part of the central claim.\n\nThat said, I don't think this is a cynical paper. The authors report the weak R² openly; they just over-interpret it. The descriptive pipeline is sound. The flaws are addressable: add an explicit train/test split, describe how the FAISS index was constructed, add error bars, and rewrite the abstract to what the system can actually do.\n\nI'd send this to a workshop or a short-paper venue, not a top journal. A serious referee should focus on the holdout protocol. With a proper evaluation and more measured claims, it could be a legitimate small contribution. Without that, it is just a demo. Worth engaging, but only after the ML section is fixed.","headline":"A competent data-engineering case study with a weak, possibly in-sample ML evaluation; the knowledge-graph part is sound but the abstract's risk-assessment claim is unsupported.","tokens_in":681,"tokens_out":1817,"would_cite":false,"duration_ms":39398,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A knowledge-graph pipeline turns North Macedonia's procurement contracts into queryable, predictable data.","keywords":["public procurement","ontology","knowledge graph","RDF","SPARQL","SHACL","semantic web","contract value prediction"],"falsifier":"Run the same prediction experiment with a documented leave-one-out split, where each test contract is removed from the similarity index before retrieving neighbors; if the reported RMSE of 39.92 and $R^2$ of 0.056 are reproduced only when test contracts remain in the index, the predictive claim fails.","tokens_in":4136,"feed_emoji":"📊","tokens_out":7014,"duration_ms":67323,"temperature":0.7,"pith_summary":"This paper presents a methodological framework for transforming twelve years of North Macedonian high-value public procurement contracts from spreadsheet files into a semantic knowledge graph. The framework defines a small public procurement ontology, converts the data to RDF with RML mappings, validates the result with SHACL shapes, and answers analytical questions through SPARQL. It further claims that a text-embedding similarity search over contract descriptions can estimate new contract amounts, reporting a modest improvement over the median baseline. If the framework works as described, procurement records become more transparent and queryable, and the pipeline offers a reusable template for other public datasets.","feed_headline":"Public procurement contracts become a queryable knowledge graph","feed_subtitle":"A semantic workflow turns 896 high-value contracts into a queryable graph for analytics and value prediction.","key_machinery":"The load-bearing object is a small public procurement ontology: the Contract class linked to Institution and Supplier through the object properties hasInstitution and hasSupplier, with datatype properties hasAmount, hasDate, and hasDescription, enforced by OWL cardinality restrictions. This schema dictates how RML maps spreadsheet attributes to RDF, what SHACL shapes validate, and what SPARQL queries can express. For prediction, the mechanism is a nearest-neighbor estimator: contract descriptions are embedded, compared by cosine similarity against a pre-indexed corpus, and the median amount of the top-nine similar contracts becomes the value estimate.","core_discovery":"The central claim is that public procurement data need not remain trapped in rigid tabular files: by modeling contracts through a small OWL ontology, converting CSV records to RDF with RML, validating them with SHACL, and loading them into a triple store, the dataset becomes a knowledge graph that supports flexible semantic queries. On this graph, the paper executes trend queries over 896 contracts from 2009 to 2021, identifying annual totals, largest contracts, most active institutions, and dominant supplier pairs. The paper also claims that an embedding-based retrieval system, which encodes contract descriptions, finds the nine nearest historical contracts, and takes the median of their values, predicts contract amounts with RMSE 39.92, MAE 12.77, and $R^2$ 0.056, compared with the median baseline RMSE 42.21, MAE 13.23, and $R^2$ -0.057.","pith_inferences":["Beyond the paper, the same ontology-and-SHACL workflow could be applied to other tabular public datasets such as budget lines, permits, or subsidies, giving each a queryable semantic layer without redesign.","A natural extension is to expand the ontology to include tender stages such as announcements, bids, and awards, so that risk analysis can look at competition and bid spreads rather than only final contract values.","The nearest-neighbor estimator suggests a testable pattern: if contract descriptions are informative, prediction error should shrink as the similarity index grows; this could be checked on the full low-value contract portal.","The framework implicitly proposes that semantic validation and predictive analytics belong together, so a practical next step would be publishing SHACL validation reports alongside each procurement dataset."],"forward_implications":["Anyone can query total public spending, top suppliers, and unusual contracts directly, without waiting for pre-built reports.","The ontology's required relations, every contract must have a supplier and an institution, act as a data-quality check that can be run before publication.","Embedding-based estimates give procurement planners a starting point for contract values before formal bids arrive.","The twelve-year knowledge graph supports longitudinal studies of spending shifts across governments and economic cycles.","Because the pipeline starts from XLSX files, it can be re-run whenever new data is published, keeping the graph current."],"supporting_citations":[{"why":"Defines the high-value contract dataset (over 1,000,000 euros) that is transformed into the knowledge graph.","marker":"[6]"},{"why":"Supplies the ontology-development method used to define Contract, Institution, and Supplier.","marker":"[7]"},{"why":"Provides the OWL language constructs behind the ontology restrictions.","marker":"[8]"},{"why":"Defines RML, the mapping language that converts the CSV data to RDF.","marker":"[9]"},{"why":"Defines SHACL, used to validate the resulting RDF against ontology constraints.","marker":"[10]"},{"why":"Defines SPARQL 1.1, the query language that produces the reported procurement statistics.","marker":"[11]"},{"why":"Supplies the transformer architecture behind the sentence embedding model.","marker":"[14]"},{"why":"Provides the similarity-search index used for nearest-neighbor retrieval in prediction.","marker":"[15]"}],"fun_headline_variants":["North Macedonia's procurement data becomes a queryable graph","Semantic graph makes 896 procurement contracts queryable","From spreadsheets to knowledge graph: procurement intelligence","Ontology-driven graph unlocks insight into public contracts","Public procurement meets ontology: a queryable knowledge graph"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The machine-learning result depends on each test contract being absent from the similarity index built from historical contracts, but the paper does not describe how training and test contracts were separated.","fun_headline_variants_meta":{"raw":{"variants":["North Macedonia's procurement data becomes a queryable graph","Semantic graph makes 896 procurement contracts queryable","From spreadsheets to knowledge graph: procurement intelligence","Ontology-driven graph unlocks insight into public contracts","Public procurement meets ontology: a queryable knowledge graph"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00091,"raw_usage":{"total_tokens":3869,"prompt_tokens":863,"completion_tokens":3006,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":479,"completion_tokens_details":{"reasoning_tokens":2932}},"tokens_in":479,"tokens_out":3006,"duration_ms":21235,"temperature":1.0,"reasoning_tokens":2932,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T21:23:20.766094+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same prediction experiment with a documented leave-one-out split, where each test contract is removed from the similarity index before retrieving neighbors; if the reported RMSE of 39.92 and $R^2$ of 0.056 are reproduced only when test contracts remain in the index, the predictive claim fails.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the high-value contract dataset (over 1,000,000 euros) that is transformed into the knowledge graph."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the ontology-development method used to define Contract, Institution, and Supplier."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the OWL language constructs behind the ontology restrictions."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines RML, the mapping language that converts the CSV data to RDF."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines SHACL, used to validate the resulting RDF against ontology constraints."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines SPARQL 1.1, the query language that produces the reported procurement statistics."},{"cited_title":"30, 2017","cited_arxiv_id":null,"evidence_quote":"Supplies the transformer architecture behind the sentence embedding model."}],"review_version":1}