{"id":"31f10735-8e31-45b3-8cdc-2c840af74508","arxiv_id":"2501.05821","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"About 36% of the 402,505 records in the University of Bologna's IRIS system could be matched into OpenCitations, but citations per record there (35.3) are close to Scopus (36.0) and Web of Science (34.7).","lead":"One of Italy's largest universities checked how much of its publication database appears in OpenCitations, an open database of research papers and citations, and found roughly a third of it. The study then showed that this open source counts about as many citations for the university's papers as the paid databases Scopus and Web of Science.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 36% coverage headline rests on searching each IRIS record under only one prioritized PID and never searching 34.5% of records that lack DOI/PMID/ISBN, so it is an unquantified lower bound presented as a precise rate.","rationale":"The central claim is the 36% coverage rate and the conclusion that OpenCitations has comparable coverage to Scopus and Web of Science. The 36% is not a direct measurement of presence in OC Meta but a counting of records that match under a restrictive identifier-selection rule. The reader's weakest assumption identifies the no-PID denominator; I agree and sharpen it: the single-PID priority affects not only no-PID records but also records with multiple PIDs, which are searched under only the first PID. The paper itself acknowledges by-construction limits of OC Meta (only resources participating in citations), but does not quantify how many additional IRIS records would match if all PIDs were tried or if no-PID records were reconciled. The companion study proves the lower-bound nature with 10,387 additional matches. This concern is load-bearing because the abstract presents 36% as a precise rate and the 'comparable coverage' conclusion depends on that number. The proposed test would settle it by producing a corrected coverage under less restrictive matching. My recommendation is UNCHANGED: the reader's CONDITIONAL verdict already requires the abstract to be qualified and sensitivity analyses to be reported; the test here is the concrete check for those conditions.","tokens_in":19609,"tokens_out":8532,"duration_ms":79436,"concrete_test":"Re-run the Comparator on the released IRIS dump and OC Meta CSV dumps, first querying all PIDs present per IRIS record (DOI, PMID, ISBN, in any combination) rather than the single prioritized PID, and then, for the 59,856 unmatched and 138,926 no-PID records, applying title+author reconciliation against OC Meta as in Andreose & Zilli (2025). Compute the corrected coverage = (matched all-PID + title+author matches) / 402,505. If the corrected coverage exceeds 41% (i.e., >5 percentage points above 36%), the abstract must be revised to report a range or a lower bound with the no-PID and multi-PID caveats.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's headline figure treats 'coverage' as the fraction of IRIS records found in OC Meta, but the Comparator only ever queries OC Meta with one PID per record. In the Validator step, 'we select one of the three PIDs, prioritising DOIs, PMIDs, and ISBNs' and 'the first is picked' when multiple identifiers exist. Consequently, a record carrying both a DOI and an ISBN is searched only by its DOI; if that DOI is absent from OC Meta while the ISBN is present, it is counted as unmatched even though a valid match exists. The no-PID group is worse: 138,926 records (34.5% of the dump) are never searched at all because they lack all three PIDs, yet they remain in the denominator of the 36% figure. Thus the headline 'only 36% of IRIS is covered' is a lower bound with two unquantified holes, not a measured rate. The reader's conditions (1) and (2) capture this, and the companion study's 10,387 additional Crossref/OC Meta matches for the no-PID group demonstrates the lower-bound nature. If a material fraction of the 59,856 unmatched deduplicated PIDs also have alternate PIDs that do match, the headline changes by more than a rounding error.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper describes a methodology to measure the coverage of the University of Bologna's IRIS bibliographic records in OpenCitations Meta and the number of citations involving them in OpenCitations Index. Using dumps from May–July 2025, the workflow filters IRIS records to those with DOIs, PMIDs, or ISBNs, selects one PID per record (DOI > PMID > ISBN), deduplicates, and matches against OC Meta. The authors report that 145,143 of 402,505 IRIS records (36%) are present in OC Meta, that journal articles have the highest coverage (91.6%), and that 5,129,406 citations in OC Index point to IRIS records, with an average of 35.34 citations per record, similar to Scopus (36.05) and Web of Science (34.71). All datasets and code are publicly available.","tokens_in":19753,"tokens_out":5955,"duration_ms":52917,"significance":"If the results were taken at face value, the paper would provide a reproducible, open methodology for institutional coverage analysis and evidence that open citation infrastructure can offer quantitative coverage comparable to proprietary databases. The paper is transparent about data provenance, publishes all intermediate datasets and code, and discloses the OpenCitations affiliation of two authors. The central coverage estimate, however, is an identifier-based lower bound rather than a precise rate; the companion study cited in the paper shows that at least 10,387 additional IRIS records without DOIs/PMIDs/ISBNs can be matched to OC Meta, so the true coverage is higher than 36%. This does not invalidate the methodological contribution, but it means the headline finding and the comparison to Scopus/WoS are not yet established at the claimed precision.","major_comments":[{"comment":"The headline 'only 36% of IRIS is covered' is presented as a precise rate, but it is a lower bound. By construction of the Trimmer step, 138,926 IRIS records without DOI, PMID, or ISBN (34.5% of the dump) are never queried against OC Meta, and the paper's own companion study (Andreose & Zilli, 2025) reports that 10,387 of those records can be matched to OC Meta via Crossref-derived DOIs. Adding these alone raises the coverage to at least (145,143 + 10,387) / 402,505 = 38.6%, and the true figure may be higher still. The abstract, RQ1 answer, and Discussion should either quantify this as an identifier-based lower bound or incorporate the reconciliation results. This is load-bearing for the paper's central claim.","section":"Results / Bibliographic Records types; Abstract"},{"comment":"The Validator selects exactly one PID per record (DOI > PMID > ISBN). A record with a DOI that is absent from OC Meta but with an ISBN or PMID that is present is counted as unmatched, even though a valid match exists. The paper does not report how many of the 59,856 unmatched deduplicated PIDs have alternative identifiers, so the magnitude of this second undercount is unknown. Please add a quantification (for example, by querying OC Meta with all PIDs per record or by reporting the number of records with multiple PID types and the match rates by fallback identifier). Without this, the 36% figure cannot be interpreted as an accurate coverage estimate.","section":"Validator; Bibliographic Records not included in OpenCitations Meta"},{"comment":"The explanation that 'OC Meta only includes bibliographic resources that take part in citations... resulting in a missing value for the present study' applies to PID-bearing records that are absent from OC Meta, but it does not address the 138,926 no-PID records that were never looked up. The paper should distinguish between 'coverage by PID-based lookup' and 'coverage by any available metadata (title/author/other identifiers).' Since RQ1 asks about 'the current coverage of the publications... in OpenCitations', the operational definition of coverage should be stated explicitly, and the reported rate should be qualified accordingly.","section":"Discussion, second paragraph"}],"minor_comments":[{"comment":"The methodology text says the workflow comprises five steps, but Figure 3 and the following description list six (Trimmer, Validator, Deduplicator, Comparator, Citation Scanner, Citation Counter). Please align the count.","section":"Methodology"},{"comment":"Table 13's heading contains the typo 'Iris in Idex' — it should be 'Iris in Index'.","section":"Table 13"},{"comment":"The Validator states that 'the first is picked' when multiple identifiers exist; the order in which identifiers appear in the IRIS CSV is not documented, so this selection is not fully reproducible. Specify that the priority scheme is applied to an explicit ordered list.","section":"Validator"},{"comment":"The data availability statement for Scopus and Web of Science notes that raw data cannot be published; the paper still reports exact citation counts and averages. A brief statement on the provenance of the Scopus/WoS snapshots (query date, API version) would improve transparency.","section":"Data Availability Statement"}],"recommendation":"major_revision","confidential_remarks":"The paper's conflict-of-interest disclosure is thorough and appropriate, and the authors' affiliation with OpenCitations is relevant context. The reported measurements are direct counts, and I do not see evidence of circular reasoning. The main concern is presentation: the headline figure overstates precision and understates coverage. I recommend major revision primarily to fix the lower-bound issue and add a sensitivity analysis for the multi-PID case; the reproducibility and open-data practices are strong points and the paper is otherwise publishable in a bibliometrics venue."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The genuinely new thing here is the full-scale OpenCitations audit of one institution's CRIS output, with all code, derived datasets, and a protocol released. The pipeline is transparent, the counts are internally consistent, and the companion study on the no-PID records is a good-faith attempt to quantify one of the biggest gaps. As an institutional case study for the Barcelona Declaration agenda, it is useful and mostly well executed.\n\nThe main soft spot is the headline. \"Only 36% of IRIS is covered\" treats 138,926 records without DOI/PMID/ISBN as uncovered without ever searching for them in OC Meta. That makes 36% an identifier-based lower bound, not a measured rate. The paper itself reports the 34.5% no-PID share and even cites the companion study's 10,387 additional Crossref/OC Meta matches, so a careful reader can reconstruct the true state of affairs. But the abstract still presents 36% as a precise figure, which is misleading. The same concern applies to the PID-selection heuristic: a record with both a DOI and an ISBN is searched only by the DOI, so some matches are missed. This is a presentation-and-robustness issue, not a demonstrated error; the direction of the bias is clear and upward.\n\nTwo smaller points. The deduplication priority tables look reasonable but are ad hoc, and there is no sensitivity analysis showing how the headline counts move under different priorities. That is worth adding. And the Scopus/WoS citation counts cannot be published because of license agreements; the paper says so explicitly, and the counts are only used for aggregate comparison, so I do not treat this as a flaw, just a limitation.\n\nThe citation pattern is fine. The authors include OpenCitations staff and a Barcelona Declaration representative, but the conflict is disclosed, and the measurements are direct counts over externally hosted dumps plus proprietary queries, not fitted results. Circularity burden is low.\n\nBottom line: this deserves a serious referee. I would send it out, with the request that the abstract and discussion clearly label 36% as a lower bound from identifier-based matching and that the authors add a simple sensitivity check on deduplication priorities. A revised version would be a solid contribution to the open-metadata coverage literature.","headline":"A transparent, reproducible coverage audit of a large Italian CRIS against OpenCitations, with a headline 36% that is really a lower bound because a third of IRIS records were never searched.","tokens_in":20485,"tokens_out":1405,"would_cite":true,"duration_ms":15541,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper shows that OpenCitations holds a comparable share of the University of Bologna's publications and citations as Scopus and Web of Science, despite matching only 36% of the full IRIS record set by persistent identifier.","keywords":["Bibliographic metadata","Citation data","CRIS systems","IRIS","OpenCitations","Open research information","Coverage analysis","Barcelona Declaration"],"falsifier":"Take a random sample of the 138,926 records in Iris No ID, search OpenCitations Meta by title and author, and recompute the coverage fraction; if more than a small fraction of that sample is found, the reported 36% understates OpenCitations' actual coverage. Separately, requesting Scopus and Web of Science citation counts for exactly the 145,143 records matched in OpenCitations would show whether the three systems' per-record averages still align when computed on an identical record set.","tokens_in":1875,"feed_emoji":"📚","tokens_out":2455,"duration_ms":68343,"temperature":0.7,"pith_summary":"The paper asks whether an open scholarly infrastructure can stand in for the closed commercial databases that universities currently rely on for research assessment. Taking the University of Bologna's IRIS system as a test case, it measures how many of the university's 402,505 bibliographic records appear in OpenCitations and how many citations those records receive. The headline result is that 36% of all IRIS records, or 145,143, are found in OpenCitations Meta, with journal articles covered at 91.6%; the same matched records receive 5,129,406 citations in the OpenCitations Index. When compared by citation count, OpenCitations performs on par with Scopus and Web of Science, with roughly 35 citations per matched record in all three systems. If this holds, institutions have a concrete quantitative basis for moving research-information workflows toward open data.","feed_headline":"OpenCitations matches Scopus and Web of Science coverage at Bologna","feed_subtitle":"Identifier-matched records and per-paper citation counts align with commercial indexes, supporting open research assessment.","key_machinery":"The load-bearing object is a five-stage identifier-matching pipeline. The Trimmer extracts IRIS records carrying a DOI, PMID, or ISBN; the Validator normalizes and discards malformed identifiers; the Deduplicator collapses 86,306 duplicate identifier assignments using a type-priority table that prefers journal articles over books, book chapters, and proceedings articles; the Comparator aligns the remaining 204,999 unique identifiers against OpenCitations Meta with a publication-year cutoff, producing the Iris in Meta and Iris Not in Meta datasets; and the Citation Scanner and Citation Counter extract all OpenCitations Index citations involving the matched OMIDs and compare per-record averages to Scopus and Web of Science counts. The OMID and OCI identifiers are the connective tissue: each IRIS record is matched to one OMID through its persistent identifier, and that OMID is then used to pull every citation from a 2.2-billion-citation index.","core_discovery":"The central claim is that, in the local context of the University of Bologna, open citation data are quantitatively equivalent to proprietary data for describing institutional research output. The authors arrive at this through a set of measurements: 145,143 of 402,505 IRIS records match OpenCitations Meta, which is 70.8% of the 204,999 deduplicated identifier-carrying records; 5,129,406 citations in the OpenCitations Index point to IRIS records; and the average number of citations per matched record is 35.34 for OpenCitations, 36.05 for Scopus, and 34.71 for Web of Science. The paper interprets the similarity of these averages as evidence that the open infrastructure can replace closed systems for quantitative purposes, at least within a single institutional context.","pith_inferences":["The 36% figure is best read as a lower bound: records without DOI, PMID, or ISBN are never searched in OpenCitations Meta, so the true fraction of Bologna output present in OpenCitations is likely higher than reported.","The comparability claim rests on per-record averages computed over different sets of matched records in each system; a stricter test would compare citation counts for exactly the records present in all three systems, but the proprietary data needed for that test is not disclosed.","If the same pipeline were run at other Italian universities using IRIS, the result would show whether Bologna's quantitative parity is a general property of OpenCitations or specific to this institution's publication mix.","A useful next experiment is to measure the overlap of the actual citing entities, not just citation counts; the paper notes that this is impossible with the aggregated data available from Scopus and Web of Science."],"forward_implications":["For the University of Bologna, OpenCitations already provides article-level coverage sufficient for many assessment workflows, since 91.6% of journal articles in IRIS are matched.","The published CC0 IRIS dump and the open pipeline give other institutions a reusable template for measuring their own coverage in open infrastructures.","The main barrier to higher coverage is the 34.5% of IRIS records that carry none of the three matched identifiers; the paper's Crossref experiment shows at least 10,387 of these could be recovered through metadata-based reconciliation.","Citation counts extracted from OpenCitations can be used to benchmark against Scopus and Web of Science in annual reporting, with known comparability of the per-record averages.","The new OpenCitations ingestion workflow could, in principle, absorb IRIS-compliant data directly, turning the current one-way matching into a mechanism for closing the coverage gap."],"supporting_citations":[{"why":"Supplies the UNIBO IRIS bibliographic data dump of 402,505 records that is the input to the entire coverage analysis.","marker":"Amurri et al., 2025"},{"why":"Provides the OpenCitations Meta dump against which IRIS records are matched, defining the coverage denominator and the OMID identifiers used.","marker":"OpenCitations, 2025b"},{"why":"Provides the OpenCitations Index dump of 2,216,426,689 citations from which the 5,129,406 citations to IRIS records are extracted.","marker":"OpenCitations, 2025a"},{"why":"Describes the OpenCitations Meta collection and its data model, establishing the metadata structure used for matching and type alignment.","marker":"Massari et al., 2024"},{"why":"Describes the OpenCitations Index and the OCI citation identifier, grounding the citation-scanner step.","marker":"Heibi et al., 2024"},{"why":"Introduces OpenCitations as an open scholarly infrastructure, motivating its selection as the authoritative open source for the comparison.","marker":"Peroni & Shotton, 2020"},{"why":"Provides the description of Scopus as a curated bibliometric data source, supporting the use of Scopus citation counts as a comparison baseline.","marker":"Baas et al., 2020"},{"why":"Provides the description of Web of Science as a data source, supporting the use of Web of Science citation counts as a comparison baseline.","marker":"Birkle et al., 2020"},{"why":"Supplies the companion study of the Iris No ID records, including the Crossref reconciliation experiment that estimates how many identifier-free records could be recovered.","marker":"Andreose & Zilli, 2025"}],"fun_headline_variants":["OpenCitations rivals Scopus and Web of Science at Bologna","Bologna research coverage: OpenCitations matches paid indexes","Open data equals commercial citation indexes for UNIBO","OpenCitations on par with Scopus and WoS for Bologna output","Bologna's OpenCitations coverage mirrors commercial databases"],"cache_read_input_tokens":22400,"weakest_assumption_plain":"The entire coverage rate depends on matching IRIS records to OpenCitations using only one of three identifiers (DOI, PMID, or ISBN), and the 138,926 records that have none of these are counted as uncovered without ever being looked up in OpenCitations, so the reported 36% figure is really a lower bound.","fun_headline_variants_meta":{"raw":{"variants":["OpenCitations rivals Scopus and Web of Science at Bologna","Bologna research coverage: OpenCitations matches paid indexes","Open data equals commercial citation indexes for UNIBO","OpenCitations on par with Scopus and WoS for Bologna output","Bologna's OpenCitations coverage mirrors commercial databases"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000485,"raw_usage":{"total_tokens":2376,"prompt_tokens":910,"completion_tokens":1466,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":526,"completion_tokens_details":{"reasoning_tokens":1387}},"tokens_in":526,"tokens_out":1466,"duration_ms":9931,"temperature":1.0,"reasoning_tokens":1387,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T21:07:09.075232+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a random sample of the 138,926 records in Iris No ID, search OpenCitations Meta by title and author, and recompute the coverage fraction; if more than a small fraction of that sample is found, the reported 36% understates OpenCitations' actual coverage. Separately, requesting Scopus and Web of Science citation counts for exactly the 145,143 records matched in OpenCitations would show whether the three systems' per-record averages still align when computed on an identical record set.","supporting_citations":[{"cited_title":"Introduction","cited_arxiv_id":null,"evidence_quote":"Provides the description of Web of Science as a data source, supporting the use of Web of Science citation counts as a comparison baseline."}],"review_version":1}