REVIEW 3 major objections 4 minor 15 references
Bridging the Digital Divide: Approach to Documenting Early Computing Artifacts Using Established Standards for Cross-Collection Knowledge Integration Ontology
T0 review · 3 major / 4 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read The paper argues that a small, carefully chosen set of CIDOC-CRM classes and properties can give community archivists a human-readable, standards-based way to document early computing artifacts so that separate collections can be…
desk verdict A modest, honest pilot applying CIDOC-CRM to community computer archives; the mapping is illustrative and the central human-readability claim outruns the evidence, but the paper's own limitations keep it credible. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The key machinery is CIDOC-CRM itself, an ISO-standard formal ontology for cultural heritage, and the proposed approach of selecting a minimal set of its 'building blocks.' These blocks—E22 Human-Made Object, E42 Identifier, E55 Type, E39 Actor, E7 Activity, E52 Time-Span, E73 Information Object, E53 Place, E41 Name/FilePath—are instantiated as atomic nodes and connected by relations like P2 'has type', P106 'forms part of', P14 'carried out by', P4 'has time-span', P62 'depicts', P16 'used specific object', P53 'has former or current location', and P1 'is identified by'. The model's work is to capture not just object attributes but also the chain of human actions and provenance, so that every digitization step is attributable and the structure can later integrate with other ontologies.
What would settle it
If volunteers who are not already fluent in ontologies cannot correctly encode a simple cassette-digitization workflow using the proposed minimal CIDOC-CRM model (e.g., higher error rates or longer completion times than with a simple file-naming scheme), the claim that this is an effective minimal building set would be refuted.
Extended reading notes
Core claim
The central claim is that CIDOC-CRM—despite its reputation for complexity—can be applied to early computing artifacts using only a small selection of its classes and properties, and that this minimal set is enough to capture the essential archival workflow while preserving full accountability. Through a concrete cassette-tape example, the paper shows how inventorying, photographing, and digitizing can be modeled with objects like E22 Human-Made Object, E42 Identifier, E55 Type, E39 Actor, E7 Activity, and E52 Time-Span, linked by properties such as P2 has type, P106 forms part of, P14 carried out by, and P62 depicts. The resulting graph is said to be intuitive to read and expandable: later actors can add decoded binaries, game titles, or links to external ontologies, building knowledge step by step without discarding earlier work.
Load-bearing premise
The survey's 20 participants, recruited through invite-only hobbyist channels, are assumed to represent the broader community of archive users and volunteer archivists.
Editorial extensions
If this is right
- A volunteer can document a cassette (or similar artifact) with a handful of object types, and the record is immediately machine-readable and linkable.
- Because every activity is tied to an actor and a timestamp, the data carries an audit trail that supports accountability and future verification.
- If multiple community collections adopt this minimal CIDOC-CRM core, queries across collections become feasible without requiring each project to re-catalog its holdings.
- The model is extensible: later work (e.g., decoding audio, verifying software titles, linking to game databases) can be added as new objects and relations on top of the existing graph.
- Switching to a richer CIDOC-CRM extension in the future would not invalidate data already captured, since the minimal core is compatible with the full model.
Reading between the lines
- The same minimal-block approach could be tested on other born-digital artifacts such as floppy disks, CD-ROMs, or web archives to see whether the building-block set generalizes beyond cassette tapes.
- The claim that CIDOC-CRM is 'human-readable' rests on the authors' own experience and a small sample; a controlled usability study with naive volunteers would be a natural next test.
- If the minimal model were adopted, the graph structure would enable automated consistency checks—for example, detecting mismatched tape/inlay pairings by querying P106 relations—something a flat file-naming scheme cannot do.
- The integration potential goes beyond early computing: the same pattern could help community archives of other niche technologies (e.g., amateur radio, vintage instruments) plug into broader cultural-heritage data networks.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper investigates how community-driven collections of early computing artifacts can be documented using the CIDOC Conceptual Reference Model (CRM), an ISO-standard ontology for cultural heritage. The authors report a small survey (N=20) of enthusiasts and volunteer archivists, identifying perceived needs such as bibliographic information, software, technical documentation, historical context, and personal stories. They then present a worked example in which CIDOC-CRM v7.3 is used to model an audio cassette, its inlay, its inventory identifier, and the digitization/photography activities around it (Sections 4 and 5). The paper concludes that, despite its complexity, CIDOC-CRM is logical, human-readable, and adaptable, and that minimal building blocks can empower community-led heritage projects. The limitations section, however, concedes that the current experiment does not yet show clear advantages over a smart file-naming scheme.
Significance. If the central claim were fully supported, the paper would provide a valuable template for adopting a formal ontology in grassroots preservation efforts, enabling cross-collection integration for early digital artifacts. The community-needs data, while limited, are a useful contribution to the human-centered design of such archives, and the worked example is a concrete, reproducible illustration of a CIDOC-CRM mapping for this niche domain. The paper is also transparent in its limitations, which strengthens its credibility. However, the strong claims in the abstract and conclusion are not backed by the empirical evidence presented; the survey does not test the ontology, and the example is researcher-authored. As an exploratory mapping proposal with clearly framed future work, this is a meaningful contribution; as a demonstration of usability and empowerment, it falls short.
major comments (3)
- [§5 and §6] The central claim is contradicted by the paper's own limitation statement. Section 5 explicitly says "the experiment we have presented, at this stage, does not seem to have many advantages" over a smart file-naming scheme, yet the abstract and Section 6 assert that CIDOC-CRM "proves logical, human-readable, and adaptable" and "empower[s] community-led heritage projects." The survey in Section 2 measures user needs and current practices, but no participant was asked to read, create, or evaluate a CIDOC-CRM model. The worked example in Section 4 is authored by the researchers, not by the target volunteers. Consequently, the conclusion exceeds the evidence and conflicts with the stated limitation. Please rewrite the abstract and conclusion to describe this as an exploratory mapping proposal with testable hypotheses, rather than a demonstrated usability or empowerment result.
- [§2, Table 1] The N=20 sample was recruited through invite-only Facebook and Discord channels, which likely over-represents highly engaged enthusiasts and may not reflect the broader community of potential archive users or volunteers. The paper does not report response rates, demographics, or any sampling strategy, and Table 1 aggregates raw counts without discussing generalizability. Since the motivation for adopting CIDOC-CRM rests in part on the transferability of these identified needs, this is a load-bearing limitation. Please add an explicit discussion of sample representativeness and temper claims about "the community" and general needs accordingly.
- [§4, Figures 1 and 2] The worked example demonstrates that CIDOC-CRM can encode the described scenario, but it does not establish "human-readable" or low-overhead documentation. Figure 1 alone uses at least ten distinct classes and properties for a cassette and its inlay, and the authors themselves acknowledge in Section 3 that "data that needs to be entered should be strictly related to the activity" and in Section 5 that a simple naming scheme could encode the same information. Without a comparative evaluation with target users (e.g., comprehension, task completion time, error rate, or perceived burden relative to a TOSEC-style scheme), the claim of reduced overhead and human readability is unsupported. Either add such an evaluation or remove these claims from the conclusions.
minor comments (4)
- [Affiliations, page 1] There are spacing/typo issues in the affiliations: "Warsa w, Poland" and "P o land" should be corrected to "Warsaw, Poland."
- [References] Reference [4] is cited as a directly relevant 2023 study on preserving early digital artifacts, but the reference list gives only a title without publication venue or DOI. Please provide full bibliographic details so readers can locate it.
- [Figure 2 caption] The caption "Redundant objects removed for clarity" is vague; please specify which objects are redundant and why their removal does not affect the model's correctness.
- [§3, footnote 9] The URL to the CIDOC-CRM specification is cited in a footnote; consider also stating the version number (v7.3) in the main text where it is first used, so readers who skip footnotes understand the version context.
Circularity Check
No significant circularity: the paper applies an external standard (CIDOC-CRM) to an author-built example; its strongest claims are under-supported but not circular.
full rationale
The paper's chain is: (1) an N=20 survey of enthusiast needs and practices (Section 2); (2) a reasoned choice to consider CIDOC-CRM because it is an ISO standard with existing deployments (Section 3); (3) an author-constructed CIDOC-CRM v7.3 model of a cassette and its digitization (Section 4, Figures 1-2); and (4) a qualitative conclusion that the model is 'intuitive and clear to the recipient, even human-readable' (Section 6). None of these steps reduces to its own input. The CIDOC-CRM model is not fitted to the survey data, and the claim that the model is human-readable is not derived by definition from CIDOC-CRM; it is the authors' self-assessment of their own example. The self-citation in Section 2 (reference [4], used to say the survey results are consistent with a broader 2023 study) is not load-bearing for the central CIDOC-CRM claim. The paper's Section 5 explicitly concedes that the current experiment 'does not seem to have many advantages' over a smart file-naming scheme, which undercuts the abstract's claim of empowerment but is an honest limitation rather than circular reasoning. The primary weakness is evidentiary overclaim: the strongest conclusion is asserted from a single untested worked example, not demonstrated by user evaluation. That is a correctness or support concern, not circularity.
Assumptions & free parameters
assumptions (3)
- domain assumption Survey responses from N=20 participants represent the needs of the broader community of early computing archivists and users.
- domain assumption CIDOC-CRM is suitable for modeling early computing artifacts without domain-specific extension.
- domain assumption The worked example is generalizable to other early computing artifacts and workflows.
Cite this review
Pith. "Pith review of Bridging the Digital Divide: Approach to Documenting Early Computing Artifacts Using Established Standards for Cross-Collection Knowledge Integration Ontology." pith.science (2026). https://pith.science/paper/WXKWY7DH
@misc{pith2026250112603,
author = {Pith},
title = {Pith review of: Bridging the Digital Divide: Approach to Documenting Early Computing Artifacts Using Established Standards for Cross-Collection Knowledge Integration Ontology},
year = {2026},
howpublished = {\url{https://pith.science/paper/WXKWY7DH}},
note = {Machine review of arXiv:2501.12603}
}
read the original abstract
In this paper we address the challenges of documenting early digital artifacts in collections built to offer historical context for future generations. Through insights from active community members (N=20), we examine current archival needs and obstacles. We assess the potential of the CIDOC Conceptual Reference Model (CRM) for categorizing fragmented digital data. Despite its complexity, CIDOC-CRM proves logical, human-readable, and adaptable, enabling archivists to select minimal yet effective building blocks set to empower community-led heritage projects.
Figures
Reference graph
Works this paper leans on
-
[1]
IFLA Journal 43, 379– 390 (2017)
Cadavid, J.A.P.: Evolution of legal deposit in new zealan d. IFLA Journal 43, 379– 390 (2017). https://doi.org/10.1177/0340035217713763
-
[2]
Cultural Her- itage in a Changing World pp
Constantinidis, D.: Crowdsourcing culture: challenges to change. Cultural Her- itage in a Changing World pp. 215–234 (2016). https://doi.o rg/10.1007/978-3-319- 29544-2_13
-
[3]
Gonzalez-Perez, C., Martín-Rodilla, P., Parcero-Oubiñ a, C., Fábrega-Álvarez, P., Güimil-Fariña, A.: Extending an abstract reference mod el for trans- disciplinary work in cultural heritage. vol. 343, pp. 190–2 01 (11 2012). https://doi.org/10.1007/978-3-642-35233-1_20 8 M. Grzeszczuk et al
-
[4]
Grzeszczuk, M., Skorupska, K.: Preserving the artifacts of the early digital era: A study of what, why and how? (2023)
work page 2023
-
[5]
Grzeszczuk, M., Skorupska, K., Grabarczyk, P., Fuchs, W. , Aubin, P.F., Dietrick, M.E., Karpowicz, B., Masłyk, R., Zinevych, P., Stawski, W., et al.: Preserving Tangible and Intangible Cultural Heritage: The Cases of Vol terra and Atari. In: Machine Intelligence and Digital Interaction Conference. pp. 351–358. Springer (2023)
work page 2023
-
[6]
Journal of the Ameri- can Society for Information Science and Technology 63, 2153–2164 (2012)
Lor, P.J., Britz, J.: An ethical perspective on political -economic issues in the long-term preservation of digital heritage. Journal of the Ameri- can Society for Information Science and Technology 63, 2153–2164 (2012). https://doi.org/10.1002/asi.22725
-
[7]
Journal of the South Afr ican Society of Archivists 54, 55–70 (2021)
Masenya, T.M.: The use of metadata systems for the preserv ation of digital records in cultural heritage institutions. Journal of the South Afr ican Society of Archivists 54, 55–70 (2021). https://doi.org/10.4314/jsasa.v54i1.5
-
[8]
In: Proceedings of CIDOC 2011 Knowled ge Management and Museums Conference, Sibiu, Romania (2011)
Mazurek, C., Sielski, K., Walkowska, J., Werla, M.: Appli cability of cidoc crm in digital libraries. In: Proceedings of CIDOC 2011 Knowled ge Management and Museums Conference, Sibiu, Romania (2011)
work page 2011
Show all 15 references
-
[9]
Semantic W eb 14, 553–584 (04 2023)
Melo, D., Rodrigues, I., Varagnolo, D.: A strategy for arc hives metadata repre- sentation on cidoc-crm and knowledge discovery. Semantic W eb 14, 553–584 (04 2023). https://doi.org/10.3233/SW-222798
2023 doi
-
[10]
Roudik, P., Buchanan, K., Ahmad, T., Zhang, L., Isajanya n, N., Boring, N., Gesley, J., Levush, R., Figueroa, D., Umeda, S., Hofverberg, E., Rod riguez-Ferrand, G., Feikert-Ahalt, C.: Digital legal deposit in selected juris dictions: Australia, canada, china, estonia, france, ...
2018
-
[11]
Virtual Archaeology Review 9, 50 (2018)
Ruymbeke, M.V., Hallot, P., Nys, G., Billen, R.: Impleme ntation of multiple in- terpretation data model concepts in cidoc crm and compatibl e models. Virtual Archaeology Review 9, 50 (2018). https://doi.org/10.4995/var.2018.8884
2018
-
[12]
Frontiers in Artificial Intelligence and Applications (2020)
Sanfilippo, E.M., Markhoff, B., Pittet, P.: Ontological A nalysis and Modulariza- tion of CIDOC-CRM. Frontiers in Artificial Intelligence and Applications (2020). https://doi.org/10.3233/faia200664
2020 doi
-
[13]
expanded notion
Sköld, O.: Understanding the “expanded notion” of video games as archival objects: a review of priorities, methods, and conceptions. Journal of the Association for Information Science and Technology 69, 134–145 (2017). https://doi.org/10.1002/asi.23875
2017 doi
-
[14]
https://soar.si.edu (2020), accessed: 2024-10-24
Smithsonian Institution: Collections care - smithsoni an object-based research. https://soar.si.edu (2020), accessed: 2024-10-24
2020
-
[15]
Tada, H., Honda, O., Higuchi, M.: A file naming scheme using hierarchical-keywords. pp. 799– 804 (02 2002). https://doi.org/10.1109/CMPSAC.2002.1045103
2002 arXiv
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.