{"id":"11606905-9ac3-418c-ab1f-1aa408e2c43f","arxiv_id":"2501.12603","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A small qualitative study and a worked example suggest CIDOC-CRM can serve as a flexible foundation for documenting early computing artifacts in community archives.","lead":"The paper evaluates CIDOC-CRM, an existing cultural heritage ontology, for documenting early computing artifacts, based on a survey of 20 enthusiasts and a worked example of cataloging a cassette tape. It argues that the standard is logical and adaptable enough for volunteer-led archives, despite its complexity.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Central claim that CIDOC-CRM is 'human-readable' and empowering for community archivists is not tested; the only evidence is a researcher-authored example, and the paper's own limitations concede no current advantage over a file-naming scheme.","rationale":"The reader's weakest assumption is the representativeness of the N=20 sample, which is a real concern about external validity. However, the more load-bearing gap is that the central claim would remain unsupported even if the sample were perfectly representative. The survey collects needs and obstacles but never tests whether community archivists can actually use CIDOC-CRM. The authors' own Section 5 says the current experiment has no clear advantage over a smart file-naming scheme, yet the abstract and Section 6 claim the model 'proves' human-readable and empowering. This is not an internal logical contradiction in the ontology itself, but it is a mismatch between the evidence presented and the strength of the conclusion. The fix is not simply a larger survey; it is a usability study comparing CIDOC-CRM against the paper's own baseline (file naming). Since the paper is early-stage and the proposed model is plausible, I would keep the reader's CONDITIONAL verdict: accept only with the required usability evaluation. My concern does not move the verdict, so 'UNCHANGED' is appropriate, with the condition made more explicit.","tokens_in":7087,"tokens_out":2939,"duration_ms":32210,"concrete_test":"Conduct a within-subject usability study in which community archivists who did not author the model (ideally sampled outside the invite-only channels) document the same audio cassette artifact twice: once using the proposed minimal CIDOC-CRM profile (E22/E39/E7/E52/E42/E55 from Figures 1-2) and once using a hierarchical file-naming scheme like TOSEC. Measure task completion time, correctness of entity and relationship selection, and a brief comprehension test of the resulting graph. If volunteers are slower, less accurate, or rate the ontology as harder than file naming, the 'human-readable/empowering' claim fails and the conclusion should be scaled back to 'representationally adequate in one example,' not 'proven usable.'","verdict_should_be":"UNCHANGED","load_bearing_attack":"The strongest claim, stated in the abstract and conclusion, is that CIDOC-CRM 'proves logical, human-readable, and adaptable, enabling archivists to select minimal yet effective building blocks' for community-led projects. The evidence offered is one worked example (Sections 4 and 5: CIDOC-CRM v7.3 models of an audio cassette and its digitization) authored by the researchers, not by archivists or volunteers. Section 5 explicitly concedes that 'the experiment ... does not seem to have many advantages' over a smart file-naming scheme, and the listed benefits are future, hypothetical structures. The N=20 survey (Section 2) measures what users want and current documentation practices; it never asks participants to read, create, or evaluate a CIDOC-CRM description. Thus the property 'human-readable' and the claim that minimal building blocks 'empower' community-led heritage projects are asserted rather than demonstrated. The concern is not disagreement with an outside consensus; it is that the paper's central conclusion goes beyond the data in the paper and is in tension with its own stated limitation.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper investigates how community-driven collections of early computing artifacts can be documented using the CIDOC Conceptual Reference Model (CRM), an ISO-standard ontology for cultural heritage. The authors report a small survey (N=20) of enthusiasts and volunteer archivists, identifying perceived needs such as bibliographic information, software, technical documentation, historical context, and personal stories. They then present a worked example in which CIDOC-CRM v7.3 is used to model an audio cassette, its inlay, its inventory identifier, and the digitization/photography activities around it (Sections 4 and 5). The paper concludes that, despite its complexity, CIDOC-CRM is logical, human-readable, and adaptable, and that minimal building blocks can empower community-led heritage projects. The limitations section, however, concedes that the current experiment does not yet show clear advantages over a smart file-naming scheme.","tokens_in":7300,"tokens_out":3050,"duration_ms":32497,"significance":"If the central claim were fully supported, the paper would provide a valuable template for adopting a formal ontology in grassroots preservation efforts, enabling cross-collection integration for early digital artifacts. The community-needs data, while limited, are a useful contribution to the human-centered design of such archives, and the worked example is a concrete, reproducible illustration of a CIDOC-CRM mapping for this niche domain. The paper is also transparent in its limitations, which strengthens its credibility. However, the strong claims in the abstract and conclusion are not backed by the empirical evidence presented; the survey does not test the ontology, and the example is researcher-authored. As an exploratory mapping proposal with clearly framed future work, this is a meaningful contribution; as a demonstration of usability and empowerment, it falls short.","major_comments":[{"comment":"The central claim is contradicted by the paper's own limitation statement. Section 5 explicitly says \"the experiment we have presented, at this stage, does not seem to have many advantages\" over a smart file-naming scheme, yet the abstract and Section 6 assert that CIDOC-CRM \"proves logical, human-readable, and adaptable\" and \"empower[s] community-led heritage projects.\" The survey in Section 2 measures user needs and current practices, but no participant was asked to read, create, or evaluate a CIDOC-CRM model. The worked example in Section 4 is authored by the researchers, not by the target volunteers. Consequently, the conclusion exceeds the evidence and conflicts with the stated limitation. Please rewrite the abstract and conclusion to describe this as an exploratory mapping proposal with testable hypotheses, rather than a demonstrated usability or empowerment result.","section":"§5 and §6"},{"comment":"The N=20 sample was recruited through invite-only Facebook and Discord channels, which likely over-represents highly engaged enthusiasts and may not reflect the broader community of potential archive users or volunteers. The paper does not report response rates, demographics, or any sampling strategy, and Table 1 aggregates raw counts without discussing generalizability. Since the motivation for adopting CIDOC-CRM rests in part on the transferability of these identified needs, this is a load-bearing limitation. Please add an explicit discussion of sample representativeness and temper claims about \"the community\" and general needs accordingly.","section":"§2, Table 1"},{"comment":"The worked example demonstrates that CIDOC-CRM can encode the described scenario, but it does not establish \"human-readable\" or low-overhead documentation. Figure 1 alone uses at least ten distinct classes and properties for a cassette and its inlay, and the authors themselves acknowledge in Section 3 that \"data that needs to be entered should be strictly related to the activity\" and in Section 5 that a simple naming scheme could encode the same information. Without a comparative evaluation with target users (e.g., comprehension, task completion time, error rate, or perceived burden relative to a TOSEC-style scheme), the claim of reduced overhead and human readability is unsupported. Either add such an evaluation or remove these claims from the conclusions.","section":"§4, Figures 1 and 2"}],"minor_comments":[{"comment":"There are spacing/typo issues in the affiliations: \"Warsa w, Poland\" and \"P o land\" should be corrected to \"Warsaw, Poland.\"","section":"Affiliations, page 1"},{"comment":"Reference [4] is cited as a directly relevant 2023 study on preserving early digital artifacts, but the reference list gives only a title without publication venue or DOI. Please provide full bibliographic details so readers can locate it.","section":"References"},{"comment":"The caption \"Redundant objects removed for clarity\" is vague; please specify which objects are redundant and why their removal does not affect the model's correctness.","section":"Figure 2 caption"},{"comment":"The URL to the CIDOC-CRM specification is cited in a footnote; consider also stating the version number (v7.3) in the main text where it is first used, so readers who skip footnotes understand the version context.","section":"§3, footnote 9"}],"recommendation":"major_revision","confidential_remarks":"The paper may be a better fit for a short/conference paper than a full journal article in its current form, given the preliminary nature of the evaluation. The authors' affiliation with The Foundation for the History of Home Computers is a relevant context for their qualitative recommendations, but I do not see it as a conflict; it does, however, reinforce the need for independent or user-based validation of the central claims. I also note that reference [4] is a self-citation that is used to position the study, but its exact content is not summarized; the authors should at least describe its findings so readers can judge the continuity of their work."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a small, honest study, not a breakthrough. It does something useful — takes a standard cultural-heritage ontology, CIDOC-CRM, and shows what a volunteer-driven archive of early computing artifacts could look like with it. The worked example (audio cassette processing and digitization) is clearly presented. The survey, though tiny, gives a reasonable snapshot of what volunteer archivists want and where they struggle, and the paper is refreshingly candid in Section 5: it admits the current mapping 'does not seem to have many advantages' over a smart file-naming scheme.\n\nThe main soft spot is the gap between the data and the abstract. Nobody in the study ever used or evaluated CIDOC-CRM; the mapping was produced by the authors. So the claim that the model 'proves logical, human-readable, and adaptable' is asserted, not demonstrated. That is a real overreach, but it is an overreach the authors partially disown in their limitations. Similarly, the N=20 convenience sample from invite-only channels may over-represent enthusiasts, so the motivational numbers should be read as qualitative signals, not population estimates.\n\nWhat’s genuinely new is narrow: applying CIDOC-CRM to home-computer artifacts with a concrete minimal-building-blocks example, plus a few survey insights about documentation practices in that community. That is enough to be a worthwhile exploratory paper, not enough to change practice yet.\n\nWho’s it for: digital preservation practitioners, museum informatics people, and HCI researchers interested in community-led archiving. A serious referee should see it, and with revisions — toning down the human-readability claim or adding a small usability study, and tempering the abstract — it could be a useful contribution. I’d take it for peer review; as is, it’s a conditional accept at best.","headline":"A modest, honest pilot applying CIDOC-CRM to community computer archives; the mapping is illustrative and the central human-readability claim outruns the evidence, but the paper's own limitations keep it credible.","tokens_in":7792,"tokens_out":2055,"would_cite":false,"duration_ms":20679,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that a small, carefully chosen set of CIDOC-CRM classes and properties can give community archivists a human-readable, standards-based way to document early computing artifacts so that separate collections can be…","keywords":["digital heritage","software preservation","CIDOC-CRM","ontology","community archives","early computing artifacts","metadata","knowledge integration"],"falsifier":"If volunteers who are not already fluent in ontologies cannot correctly encode a simple cassette-digitization workflow using the proposed minimal CIDOC-CRM model (e.g., higher error rates or longer completion times than with a simple file-naming scheme), the claim that this is an effective minimal building set would be refuted.","tokens_in":6917,"feed_emoji":"💾","tokens_out":3448,"duration_ms":35948,"temperature":0.7,"pith_summary":"Early home-computer artifacts are scattered across hobbyist collections with thin, inconsistent metadata. This paper asks whether a formal cultural-heritage ontology, CIDOC-CRM, can give community archivists a practical documentation structure that stays lightweight and still supports cross-collection knowledge integration. Based on a survey of 20 enthusiasts and a worked example of cataloging and digitizing a cassette tape, the authors claim that a minimal set of CIDOC-CRM building blocks is logical, human-readable, and adaptable. If true, community-led archives could adopt a standards-based model without incurring heavy cataloging overhead, and their records would become linkable to other heritage data sources.","feed_headline":"CIDOC-CRM tames early-computing archives with minimal building blocks","feed_subtitle":"Community volunteers can document tapes and software with a slim CIDOC-CRM subset, ready to link across collections.","key_machinery":"The key machinery is CIDOC-CRM itself, an ISO-standard formal ontology for cultural heritage, and the proposed approach of selecting a minimal set of its 'building blocks.' These blocks—E22 Human-Made Object, E42 Identifier, E55 Type, E39 Actor, E7 Activity, E52 Time-Span, E73 Information Object, E53 Place, E41 Name/FilePath—are instantiated as atomic nodes and connected by relations like P2 'has type', P106 'forms part of', P14 'carried out by', P4 'has time-span', P62 'depicts', P16 'used specific object', P53 'has former or current location', and P1 'is identified by'. The model's work is to capture not just object attributes but also the chain of human actions and provenance, so that every digitization step is attributable and the structure can later integrate with other ontologies.","core_discovery":"The central claim is that CIDOC-CRM—despite its reputation for complexity—can be applied to early computing artifacts using only a small selection of its classes and properties, and that this minimal set is enough to capture the essential archival workflow while preserving full accountability. Through a concrete cassette-tape example, the paper shows how inventorying, photographing, and digitizing can be modeled with objects like E22 Human-Made Object, E42 Identifier, E55 Type, E39 Actor, E7 Activity, and E52 Time-Span, linked by properties such as P2 has type, P106 forms part of, P14 carried out by, and P62 depicts. The resulting graph is said to be intuitive to read and expandable: later actors can add decoded binaries, game titles, or links to external ontologies, building knowledge step by step without discarding earlier work.","pith_inferences":["The same minimal-block approach could be tested on other born-digital artifacts such as floppy disks, CD-ROMs, or web archives to see whether the building-block set generalizes beyond cassette tapes.","The claim that CIDOC-CRM is 'human-readable' rests on the authors' own experience and a small sample; a controlled usability study with naive volunteers would be a natural next test.","If the minimal model were adopted, the graph structure would enable automated consistency checks—for example, detecting mismatched tape/inlay pairings by querying P106 relations—something a flat file-naming scheme cannot do.","The integration potential goes beyond early computing: the same pattern could help community archives of other niche technologies (e.g., amateur radio, vintage instruments) plug into broader cultural-heritage data networks."],"forward_implications":["A volunteer can document a cassette (or similar artifact) with a handful of object types, and the record is immediately machine-readable and linkable.","Because every activity is tied to an actor and a timestamp, the data carries an audit trail that supports accountability and future verification.","If multiple community collections adopt this minimal CIDOC-CRM core, queries across collections become feasible without requiring each project to re-catalog its holdings.","The model is extensible: later work (e.g., decoding audio, verifying software titles, linking to game databases) can be added as new objects and relations on top of the existing graph.","Switching to a richer CIDOC-CRM extension in the future would not invalidate data already captured, since the minimal core is compatible with the full model."],"supporting_citations":[{"why":"Provides the comparison with CHARM and asserts that CIDOC-CRM is deployable without a domain-specific extension from the start.","marker":"[3]"},{"why":"A broader international study by the authors that this survey builds on, establishing the diverse interests of archive users and preservationists.","marker":"[4]"},{"why":"Documents applicability of CIDOC-CRM in digital libraries, serving as evidence of successful deployments in related settings.","marker":"[8]"},{"why":"Offers a strategy for representing archive metadata in CIDOC-CRM and enabling knowledge discovery, supporting the paper's claim that integration and knowledge creation are feasible.","marker":"[9]"},{"why":"Provides an ontological analysis and modularization of CIDOC-CRM, backing the idea that a minimal subset can be selected without losing coherence.","marker":"[12]"},{"why":"Describes a hierarchical-keyword file naming scheme that the authors use as the baseline alternative; the paper argues the ontology approach goes beyond what such a scheme can express.","marker":"[15]"}],"fun_headline_variants":["Slim CIDOC-CRM subset documents vintage computing artifacts","Minimal ontology model for community-led computer heritage","Cassette tapes and software cataloged with lean CIDOC-CRM","Taming computer history with a handful of CIDOC-CRM classes","Cross-collection knowledge without the CRM complexity"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The survey's 20 participants, recruited through invite-only hobbyist channels, are assumed to represent the broader community of archive users and volunteer archivists.","fun_headline_variants_meta":{"raw":{"variants":["Slim CIDOC-CRM subset documents vintage computing artifacts","Minimal ontology model for community-led computer heritage","Cassette tapes and software cataloged with lean CIDOC-CRM","Taming computer history with a handful of CIDOC-CRM classes","Cross-collection knowledge without the CRM complexity"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000173,"raw_usage":{"total_tokens":1207,"prompt_tokens":805,"completion_tokens":402,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":421,"completion_tokens_details":{"reasoning_tokens":322}},"tokens_in":421,"tokens_out":402,"duration_ms":4562,"temperature":1.0,"reasoning_tokens":322,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T16:59:09.337851+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"If volunteers who are not already fluent in ontologies cannot correctly encode a simple cassette-digitization workflow using the proposed minimal CIDOC-CRM model (e.g., higher error rates or longer completion times than with a simple file-naming scheme), the claim that this is an effective minimal building set would be refuted.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the comparison with CHARM and asserts that CIDOC-CRM is deployable without a domain-specific extension from the start."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"A broader international study by the authors that this survey builds on, establishing the diverse interests of archive users and preservationists."},{"cited_title":"In: Proceedings of CIDOC 2011 Knowled ge Management and Museums Conference, Sibiu, Romania (2011)","cited_arxiv_id":null,"evidence_quote":"Documents applicability of CIDOC-CRM in digital libraries, serving as evidence of successful deployments in related settings."},{"cited_title":"Semantic W eb 14, 553–584 (04 2023)","cited_arxiv_id":null,"evidence_quote":"Offers a strategy for representing archive metadata in CIDOC-CRM and enabling knowledge discovery, supporting the paper's claim that integration and knowledge creation are feasible."},{"cited_title":"Frontiers in Artiﬁcial Intelligence and Applications (2020)","cited_arxiv_id":null,"evidence_quote":"Provides an ontological analysis and modularization of CIDOC-CRM, backing the idea that a minimal subset can be selected without losing coherence."},{"cited_title":"Neural Lyapunov Model Predictive Control: Learning Safe Global Controllers from Sub-optimal Examples","cited_arxiv_id":"2002.10451","evidence_quote":"Describes a hierarchical-keyword file naming scheme that the authors use as the baseline alternative; the paper argues the ontology approach goes beyond what such a scheme can express."}],"review_version":1}