{"id":"e7791739-b0ef-4dd5-89dd-4cdc6d88083d","arxiv_id":"2501.14976","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":3.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A literature review proposes a four-way taxonomy of annotation classification mechanisms used by educational annotation tools and reports that about 85 percent of surveyed tools include some classification support.","lead":"This paper reviews 38 educational annotation tools and groups them into four types by how annotations can be classified: no classification, pre-set tags, user-extensible tags, and structured vocabularies such as ontologies. Most of the tools surveyed, by the authors' count, provide some classification mechanism, with ontologies described as the most flexible option.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper's 84.84% classification claim is unsupported as written: Section 1 says 38 tools were considered, but the percentages in Section 6 divide into 33 tools (5+13+7+8); the missing five tools are never reconciled.","rationale":"Good-faith reading: the paper is an exploratory review whose value is the four-way taxonomy with named example tools; that taxonomy is plausible and the examples are useful. However, the central claim as summarized includes a precise quantitative distribution, and that distribution is not internally reproducible. The reader's verdict of CONDITIONAL is appropriate, and my concern does not move it; it strengthens the condition that the sample table and percentages be corrected. I agree partially with the reader's weakest_assumption: the unverifiable search is a real limitation, but the more decisive problem is the arithmetic mismatch between the stated sample size and the displayed percentages, which is independent of any external representativeness debate. The concrete check is a simple recount of the paper's own tables and text; it settles whether the reported distribution is self-consistent. If the recount confirms 33, the paper's quantitative summary is wrong as written; if it reveals unlisted tools in Section 6, those omissions should be documented. The taxonomy itself would remain conditionally acceptable.","tokens_in":11462,"tokens_out":4791,"duration_ms":43320,"concrete_test":"Recount every distinct annotation tool named in Sections 2–5 and in Tables 1–5 (including PDF Annotator and the Chang et al. clustering tool in §3.1). Compare the count with the '38' in Section 1 and with the denominators implied by each percentage in Section 6. A correct reproduction must produce one of: (a) a list of all 38 tools with one category each, yielding the stated percentages; or (b) corrected percentages based on the actual count, with the '38' text fixed. If the percentage values require a 33-tool denominator, the source of the extra five tools must be identified or the headline '84.84%' withdrawn.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative claim—that 84.84% of analyzed annotation tools use some classification mechanism—depends on the denominator and category counts. Section 1 states '38 different tools have been considered', but the four percentages in Section 6 (15.15%, 39.39%, 21.21%, 24.24%) are exactly the values produced by 5/33, 13/33, 7/33, and 8/33. Counting the tools actually listed in Tables 1–5 gives 33 rows (5 without classification, 6 style-tag tools, 7 semantic-tag tools, 7 folksonomy tools, 8 ontology tools), not 38. Even including the clustering tool mentioned in §3.1 and PDF Annotator mentioned in §3.1 text reaches only 35. No table or explanation reconciles the 38-tool claim with the 33-tool denominator used in the percentages. Because the taxonomy's breadth and the reported distribution are the paper's main results, an unexplained five-tool discrepancy is a load-bearing internal inconsistency: either the sample is 33 and the text '38' is wrong, or the sample is 38 and all percentages must be recomputed. The paper also does not specify which tools fall in which category for the missing five, so the reader cannot independently reconstruct the distribution.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a review of annotation tools used in educational settings, focusing on how annotations are classified. It proposes a four-way taxonomy: tools with no classification mechanisms; tools with pre-established vocabularies (style tags and semantic tags); tools with extensible vocabularies (folksonomies); and tools with structured vocabularies (taxonomies, thesauri, and mainly ontologies). The paper assigns 38 (or, per the tables, 33) tools to these categories and reports that 84.84% of the analyzed tools use some classification mechanism, with 15.15% using no classification, 39.39% using controlled vocabularies, 21.21% using folksonomies, and 24.24% using ontologies. It then discusses the educational implications of each category and concludes that ontologies are the preferred structured vocabulary.","tokens_in":11734,"tokens_out":6438,"duration_ms":48733,"significance":"The taxonomic framework is simple and could be useful to researchers and practitioners as an initial orientation to annotation-tool design. The paper's qualitative descriptions of individual tools are informative, and the four categories are intuitively meaningful. However, the contribution is primarily descriptive: the central quantitative claim (84.84%) is not currently verifiable because the sample size is inconsistent, the search protocol is not described, and some tools mentioned in the text are not counted in the tables. If the sample and counts are corrected, the distribution claim may still hold, but as written the paper does not provide enough evidence for its headline numbers.","major_comments":[{"comment":"Section 1 states that '38 different tools have been considered,' but the percentages reported in Section 6 (15.15%, 39.39%, 21.21%, 24.24%) correspond exactly to 5/33, 13/33, 7/33, and 8/33, and the rows in Tables 1-5 sum to 33 tools. The five-tool discrepancy is never resolved. Because the paper's central claim is that 84.84% of tools use some classification mechanism, the denominator matters: if the sample is 38, all percentages must be recomputed; if it is 33, the '38' in Section 1 is wrong. Please provide a complete list of all tools considered, with their category assignments, and ensure that the counts, tables, and percentages are consistent.","section":"Section 1 vs Section 6 and Tables 1-5"},{"comment":"The description of the search ('an exhaustive bibliographic search in several current reviews of this topic was conducted, as well as searches in repositories of academic articles') provides no search strings, date range, inclusion/exclusion criteria, or coding rules. Without this information, the reader cannot judge whether the sample is representative or complete, and the reported distribution is not reproducible. Please add a methodology paragraph or appendix detailing the search process, the screening criteria, and the rules used to assign tools to categories.","section":"Section 1 (search protocol)"},{"comment":"Section 5 states that the paper 'focuses on the tools that use ontologies' among structured vocabularies, yet Section 7 concludes that 'of the possibilities available (taxonomies, thesauri, and ontologies), ontologies are preferably used.' If the search deliberately limited the structured-vocabulary category to ontologies, the conclusion about preference is not supported; if the search covered all three types and found only ontologies, that should be stated explicitly. Please clarify the scope and adjust the conclusion accordingly.","section":"Section 5 and Section 7"}],"minor_comments":[{"comment":"Table labels are inconsistent: the semantic-tag table is labeled 'Table 2' (should be Table 3), the folksonomy table is labeled 'Table 4', and the ontology table is labeled 'Table 3' (should be Table 5). Please renumber all tables and update in-text references.","section":"Tables"},{"comment":"Section 3.1 mentions PDF Annotator and the clustering tool of Chang et al. (2015) as examples of style-tag tools, but neither appears in Table 2; clarify whether these are included in the counts and, if so, add them to the table or explain their omission.","section":"Section 3.1"},{"comment":"The bibliography contains duplicate entries (e.g., Kawase et al. 2009, Jan et al. 2016, Chen et al. 2014, Hwang et al. 2007 appear twice) and inconsistent citation formats; please deduplicate and normalize.","section":"Bibliography"},{"comment":"There are several typos and stylistic errors, e.g., 'users them to highlight' (Section 3.1), 'which makes it possible allows to classify' (Section 4), and 'about the focus is placed' (Section 6); a thorough language edit is recommended.","section":"Language"}],"recommendation":"major_revision","confidential_remarks":"The paper is more of a position/survey note than a systematic review; its fit to the journal depends on whether the journal publishes exploratory surveys. The authors' own prior work (Gayoso-Cabada et al. 2018) is the direct predecessor, which is fine, but the novelty beyond that conference paper should be clarified. Also, the sample is largely drawn from references around 2004-2018, despite the 2025 submission; recent tools (e.g., Hypothesis, Perusall) appear absent, which may reflect the search limits. Editors may wish to consider whether the missing methodology can be supplied in revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThis is a survey of annotation tools in education, grouped into four buckets: no classification, fixed vocabularies, folksonomies, and ontologies. The taxonomy makes sense and the tool examples are handy, but the novelty is thin: the four-case scheme comes from the authors' own 2018 paper, and this version mostly adds more examples and discussion.\n\nThe real selling point is the clean presentation. Each category is clearly defined, the tables give a quick overview, and Section 6's discussion of what each vocabulary type means for teaching is thoughtful. The paper is also honest about being an initial study.\n\nThe soft spots are real. Section 1 says 38 tools were considered, but the tables only list 33, and the percentages in Section 6 are exactly 5/33, 13/33, 7/33, 8/33. The missing five tools are never accounted for. So the headline number, 84.84%, is wrong or at least unverifiable. There is no search protocol: no inclusion/exclusion criteria, no date range, no coding rules. And the bibliography is mostly pre-2018, which is odd for a 2025 submission. Minor issues include two tables with the same number and a mislabeled ontology table.\n\nWho is this for? An educator or developer wanting a compact way to compare annotation tools will get some value. A researcher reading for theoretical novelty won't. It's a plausible starting taxonomy, not a strong empirical claim.\n\nI'd send it to peer review, but with a request for major revision. The referee should push for an explanation of the 38/33 discrepancy, a clear search protocol, and an updated tool list. If those are fixed, it becomes a usable short survey.","headline":"A sensible four-bucket taxonomy of annotation tools, but the paper's own 38-vs-33 count makes the central percentage unverifiable.","tokens_in":12223,"tokens_out":3681,"would_cite":false,"duration_ms":32266,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This review claims that educational annotation tools fall into exactly four cases according to how they classify annotations, and that 84.84% of the 38 tools examined have some classification mechanism.","keywords":["annotation tools","annotation classification","folksonomy","ontology","educational technology","controlled vocabulary","structured vocabulary","review"],"falsifier":"Build a systematic, reproducible corpus of educational document-annotation tools with explicit search queries, inclusion criteria, and a coding scheme, then compare the share with no classification mechanism and the number of tools that do not fit the four cases; if the no-classification share differs substantially from 15.15% or a meaningful class of hybrid or non-document tools appears, the four-case taxonomy and its distribution are not robust.","tokens_in":11286,"feed_emoji":"🏷️","tokens_out":7762,"duration_ms":62287,"temperature":0.7,"pith_summary":"This paper tries to establish that classifying annotations is a standard, not marginal, feature of software tools for educational reading. Reviewing 38 document-annotation tools, it proposes that each tool belongs to one of four cases: no classification mechanism, a closed set of predefined tags, an open user-extensible tag vocabulary (a folksonomy), or a structured vocabulary such as a taxonomy, thesaurus, or mostly an ontology. The authors report that 84.84% of the surveyed tools provide some classification mechanism, with 39.39% using pre-established vocabularies and 24.24% using ontologies. They argue that classification matters because it is what allows teachers and students to exploit annotations to reveal content comprehension, annotation styles, and intellectual maturity.","feed_headline":"84.84% of annotation tools classify annotations","feed_subtitle":"A review of 38 tools finds fixed tags, folksonomies, and ontologies are the standard ways to organize annotations.","key_machinery":"The central machinery is the four-case classification of annotation tools by the structure of the vocabulary used to label annotations, ordered from no vocabulary to increasingly structured vocabulary. Each case determines what educational information can be extracted: no classification yields only raw annotations, flat controlled vocabularies allow style or semantic tagging and simple clustering, uncontrolled folksonomies measure students' creative tagging but suffer from synonymy, and structured vocabularies, especially ontologies, allow relationships among tags to be inherited and exploited, supporting recommenders, annotation models, and even extraction of new conceptual structures. This ordering is what carries the review's argument that classification, not annotation itself, is the feature that turns annotation activity into learning evidence.","core_discovery":"On the paper's own terms, the discovery is that the classification mechanisms of educational annotation tools can be organized into a small, principled taxonomy of four types. Tools without classification mechanisms (15.15% of the 38) simply record annotations; tools with pre-established vocabularies (39.39%) label annotations with a closed set of style or semantic tags; tools with extensible vocabularies (21.21%) let annotators create and share their own tags, forming folksonomies; and tools with structured vocabularies (24.24%) label annotations with concepts from taxonomies, thesauri, or, in practice, ontologies. The paper further claims that among structured vocabularies, ontologies dominate because they can be chosen to fit the content, inherit relationships through subject-predicate-object triplets, and give teachers the richest view of how students annotate.","pith_inferences":["My inference: the four-case scheme works as a design checklist for new educational annotation tools, since choosing a vocabulary structure determines in advance what learning evidence the tool can harvest.","My inference: because the surveyed references stop around 2018, how these four cases play out in newer collaborative platforms and ebook ecosystems is an open empirical question that the paper does not settle.","My inference: the high share of ontology-based tools may reflect research prototypes rather than typical classroom adoption; the paper reports that tools exist, not how widely any of the four types is used.","My inference: a natural test is to apply the same four-case coding to annotation tools for non-document content (video, audio, images), which the authors explicitly leave to future work."],"forward_implications":["Because 84.84% of the surveyed tools classify annotations, classification should be treated as a standard capability in educational annotation tools rather than an optional add-on.","In tools with pre-established vocabularies, the flat structure of the tags limits exploitation to clustering terms, so complex annotation models are out of reach.","Folksonomy tools support measuring students' creative and reflective choice of terms, but the open vocabulary makes it hard to find annotation models when different terms mean the same thing.","Ontology-based tools give teachers the most information, including relationships deduced from the ontology, and can reveal a student's reflective capacity and maturity.","As annotations become content-rich components of ebooks and interactive fiction, more semantic classification systems will be needed to exploit them."],"supporting_citations":[{"why":"Supplies the notion of annotation technology and the folksonomy label used to define the extensible-vocabulary category.","marker":"(Ovsiannikov et al, 1999)"},{"why":"Provides the taxonomy-thesaurus-ontology structure distinction that defines the structured-vocabulary category.","marker":"(Kalboussi A et al, 2015)"},{"why":"Frames the educational analysis of controlled versus uncontrolled and structured vocabularies in the Discussion.","marker":"(Kalboussi A et al, 2016)"},{"why":"Supports the claim that annotations can be exploited for patterns and recommenders, motivating why classification matters.","marker":"(Novak et al, 2012)"},{"why":"Earlier version of this study on annotation classification mechanisms that the present review extends.","marker":"(Gayoso-Cabada et al, 2018)"},{"why":"Grounds the educational value of analyzing student annotations for content comprehension.","marker":"(Cigarran et al, 2014)"}],"fun_headline_variants":["38 annotation tools, 4 ways to classify","Annotation tools: four classification styles","Review: how annotation tools sort student tags","Education tools: 4 types of annotation classification","From fixed tags to ontologies: annotation schemes"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the 38 tools found by the authors' informal search are a representative and complete picture of educational annotation tools; the paper gives no search strings, inclusion or exclusion criteria, or date range, so if tools were missed, the four-case classification and the reported percentages would lose support.","fun_headline_variants_meta":{"raw":{"variants":["38 annotation tools, 4 ways to classify","Annotation tools: four classification styles","Review: how annotation tools sort student tags","Education tools: 4 types of annotation classification","From fixed tags to ontologies: annotation schemes"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000634,"raw_usage":{"total_tokens":2956,"prompt_tokens":1006,"completion_tokens":1950,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":622,"completion_tokens_details":{"reasoning_tokens":1893}},"tokens_in":622,"tokens_out":1950,"duration_ms":14803,"temperature":1.0,"reasoning_tokens":1893,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T14:44:08.614566+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Build a systematic, reproducible corpus of educational document-annotation tools with explicit search queries, inclusion criteria, and a coding scheme, then compare the share with no classification mechanism and the number of tools that do not fit the four cases; if the no-classification share differs substantially from 15.15% or a meaningful class of hybrid or non-document tools appears, the four-case taxonomy and its distribution are not robust.","supporting_citations":[],"review_version":1}