{"id":"eeff8b91-4dff-45a9-a8ba-d6ceb5ea18ef","arxiv_id":"1908.04911","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"Semantic networks of linear algebra textbooks show core-periphery structure, and the persistence of topological 'knowledge gaps' in the exposition correlates negatively with textbook ratings.","lead":"Researchers built growing semantic networks from ten linear algebra textbooks, linking mathematical concepts that appear in the same sentence. They found dense early-introduced cores, modular peripheries, and persistent topological gaps whose density correlates negatively with Goodreads ratings.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The headline rating correlation appears only after the OAAT transformation, whose random within-sentence ordering is not part of the exposition; the sentence-level filtration shows no correlation, so the result may be an artifact of that arbitrary unfurling.","rationale":"The reader's weakest assumption concerned the interpretational bridge from topological cavities to cognitive knowledge gaps. My concern is upstream and more technical: the metric used for the headline correlation is constructed via a stochastic within-sentence reordering that is not part of the exposition, and the correlation is absent at the sentence level. This does not contradict the reader's CONDITIONAL verdict, but it sharpens the condition: before interpreting the correlation as evidence about knowledge gaps, one must show it is robust to the OAAT ordering rule. The descriptive core-periphery and persistence results are supported by multiple null models and remain valuable regardless of this issue. The n=7 sample and lack of multiple-comparison correction further weaken the rating correlation, but the OAAT dependence is the more specific and decisive threat. I therefore retain the CONDITIONAL/UNCHANGED verdict while emphasizing that the proposed check should be run before the central claim is cited as evidence about learning or text quality.","tokens_in":23966,"tokens_out":8370,"duration_ms":93330,"concrete_test":"Recompute the Table S5 Spearman correlations using an OAAT variant in which, within each sentence, edges are introduced in the order of the tokens that introduce them (or, as a minimal check, edges-before-nodes instead of nodes-before-edges), keeping all other pipeline choices fixed. If the significant negative correlations for dimension 0, dimension 2, and the average do not survive this change, the headline result is an artifact of the arbitrary OAAT unfurling and should be withdrawn from the central claim.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that knowledge-gap density tracks negatively with community ratings rests entirely on correlations computed with the one-at-a-time (OAAT) filtration (Fig. 5b; Table S5). The sentence-level expositional filtration, which preserves actual text order, gives no significant correlation in any dimension (Table S5: sentence NACL dim0 rho=0.143, dim1=0.036, dim2=0.0, avg=0.071; all p>0.7). The OAAT procedure 'unfurls' each sentence by adding all newly introduced nodes first, then all newly introduced edges, in random order (Supplementary Methods). This is a modeling choice, not a property of the exposition: within a sentence, readers do not encounter each new concept in isolation before the sentence's connections are present. Consequently, OAAT dimension-0 lifetimes are inflated by sentences that introduce many terms at once, because those nodes persist as isolated components until the randomly ordered edges are added; higher-dimensional lifetimes are similarly affected. The supplement reports that OAAT NACL in dimension 2 and the average correlate with sentence count and node count (Supplementary Results, Extended correlation analysis), suggesting the metric partly measures text length or term-burstiness rather than gaps. The paper's own justification that 'long cycles should still tend to be long, under the assumption that there is relatively consistent introduction of nodes and edges throughout the texts' (Supplementary Methods) is an untested assumption, and the texts differ substantially in node count (146 to 453). Thus the load-bearing condition that the OAAT metric faithfully reflects expositional gap structure is not established; if this condition fails, the headline negative correlation is an artifact of the arbitrary OAAT ordering rather than a fact about textbook exposition.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper constructs growing semantic networks from ten linear algebra textbooks, with nodes identified by a modified RAKE keyword-extraction algorithm and edges defined by sentence-level co-occurrence. It reports three families of findings: (i) the networks display core-periphery structure with an early-introduced core and a modular periphery; (ii) the expositional filtrations show persistent homology in dimensions 0, 1, and 2, with fewer and shorter-lived cavities than a random edge-order null model and more than a node-ordered null model, interpreted as 'knowledge gaps' that are created and filled; and (iii) the density of these gaps, measured by the one-at-a-time (OAAT) normalized average cycle lifetime, correlates negatively with average Goodreads ratings of seven of the texts (Spearman rho = -0.857, -0.893, and -0.821 for dimensions 0, 2, and the average across dimensions). The paper is explicitly exploratory and provides code, multiple null models, and supplementary correlation analyses.","tokens_in":24174,"tokens_out":3386,"duration_ms":36620,"significance":"If the central claim holds, the paper would introduce a novel, fully computational way to score textbook exposition by the persistence of topological cavities in a growing semantic network, with implications for learning-enhancement design. The descriptive structural findings — early introduction of core concepts, peripheral modularity, sparse persistent homology relative to random edge order — are supported by several null ensembles and are a useful contribution to the quantitative study of educational texts. The paper also openly ships its extraction and analysis pipeline, which aids reproducibility. However, the most prominent claim, the negative correlation between knowledge-gap density and community ratings, rests on a small sample (n=7), a nonstandard OAAT filtration that is not part of the actual exposition, and an interpretive bridge from topological holes to cognitive 'knowledge gaps' that is asserted rather than validated. The result is therefore best viewed as an intriguing hypothesis in need of further testing rather than an established fact.","major_comments":[{"comment":"The statistical basis for the rating correlation is thin: only seven textbooks have Goodreads ratings, the number of ratings per text ranges from 6 to 891 (Table S6), and the average rating across editions is used as a proxy for the specific edition analyzed. The Spearman tests in Table S5 are not corrected for multiple comparisons despite at least eight hypothesis tests being performed; under a Bonferroni correction for eight tests, only the dimension-2 OAAT p-value (0.00681) would remain below a 0.05 threshold, and the dimension-0 and average p-values would not. The paper acknowledges the small sample in the Results, but the abstract and Discussion present the correlation as a central finding ('density of these gaps tracks negatively with community ratings') without the same caveat. I recommend adding an explicit statement of the number of comparisons, a corrected p-value or false-discovery-rate analysis, and a more prominent caveat that these are exploratory correlations.","section":"Results, 'Evolving structure and text properties'; Supplementary Methods; Table S5"},{"comment":"The interpretive claim that topological cavities in the clique complex of a sentence co-occurrence network are 'exactly the knowledge gaps we seek' (Methods) is an assertion, not a validated equivalence. The Discussion subsequently concedes that the study 'did not deal explicitly with differential learnability of texts or in how knowledge gaps might affect the learning process.' Since the entire rating analysis depends on this bridge, the paper should distinguish the descriptive topological findings from the cognitive interpretation, either by validating the bridge (e.g., with human judgments of conceptual gaps or learning outcomes) or by clearly labeling the knowledge-gap interpretation as a hypothesis in the abstract and results, not a demonstrated fact.","section":"Methods, 'Persistent homology'; Discussion, 'Methodological considerations'"}],"minor_comments":[{"comment":"The main text refers to the 'one-at-a-time (OAAT) filtration' without defining it; the definition appears only in the Supplementary Methods. A one-sentence description should be given in the main text, along with a statement that this filtration is an analytic convenience for comparing across models with different time granularities, not a model of reading order.","section":"Results, 'Expositional development of knowledge gaps'"},{"comment":"The caption states 'introduction, persistence, and death of cycles introduced throughout exposition'; in the language of persistent homology, bars are born and die rather than being introduced, and the word 'introduced' is used inconsistently elsewhere for nodes and edges. Please use consistent terminology (birth/death) throughout.","section":"Fig. 4 caption"},{"comment":"The definition of D_k includes the convention that infinite-lifetime intervals are set to d_i = N+1; this convention is important for interpreting the metric and should be stated in the main text rather than left implicit.","section":"Methods, 'Persistent homology'"},{"comment":"The supplement reports that the number of Goodreads ratings does not correlate with the average rating (rho=0.464, p=0.294), but the analysis does not report the date of data collection or the exact edition-matching procedure. Adding these details would help readers assess the stability and reproducibility of the rating variable.","section":"Supplementary Results, 'Extended correlation analysis'"}],"recommendation":"major_revision","confidential_remarks":"The paper is honest about its exploratory nature and provides a useful computational pipeline, but the central claim about ratings is currently built on the OAAT filtration, which is a non-obvious and under-validated transformation. If the authors can show the correlation is robust to the within-sentence ordering (or to text-length controls) and temper the framing accordingly, the manuscript could be publishable. If not, the headline claim should be removed and the paper reframed as a descriptive topological study. The self-citations to Refs [26] and [50] are not problematic in themselves, but the 'knowledge gap' framing draws heavily on those prior results, and the referee report should ensure the novelty claim is not overstated."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The descriptive core of this paper is genuine and well executed. The authors build growing semantic networks from ten linear algebra textbooks, show that core concepts are introduced early while peripheral concepts come in modular bursts, and use persistent homology with several null models to argue that real expositions create and fill topological gaps more sparingly than random edge order. Those findings are new for whole-book expositions, the null-model comparisons are thoughtful, and the paper is candid about its pipeline being heuristic and its conclusions being exploratory.\n\nThe problem is the headline claim: that the density of knowledge gaps tracks negatively with community ratings. That correlation only appears after the one-at-a-time (OAAT) transformation, which unfurls each sentence by adding all new nodes first, then all new edges, in random order. The actual sentence-level filtration, which preserves the order readers encounter, shows no significant correlation in any dimension (Table S5). The OAAT ordering is a modeling choice made so that text filtrations can be compared with node-ordered null models; it is not part of the exposition. The supplement also reports that OAAT NACL correlates with sentence count and node count, so the metric partly measures text length or term burstiness, not necessarily knowledge gaps. On top of that, the rating analysis uses seven books, no multiple-comparison correction, and Goodreads ratings as a proxy for learning quality.\n\nSo the interpretive bridge is weak: topological cavities in a co-occurrence clique complex are asserted to be knowledge gaps, and the only evidence tying them to perceived text quality rests on an arbitrary unfurling. I would not call the rating result an artifact with certainty, but the paper has not established that OAAT faithfully reflects expositional structure, and the internal evidence (sentence-level null result, length correlations) points that way.\n\nWhat holds up: the core-periphery and community findings, the early-core/late-periphery development, and the comparison of persistent homology against random-index, random-sentence, random-edge, and node-ordered nulls. Those are reproducible in principle (code is on GitHub, though no commit hash and no cleaned text data). This is a useful exploratory study for network scientists working on text structure, and possibly for mathematics-education researchers, as long as the rating claim is treated as unvalidated.\n\nA serious referee should engage with this paper, but the rating-correlation section needs major revision: either test it against the sentence-level filtration, control for text length, or frame it as a hypothesis-generation finding with explicit caveats. I would not cite the rating result in my own work; I might cite the descriptive structural findings.","headline":"Solid descriptive network-science findings on textbook structure, but the headline Goodreads-rating correlation is likely an artifact of the OAAT filtration and should not be taken at face value.","tokens_in":24856,"tokens_out":1536,"would_cite":false,"duration_ms":18913,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that the evolving semantic networks of linear algebra textbooks contain topological knowledge gaps, and that longer-lived gaps correlate negatively with reader ratings, especially when the filtration is examined one step…","keywords":["semantic networks","persistent homology","linear algebra textbooks","knowledge gaps","core-periphery structure","topological data analysis","textbook exposition","co-occurrence networks"],"falsifier":"Take a set of linear algebra textbooks with known student exam performance or concept-map test scores, run the same concept-extraction and OAAT persistent-homology pipeline, and check whether normalized average cycle lifetimes still correlate negatively with those learning measures. If the correlation vanishes, reverses, or is explained by a confound such as text length, the interpretive bridge between holes and knowledge gaps fails even if the original rating correlation was real.","tokens_in":23667,"feed_emoji":"🧮","tokens_out":5648,"duration_ms":52328,"temperature":0.7,"pith_summary":"This paper argues that the way mathematical knowledge is arranged in textbooks can be measured as a growing network, and that the topology of that network tracks how well readers receive the book. Working from ten linear algebra textbooks, the authors build a network whose nodes are concepts and whose edges are sentences that mention two concepts together, then let the network grow in the order the text presents it. They claim that this growth has a consistent architecture: a dense core of foundational concepts appears early, while a sparse, modular periphery is introduced more evenly. Using persistent homology, they identify 'knowledge gaps' as topological cavities that are born, persist, and later fill as the text proceeds, and they report that longer-lived gaps, measured as normalized average cycle lifetime, correlate negatively with community ratings of the book, particularly at the sub-sentence scale. If correct, the result makes textbook exposition quantitatively comparable and opens the door to designing orderings that minimize persistent gaps.","feed_headline":"Math texts with fewer long-lived knowledge gaps earn higher ratings","feed_subtitle":"Holes in a textbook's concept network open and close as it unfolds; longer-lived holes track lower reader ratings.","key_machinery":"The central machinery is persistent homology applied to the clique complex of a growing co-occurrence network. As the text proceeds, each new sentence adds concepts (nodes) and co-occurrences (edges) to the filtration; filling in every all-to-all connected subgraph as a simplex turns the graph into a clique complex whose holes, cavities, and disconnected components can be tracked with barcodes. A hole is born when it first appears and dies when later connections fill it, and the normalized average cycle lifetime $D_k = \\frac{1}{mN}\\sum_{i=1}^{m} (d_i - b_i)$ summarizes how long such gaps persist per dimension. The one-at-a-time (OAAT) variant adds one node or edge at a time so that texts with different sentence structures and null models can be compared on equal footing. This machinery is what turns prose order into a quantitative 'gap density' that can be correlated with reader ratings.","core_discovery":"The central discovery is a negative association between the persistence of topological cavities in the evolving semantic network and the community rating of the textbook. In the paper's one-at-a-time filtration, which adds one node or edge per step to remove sentence-level granularity, the normalized average cycle lifetime correlates negatively with average rating in dimension 0 (Spearman rho = -0.857, p = 0.0137), dimension 2 (rho = -0.893, p = 0.00681), and the mean over dimensions 0, 1, and 2 (rho = -0.821, p = 0.0234). The authors interpret these cavities as knowledge gaps: places where concepts are not yet connected to the growing network. The texts also systematically fall below null models in the number and persistence of 0-dimensional gaps, suggesting that ordinary exposition deliberately connects new concepts quickly, while higher-dimensional cavities vary more across books. The paper is careful to present the rating correlation as preliminary, since only seven of the ten textbooks had enough community ratings.","pith_inferences":["A direct test the paper does not run: replace community ratings with measured learning outcomes from students reading the same texts; if the negative correlation survives, the gap metric is about learnability, not taste. The paper itself calls for classroom studies.","The relationship may be causal in the direction the authors suggest, but it could equally reflect a confound: higher-rated books may be better written in many ways that also reduce topological persistence, so the gap metric may be a proxy rather than a mechanism.","One could compute the same metrics on randomized section order within a single textbook and ask how much gap persistence changes; this would quantify how much of the signal is due to chapter ordering versus sentence-level exposition.","If the metric generalizes, it suggests a practical optimization criterion, but optimizing solely to minimize persistent gaps could suppress motivating connections, a tradeoff the authors note in their discussion of productive failure and curiosity."],"forward_implications":["If the reported correlations hold in larger samples, the persistence of topological gaps in a growing concept network becomes a measurable signal of perceived exposition quality.","Textbook authors and curriculum designers could compare alternative orderings of the same material by computing the normalized average cycle lifetime of each ordering and preferring one with fewer persistent gaps.","The finding that real texts sit between a fully connected node-ordered model and a random edge order suggests that effective exposition intentionally leaves some gaps, rather than minimizing them entirely.","The same pipeline can be applied to other well-structured expository domains, such as physics or biology textbooks, to test whether the gap-rating relationship is specific to linear algebra.","The absence of a significant correlation at sentence-level granularity, together with significance at OAAT granularity, implies that the relevant gaps are sub-sentence ordering effects, not coarse chapter-level structure."],"supporting_citations":[{"why":"Supplies the notion that growing semantic networks contain persistent gaps and motivates comparing text expositions to null models.","marker":"[26]"},{"why":"Provides persistent homology as a method for tracking the birth and death of topological cavities.","marker":"[27]"},{"why":"Gives the algorithmic foundation for computing persistent homology.","marker":"[28]"},{"why":"Reviews persistent homology theory and practice as applied to data analysis.","marker":"[29]"},{"why":"Supplies the RAKE algorithm that the paper modifies for concept extraction, defining the network nodes.","marker":"[31]"},{"why":"Provides the persistence-bar-code metric from which the normalized average cycle lifetime is adapted.","marker":"[34]"},{"why":"Software used to compute persistent homology for the empirical networks and null models.","marker":"[67]"}],"fun_headline_variants":["Long-lived knowledge gaps in math texts tie to low ratings","Math textbooks: fewer persistent concept gaps, better ratings","Semantic network gaps in textbooks predict community ratings","Textbooks with slower-filling concept gaps score lower","Math texts' knowledge gap persistence linked to ratings"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that a topological cavity in the co-occurrence clique complex is the same thing as a 'knowledge gap' that affects a reader's understanding or enjoyment, and that community star ratings measure that effect.","fun_headline_variants_meta":{"raw":{"variants":["Long-lived knowledge gaps in math texts tie to low ratings","Math textbooks: fewer persistent concept gaps, better ratings","Semantic network gaps in textbooks predict community ratings","Textbooks with slower-filling concept gaps score lower","Math texts' knowledge gap persistence linked to ratings"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000235,"raw_usage":{"total_tokens":1485,"prompt_tokens":912,"completion_tokens":573,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":528,"completion_tokens_details":{"reasoning_tokens":498}},"tokens_in":528,"tokens_out":573,"duration_ms":6321,"temperature":1.0,"reasoning_tokens":498,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T13:28:44.065370+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a set of linear algebra textbooks with known student exam performance or concept-map test scores, run the same concept-extraction and OAAT persistent-homology pipeline, and check whether normalized average cycle lifetimes still correlate negatively with those learning measures. If the correlation vanishes, reverses, or is explained by a confound such as text length, the interpretive bridge between holes and knowledge gaps fails even if the original rating correlation was real.","supporting_citations":[{"cited_title":"E., Karuza, E., Giusti, C","cited_arxiv_id":null,"evidence_quote":"Supplies the notion that growing semantic networks contain persistent gaps and motivates comparing text expositions to null models."},{"cited_title":"Topology and data","cited_arxiv_id":null,"evidence_quote":"Provides persistent homology as a method for tracking the birth and death of topological cavities."},{"cited_title":"& Morozov, D","cited_arxiv_id":null,"evidence_quote":"Reviews persistent homology theory and practice as applied to data analysis."},{"cited_title":"& Cowley, W","cited_arxiv_id":null,"evidence_quote":"Supplies the RAKE algorithm that the paper modifies for concept extraction, defining the network nodes."},{"cited_title":"& Carlsson, G","cited_arxiv_id":null,"evidence_quote":"Provides the persistence-bar-code metric from which the normalized average cycle lifetime is adapted."}],"review_version":1}