{"id":"7778f9c8-9ef5-4c5d-b15b-84c54abcb313","arxiv_id":"1908.09080","paper_version":5,"verdict":"REJECT","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"high","formal_verification":"none","parameter_count":4,"one_line_summary":"DAST defines semantic complexity via a six-tuple formal system over lattices of intuitions and reports human experiments where its judgments match majority human rankings.","lead":"This paper introduces DAST, a formal model that treats the meaning of a text as a lattice of intuitions and defines semantic complexity as a calculation over that lattice. It reports human-judgment experiments claiming the model matches people's complexity ratings better than random chance.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central DAST measure is produced by human experts, not by the model; the reported correlations therefore do not demonstrate that DAST decides semantic complexity.","rationale":"The reader's weakest_assumption and my concern align: the operationalization of semantic complexity via DASTEX is manual and non-reproducible. I agree with the REJECT verdict. This concern is load-bearing rather than an implementation detail because the abstract and Section 6.2 claim the model decides; the experiments either use hand-authored Semantic Logic for one sentence family or manual DASTEX counts. In both cases, the human who creates the rules or counts already embodies the common-sense judgments that are later treated as ground truth. An axiomatic system agreed to by 88% of participants will, by construction, produce conclusions those participants endorse; that is a consistency check on the axioms, not evidence that the model autonomously measures semantic complexity. The strongest independent evidence would be a pre-registered rule set and automated pipeline applied to new texts, with DAST scores compared against held-out human judgments. The paper does not provide this, and it explicitly reports human enumeration. Therefore the central claim is not supported; the framework may be a useful scaffold, but the current evaluation is circular and selective.","tokens_in":22242,"tokens_out":5014,"duration_ms":55025,"concrete_test":"Recruit at least five annotators who have not seen the human vote-values or readability labels. Give them only Definition 6 and the 32 Scanpath paragraphs, ask each to enumerate involving semantic theories and compute DASTEX independently, then compute inter-annotator agreement (e.g., Krippendorff's alpha) and re-estimate the correlation between independent DASTEX and both fixation-based DR and readability labels. If agreement is below the conventional cutoff or the R2 collapses, DASTEX is not reproducible and the headline evidence is unsupported; if agreement is high and the regression persists, the manual-enumeration concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that DAST can 'decide' semantic complexity rests on DASTEX (Definition 6), the count of a text's involving semantic theories. The paper never supplies an operational rule for identifying a semantic theory. In §6.3.1, 'a human expert analyzed the paragraphs to enumerate involving semantic theories'; in §6.1, one Semantic Logic took 'about 10 man-hours' to elicit. Thus the DAST values used in the headline R2=0.83 (Figure 18) and in the Scanpath difficulty-ratio comparisons are human annotations, not outputs of a decision procedure. This creates a circularity: human-judgment votes are compared with numbers that were themselves produced by human judgment under the DAST vocabulary. The result is fragile in another way: Hypothesis 2 fails on all 16 data points, and only after splitting by genre and dropping an outlier do R2=0.98/0.96 appear (§6.3.1.1). Since the formal 6-tuple (Definition 1) leaves P, SA, V, and CA unspecified, the framework does not constrain these hand-made choices. The paper's Java-automated deduction for SUS-2 is a limited reproducibility point, but it does not apply to the corpus-based DASTEX measurements.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces the DAST model, a framework for measuring the semantic complexity of text. The model is built on an intuitionistic view of semantics: a text's meaning is represented as a lattice of symbols generated by a set of rules, and semantic complexity is defined as a value computed on this lattice. The authors give a set-theoretic formal definition of a semantic complexity tuple CM=(T, P, SA, L, V, CA), define a semantic space and semantic point, and propose DASTEX, a 'first level estimation' of complexity as the number of involving semantic theories. The evaluation consists of a worked example, three human-judgment experiments on one target sentence and its mutants, sixteen experiments on Persian poem sentences, and a corpus-based comparison against the Scanpath Complexity dataset using difficulty ratios. The authors report a 70% overall precision against a 20% random baseline in the mutation experiment, R^2=0.83 for the correlation between vote-values and DAST relative-values, and R^2=0.98/0.96 for genre-split, outlier-excluded regressions in the corpus evaluation. They also propose a Markovian model for common-sense multi-step reasoning and a gate mechanism for semantic complexity agreement.","tokens_in":22562,"tokens_out":3872,"duration_ms":38976,"significance":"If the central claim were supported—that a formal, automated procedure can decide semantic complexity of text—the work would contribute a conceptually distinctive approach to text complexity, drawing on intuitionistic logic and lattice theory rather than surface/syntactic features. The paper has some genuine positive aspects: it reports a very large human-judgment collection (more than 12,000 judgments in the first scenario, 15,000+ overall), supplies a dataset DOI, and includes a Java-implemented automatic deduction for one sentence (SUS-2). These are useful resources. However, the significance is not established as written, because the core quantity DASTEX is not computed by the model but by a human expert, and the formal components of the 6-tuple are left unspecified, so the reported correlations do not demonstrate that DAST itself decides semantic complexity. The evaluation also relies on post hoc splits and outlier removal to obtain its headline corpus results. The paper is better read as a report of an expert-assisted framework with preliminary correlations, not as a validated decision procedure.","major_comments":[{"comment":"The formal system CM=(T,P,SA,L,V,CA) leaves SA, V, and CA as abstract functions without concrete definitions or instantiations, so the paper does not establish that DAST itself decides semantic complexity; the decision is effectively delegated to whoever supplies these components. The claim in the Abstract and Section 6.2 that 'DAST model is capable of deciding about semantic complexity' is therefore not supported by the formal apparatus.","section":"Section 3.4, Definition 1"},{"comment":"DASTEX is computed by a human expert, not by the model: Section 6.3.1 states that 'a human expert analyzed the paragraphs to enumerate involving semantic theories,' and Section 6.1 reports that eliciting one Semantic Logic took about 10 man-hours. Consequently, the reported R2=0.83 in Figure 18 and the difficulty-ratio comparisons in Table 5 are correlations between human judgments and human-assigned indices, not between human judgments and outputs of the DAST decision procedure. This undermines the central claim of automatic semantic complexity measurement and raises a circularity concern, since both sides of the comparison originate in human judgment.","section":"Section 6.3.1 and Definition 6"},{"comment":"Hypothesis 2 is initially a null result on all 16 data points; the supporting regressions appear only after splitting the data by genre into two 8-point classes and, for Class 2, excluding one outlier, yielding R2=0.98 and R2=0.96. With only 8 points per class, a post hoc split combined with outlier exclusion cannot support the claim that 'the general claim of Hypothesis 2 has been supported by the results of this experiment.' The degrees of freedom used in this analysis are not reported or corrected, and the result is fragile enough that the corpus evaluation does not validate DAST as a general measure.","section":"Section 6.3.1.1, Figure 20"},{"comment":"In the mutation experiment, the Semantic Logic for SUS-2 was hand-crafted so that the deduction yields semantic loads such as Wonder and Engagement, and participants were first asked to agree with the axioms before making judgments. The reported 70% precision against a 20% random baseline therefore partly re-states the axioms; the experiment does not independently validate the model's complexity ordering, because the axioms already encode the expected conclusion. A more convincing test would fix the Semantic Logic in advance or apply it to a held-out set of sentences without prior participant agreement.","section":"Section 6.2.1"}],"minor_comments":[{"comment":"The Flesch-Kincaid formula is presented with ellipses and garbled symbols (RGFK = 0.39... + 11.8...); the formula should be typeset correctly for reproducibility.","section":"Section 2"},{"comment":"The text states that 'Precision for Overall Result has a Mean 58%,' but Table 3 reports Avg.=61% with STDEV=13; the text and table are inconsistent and should be reconciled.","section":"Section 6.2.2"},{"comment":"The formula for Difficulty Ratio (DR) is missing from the manuscript; only the verbal definition appears, so the reader cannot verify the computed DR values.","section":"Section 6.3.1, Definition 9"},{"comment":"The Markovian model introduces transition probabilities αi and βi without defining them in the text; these symbols are central to the proposed model and require explicit definitions and constraints.","section":"Section 6.2.1, Figure 15"},{"comment":"The Abstract claims that DAST 'competes with other methods by less error percentage,' but the evaluation reports error percentages of DASTEX relative to three reference measures (Fixation-Time, Word-Count, Readability-Level), not a head-to-head comparison against other readability methods; the claim should be restated accordingly.","section":"Abstract and Section 6.3.1.3"}],"recommendation":"reject","confidential_remarks":"The paper is better framed as a proposal for a human-in-the-loop semantic complexity framework with preliminary evidence, not as a validated automatic decision procedure. If the authors resubmit, they would need to provide concrete instantiations of all six components of Definition 1, report DASTEX without post hoc splitting or clearly label it as expert annotation, and address the circularity in the mutation experiment. The dataset and the Java-automated deduction for SUS-2 are useful contributions, but the current central claim is not supported."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know this paper claims more than it delivers, but it isn't empty. The new thing is a formal framework treating semantic complexity as a construction over lattices of intuitions, plus a simple index (DASTEX) defined as the number of semantic theories involved in a text. That specific formalization and the Markovian model of multi-step reasoning seem genuinely new. The paper also reports unusually large human-judgment experiments (over 7,000 participants across scenarios). That scale is real.\n\nWhat it does well: the 6-tuple formal system (Definition 1) is an honest attempt to make the components explicit, and the authors are transparent about needing to hand-craft Semantic Logic rules per text — they say one logic took about 10 man-hours. The Java-automated deduction for SUS-2 is a concrete reproducibility point for that one example.\n\nThe soft spots are not minor. The central measurement, DASTEX, is described in Definition 6 and then in the corpus evaluation is actually computed by a human expert counting semantic theories. No operational rule is given for identifying a theory, so the numbers behind the headline R²=0.83 and the corpus comparisons are human annotations dressed in formal notation. On top of that, the Semantic Logic axioms in the mutation experiment are hand-authored to derive exactly the semantic loads (Wonder, Engagement) that participants are asked to judge, and the participants are asked to agree with those axioms — so part of the agreement is baked in. The corpus evaluation is also fragile: Hypothesis 2 fails on all 16 data points before the authors split by genre and drop outliers to get R²=0.98/0.96. That is post-hoc analysis, not confirmation.\n\nEven so, I would not dismiss the paper. The mutation experiment against a 20% random baseline is a genuine signal, even if only on one sentence family. The framework is a plausible starting point for future automated work. The right verdict is 'major revision': require the authors to either automate DASTEX or provide a clear annotation protocol with inter-annotator agreement, and to report the failed hypothesis without post-hoc filtering as the primary result. It deserves a serious referee, but it is not ready as is.","headline":"The formal framework is novel and the experiments are large, but DASTEX is a human expert count, so the claim that DAST 'decides' semantic complexity is not supported as stated.","tokens_in":23020,"tokens_out":2205,"would_cite":false,"duration_ms":21859,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Semantic complexity can be decided by building a lattice of intuitions, the DAST model claims, and human judgments agree with it in large experiments.","keywords":["semantic complexity","intuitionistic semantics","semantic lattice","text complexity","human judgment","readability assessment","DAST","semantic theories"],"falsifier":"Have several independent annotators enumerate the involving semantic theories for the 32 Scanpath paragraphs and the 80 Persian sentences; if their DASTEX counts disagree substantially, or if DASTEX fails to correlate with readers' complexity judgments and fixation times on a new set of texts spanning the same genres, the claim that DAST decides semantic complexity would be falsified.","tokens_in":22037,"feed_emoji":"🧠","tokens_out":7117,"duration_ms":68934,"temperature":0.7,"pith_summary":"The paper sets out to establish that semantic complexity—how much meaning-related work a text demands—can be decided by a formal, intuitionistic model rather than by word counts, syntax, or readability formulas. Meaning is modeled as a lattice grown from basic intuitions by symbolic rules, and a text's semantic complexity is a computed value on that lattice. The paper reports that DAST's complexity rankings track human comparative judgments: in a mutation experiment with 3,198 participants, DAST agreed with the majority human choice on the overall most complex sentence in 70% of cases, against a 20% random baseline, and across 80 sentences vote-values correlated with DAST's relative complexity values at $R^2=0.83$. A simpler corpus-facing index, DASTEX (the number of semantic theories a text invokes), is claimed to have distinction power close to fixation time and word count on a readability corpus. If correct, the model gives a principled bridge from theories of meaning to computable text-complexity scores.","feed_headline":"Human complexity rankings match a lattice model 70% of the time","feed_subtitle":"DAST's intuition-lattice scores echo votes from 3,198 readers four times better than chance, with $R^2=0.83$ across 80 sentences.","key_machinery":"The load-bearing object is the semantic lattice: a directed structure of symbol strings produced by applying Semantic Logic rules whenever a rule's left side is present in working memory. Rules give the lattice its edges, and a valuation function $V$ assigns each node a number based on its predecessors; the complexity calculation algorithm $CA$ then turns the valued lattice into a number. Semantic items with many combined inputs receive higher values, and the whole sentence gets an overall complexity as the distance of its semantic point from the origin in an $n$-dimensional semantic space whose axes are semantic items. The simpler DASTEX index—counting a text's involving semantic theories—is the version of the machinery used for corpus evaluation, and it is explicitly a first-level estimation rather than the full lattice calculation.","core_discovery":"The paper's central claim is that a text's semantic complexity is not a hidden quality but a computable object: for a given text, a set of principal intuitions, a Semantic Logic of derivation rules, and a valuation scheme, the construction of the semantic lattice and a calculation on that lattice yields a definite complexity value. The 6-tuple $CM=(T,P,SA,L,V,CA)$ formalizes this, and an overall sentence-level value is the distance of the text's semantic point from the origin of the semantic space spanned by its semantic items. The authors argue the model decides about semantic complexity in the sense that its judgments reproduce the comparative judgments of human readers: majority human choices in three mutation experiments matched DAST 64–78% of the time, the largest experiment showing 70% overall precision against a 20% random baseline, and vote-values for 80 sentences correlated linearly with DAST's relative complexity values at $R^2=0.83$. They further claim the pattern of human deviations from DAST follows an exponential curve, which they model as a Markovian process of multi-step common-sense reasoning, and that consensus with the axiomatized Semantic Logic has a sigmoid-like triggering effect on agreement with DAST.","pith_inferences":["The paper leaves implicit that if DASTEX enumeration were automated—say, by a classifier trained on expert-annotated semantic theories—the index would become reproducible and could be tested on much larger corpora without author intervention.","The genre split in the corpus results suggests DASTEX's relation to reading effort is not uniform across domains; one testable consequence is that a genre-aware semantic index, rather than a single global formula, may be the right target for automatic readability systems.","The sigmoid consensus-agreement curve hints at a threshold effect: below some consensus level, readers may effectively use different semantic logics, so DAST's complexity ordering would only track the majority past that threshold.","The lattice formalism also yields local node values that are never aggregated; those values could be tested directly as predictors of word- or phrase-level reading times, a prediction the paper does not make."],"forward_implications":["If DAST is correct, semantic complexity can be measured without relying on sentence length, vocabulary lists, or parse-tree features, so readability assessment could be extended to meaning-level difficulty.","A reusable Semantic Logic rule base could amortize the expert effort of building lattices, making DAST-style analysis practical for new domains once such a base exists.","The Markovian-noise model implies that human disagreement with a common-sense complexity ordering should fall exponentially with the number of deviated comparison steps, a quantitative prediction that can be checked on new judgment data.","The vote-value correlation ($R^2=0.83$) supports treating DAST's relative complexity values as an estimator of group complexity votes, with applications in text selection and simplification.","On the Scanpath corpus, DASTEX's difficulty ratio sits in the same cluster as fixation time and word count, suggesting a semantic measure can mimic both an objective and a subjective readability signal."],"supporting_citations":[{"why":"supplies the Brouwer–Heyting–Kolmogorov interpretation that underlies the paper's Semantic Logic.","marker":"(Sato, 1997)"},{"why":"provides the Scanpath Complexity Dataset and its fixation-time, readability, and formula scores used for the corpus evaluation.","marker":"(Mishra et al., 2017)"},{"why":"provides the DAST Dataset of questionnaires and vote data behind the 16 semantic-comparison experiments.","marker":"(Besharati & Izadi, 2019)"},{"why":"represents the language-modeling approach to difficulty that DAST is positioned against as a semantic alternative.","marker":"(Collins-Thompson & Callan, 2004)"},{"why":"offers an earlier neuroscientific operationalization of semantic complexity that supports treating it as an empirically measurable quantity.","marker":"(Brennan & Pylkkänen, 2010)"}],"fun_headline_variants":["Lattice model matches human text complexity 70% vs 20%","DAST computes semantic complexity, beats random by 3.5x","Text complexity as lattice: 70% human agreement, R²=0.83","Semantic complexity is computable: new model hits 70% precision","Intuition-lattice DAST predicts human complexity judgments 70%"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that a human expert can reliably enumerate all semantic theories a text involves, and that this count—or the expert-built lattice—faithfully measures semantic complexity; in the corpus study DASTEX was computed manually by the authors, so the central quantity is not yet an automated, reproducible measurement.","fun_headline_variants_meta":{"raw":{"variants":["Lattice model matches human text complexity 70% vs 20%","DAST computes semantic complexity, beats random by 3.5x","Text complexity as lattice: 70% human agreement, R²=0.83","Semantic complexity is computable: new model hits 70% precision","Intuition-lattice DAST predicts human complexity judgments 70%"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000183,"raw_usage":{"total_tokens":1379,"prompt_tokens":1073,"completion_tokens":306,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":689,"completion_tokens_details":{"reasoning_tokens":207}},"tokens_in":689,"tokens_out":306,"duration_ms":3762,"temperature":1.0,"reasoning_tokens":207,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T11:22:12.649768+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have several independent annotators enumerate the involving semantic theories for the 32 Scanpath paragraphs and the 80 Persian sentences; if their DASTEX counts disagree substantially, or if DASTEX fails to correlate with readers' complexity judgments and fixation times on a new set of texts spanning the same genres, the claim that DAST decides semantic complexity would be falsified.","supporting_citations":[],"review_version":1}