{"id":"0ce5a94d-99bb-469d-b5af-94c7511de98e","arxiv_id":"1908.07218","paper_version":5,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"CA-EHN contains 90,505 Chinese commonsense analogies covering 5,656 words and 763 relations, extracted by comparing E-HowNet definition graphs.","lead":"The authors present CA-EHN, a Chinese word analogy benchmark of 90,505 analogies across 763 commonsense relations, automatically extracted from E-HowNet definitions and refined by linguists. The dataset is offered as a test of whether word embeddings capture everyday world knowledge rather than only morphology or named entities.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 179% CA-EHN gain from E-HowNet retrofitting is confounded by source overlap, leaving the 'great indicator' claim without independent validation.","rationale":"The reader's formal weakest_assumption concerns E-HowNet's definition accuracy and whether single-node-difference graphs yield human-meaningful analogies. That is a valid dataset-quality concern, but the paper's linguist annotation of concept analogies and synsets partially mitigates it. The more load-bearing problem for the central claim is the circularity in the validation experiment: the dataset and the retrofitting source are both E-HowNet, so the 179% improvement could reflect shared vocabulary and structure rather than genuine commonsense sensitivity. The reader's rationale explicitly names this circularity as the primary reason for CONDITIONAL, so we are aligned in substance, but our identified concern is the validation confound rather than ontology noise. The HIT-Thesaurus experiment provides partial independent support because an external resource also improves CA-EHN more than other benchmarks, but it does not rule out a strong source-overlap effect, and the abstract's unqualified claim 'stands out as a great indicator' goes beyond what the experiments establish. Keeping the verdict CONDITIONAL is appropriate: the dataset is valuable and described in detail, but the headline claim requires a source-independent validation before it can be accepted as demonstrated. Our proposed cross-split test would directly settle whether the large gain is an artifact.","tokens_in":7382,"tokens_out":6955,"duration_ms":77359,"concrete_test":"Split the E-HowNet concept taxonomy into two disjoint partitions; extract a validation benchmark from partition A only using the CA-EHN algorithm, and retrofit embeddings using only same-taxon and hypo-hyper edges from partition B. If the retrofitted accuracy gain on this cross-split benchmark is far below the reported 179% (or near zero), the headline gain is largely an artifact of train/evaluation source overlap. Repeat with a fully independent commonsense resource (e.g., Chinese ConceptNet) as an external check.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline claim, that CA-EHN is a great indicator of how well word representations embed commonsense knowledge, is supported mainly by Section 5.4: retrofitting embeddings with the E-HowNet taxonomy raises CA-EHN accuracy by up to 179%, far more than existing benchmarks. This experiment is confounded by source overlap. CA-EHN is extracted from E-HowNet word sense definitions, and the retrofitting lexicon is the E-HowNet taxonomy. The two share a large fraction of surface words and concept nodes, so the injected knowledge directly moves vectors of words that define CA-EHN analogy questions. The smaller but still large gain from HIT-Thesaurus (88%) shows some general benefit of lexical injection, but it does not isolate whether CA-EHN measures commonsense semantics rather than vocabulary overlap with the injected resource. Without a source-independent control, the magnitude of the E-HowNet gain cannot be interpreted as evidence that CA-EHN is a valid commonsense benchmark rather than a benchmark that rewards memorizing E-HowNet's own structure.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces CA-EHN, a Chinese word analogy dataset automatically extracted from the E-HowNet ontology by comparing structured word-sense definition graphs that differ in exactly one concept node. The extraction is filtered to concrete concepts and common words, and the resulting concept analogies are partially validated by linguists (κ=0.76 on 1,000 items), yielding 90,505 word-level analogies covering 5,656 words and 763 derived relation types. The paper then evaluates distributed word embeddings on this benchmark and reports that retrofitting with the E-HowNet taxonomy improves CA-EHN accuracy by up to 179%, while retrofitting with HIT-Thesaurus improves it by up to 88%. On this basis it claims that CA-EHN is a 'great indicator' of how well word representations embed commonsense knowledge.","tokens_in":7539,"tokens_out":5785,"duration_ms":58925,"significance":"If validated, CA-EHN would be a valuable resource: it is substantially larger and relationally richer than existing Chinese analogy benchmarks, it is derived from a structured ontology rather than hand-written templates, and the dataset is publicly released. The extraction pipeline is described concretely with transparent filters, the inter-annotator agreement is reported, and the benchmark is evaluated across multiple embedding configurations. However, the central claim that CA-EHN specifically measures commonsense embedding quality is currently supported by a retrofit experiment that shares its source ontology with the benchmark, so the significance hinges on whether additional source-independent validation confirms the claim.","major_comments":[{"comment":"The central claim that CA-EHN is 'a great indicator of how well word representations embed commonsense knowledge' is supported mainly by the retrofit experiment, but that experiment is confounded by source overlap. CA-EHN analogies are extracted from E-HowNet definition graphs (Section 4.1), while the retrofitting lexicon is the E-HowNet taxonomy (Section 3.2). Because the definition graphs and the taxonomy are two views of the same ontology, the injected knowledge directly moves the representations of words and concepts that appear in the analogy questions; the up-to-179% improvement may therefore reflect the embedding's exposure to E-HowNet structure rather than a general sensitivity to commonsense knowledge. The HIT-Thesaurus result (up to 88%) is a useful partial control, but it does not rule out sharing of words or concepts with CA-EHN. I ask for a source-independent validation: for example, retrofit with a lexicon that has no ontology overlap with E-HowNet, or evaluate on a split of CA-EHN from which all words and concepts appearing in the retrofitting resource have been removed, and report the gains on that split. Without such a control, the differential gain cannot be interpreted as evidence for the benchmark's construct validity.","section":"Section 4.1 and Section 4.3"},{"comment":"The extraction criterion that two definition graphs differing in exactly one concept node form an analogy is a syntactic rule; its validity as a proxy for human commonsense analogies is only partially established. The paper reports inter-annotator agreement of κ=0.76 on 1,000 of 36,100 concept analogies, but it does not state how the 25,010 surviving concept analogies were selected from the annotations (e.g., majority vote, unanimous agreement, or some other threshold), nor the distribution of labels. It also does not explicitly state whether the final 90,505 word-level analogies were re-validated after left/right expansion, or only the concept-level analogies. Please provide the annotation guideline, label distribution, selection rule, and a clear statement of which units (concept analogies vs. word analogies) were validated. This is necessary for users to judge the reliability of the dataset's relation labels.","section":"Section 4.1 and Section 4.3"}],"minor_comments":[{"comment":"The text says 'the later is translated from English'; 'later' should be 'latter'.","section":"Section 2"},{"comment":"The '763 relations' statistic is computed by grouping word pairs into equivalence classes, not by annotating each analogy with a relation label; the manuscript should state this explicitly in the abstract and in the Table 3 caption to avoid implying that the dataset contains 763 labelled relation types.","section":"Section 5.2 and Table 3"},{"comment":"In the synset annotation example, the paper says 'the annotator' judged camellia, lavender, and iris as hyponyms; please specify how many annotators performed this step, whether the synset refinement was applied to all synsets or a sample, and report the number of synsets that were refined.","section":"Section 4.3"},{"comment":"The coverage row for CA-EHN reports 90,505 for every embedding; please clarify whether this means all CA-EHN analogy questions were fully covered by all embedding vocabularies, and if so, how this follows from the construction pipeline.","section":"Section 5.3, Table 4"},{"comment":"The example 'Beijin : Peking University' contains a typo: 'Beijin' should be 'Beijing'.","section":"Section 5.2"},{"comment":"The phrase 'infused some structure knowledge' might be better expressed as 'retrofitted with structured knowledge' for consistency with the retrofitting terminology used elsewhere.","section":"Section 5.4"}],"recommendation":"major_revision","confidential_remarks":"The dataset is a concrete contribution and the extraction pipeline is clearly described, but the headline claim depends on a retrofit experiment whose control condition shares the source ontology with the benchmark. I believe this can be addressed with additional experiments rather than rejection, and I would encourage the editor to treat the source-overlap concern as the primary revision requirement."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things worth knowing. The dataset is genuine and useful: 90,505 Chinese analogies, 763 relations, automatically extracted from E-HowNet definition graphs with linguist filtering, and it is public. That alone is a solid contribution to Chinese embedding evaluation. The second thing is that the paper's central claim that CA-EHN is 'a great indicator' of commonsense embedding quality is not backed by the experiments as reported, because the main supporting experiment (Section 5.4) retrofits embeddings with the E-HowNet taxonomy, the same ontology the benchmark was built from. The up-to-179% accuracy gain on CA-EHN is partly by construction. The HIT-Thesaurus retrofit (up to 88%) is a nice control showing that lexical injection helps in general, but it does not isolate the source-overlap effect. So the benchmark's validity as a commonsense probe remains conditional on external, source-independent validation.\n\nWhat the paper does well: the extraction pipeline is described concretely (definition expansion, graph parsing, graph comparison, left/right expansion), and the filters are sensible. The concept-analogy annotation with four annotators and κ=0.76 is real evidence of quality, and the synset re-annotation catches a genuine issue (hyponyms wrongly treated as synonyms). The benchmark statistics in Table 3 make the coverage point convincingly: existing Chinese sets have dozens of relations, mostly morphological or named-entity; CA-EHN's 763 relations are a different scale.\n\nSoft spots, in proportion. The circularity in validation is the main one; it's a real flaw but not a fatal one. No error bars, no significance tests, and no extraction code are released, which makes it harder to judge robustness or reproducibility. The frequency threshold in ASBC is stated but not analyzed for sensitivity. These are minor-to-moderate. The paper does not claim to solve the confound, but it also does not flag it; that is a fairness issue.\n\nWho should read it: anyone working on Chinese lexical semantics, word embedding evaluation, or benchmark construction more generally. It is a worthwhile resource paper. It deserves a serious referee, and I would recommend the editor send it out. In revision, the authors should add a source-independent validation (e.g., a held-out human-constructed analogy set, or a retrofit with a different ontology of comparable size) and report variance across embedding seeds or bootstrapped confidence intervals.","headline":"The CA-EHN dataset is a genuinely useful new Chinese analogy resource, but the headline 'great indicator' claim rests on a confounded retrofit experiment and needs independent validation.","tokens_in":8072,"tokens_out":2687,"would_cite":false,"duration_ms":25169,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"CA-EHN, a 90,505-item Chinese commonsense word analogy benchmark, is built by comparing definitions in the E-HowNet ontology, and retrofitting commonsense knowledge into embeddings lifts its accuracy by up to 179 percent.","keywords":["commonsense reasoning","word analogy","word embeddings","E-HowNet","ontology","Chinese","benchmark","lexical database"],"falsifier":"Take a random sample of CA-EHN analogies, remove the ontology definitions, and have native speakers judge whether the relation between the two word pairs is a meaningful commonsense analogy; if judged correctness falls far below the paper's reported annotation agreement for concept analogies, the extraction criterion is admitting spurious relations. A second check is to build a control benchmark from random single-node perturbations of the same definition graphs and see whether retrofitting improves accuracy on the control as much as it does on CA-EHN.","tokens_in":7152,"feed_emoji":"🧠","tokens_out":7259,"duration_ms":63278,"temperature":0.7,"pith_summary":"The paper tries to establish that commonsense knowledge can be reduced to word-level analogical reasoning and measured by a dedicated benchmark. It introduces CA-EHN, which it calls the first commonsense word analogy dataset, with 90,505 analogies over 5,656 Chinese words and 763 relations, extracted by comparing structured definitions in the E-HowNet ontology. Existing Chinese analogy datasets are mostly morphological or named-entity relations; CA-EHN instead draws on ordinary lexical knowledge such as sound-origin, organ-disabled, and juvenile-adult relations. The paper's central evidence is that retrofitting E-HowNet taxonomy into word embeddings raises CA-EHN accuracy by up to 179 percent while improving or hardly changing other benchmarks, which it interprets as showing that the dataset tests how well embeddings store commonsense knowledge.","feed_headline":"Ontology mining yields 90,505 commonsense analogies","feed_subtitle":"CA-EHN rewards embeddings that absorb commonsense knowledge; retrofitted vectors gain up to 179 percent.","key_machinery":"The central object is the E-HowNet definition graph: each word sense is parsed into a directed graph whose nodes are words, concepts, or functions and whose edges are attribute modifiers, so a sense like 'laboratory' becomes an InstitutePlace concept modified by a telic relation to research or experiment. The extraction mechanism is graph comparison: two definition graphs that differ in exactly one concept node yield the analogy w1:c1 = w2:c2. Synonym expansion then turns the concept analogy into word-level questions, with the right-hand answer kept as a synset so that the standard vector arithmetic v1 + v2 - v3 can be evaluated as a membership test.","core_discovery":"The paper's central claim is that a commonsense analogy can be read off an ontology by comparing definition graphs: two word senses form an analogy when their parsed definitions differ in exactly one concept node. Using E-HowNet's structured sense definitions and taxonomy, the authors extract concept analogies, expand the left concept into synonymous words to form analogy questions, and keep the right expansion as an accepted synset so embeddings can be scored by whether their nearest vector falls in that synset. After filtering for concrete concepts and frequent words and after linguist annotation of concept analogies and synsets, the result is a benchmark of 90,505 analogies covering 5,656 words and 763 relations. The paper further claims this benchmark is a useful indicator of commonsense knowledge in embeddings because injecting commonsense ontology structure via retrofitting produces much larger gains on CA-EHN than on existing analogy sets.","pith_inferences":["An implication the paper leaves implicit is that the single-node-difference criterion could be run in reverse: automatically propose new commonsense relations by clustering analogies whose differing concept pairs share a taxonomy path, turning CA-EHN into a relation-induction resource.","The benchmark's sensitivity to retrofitting suggests CA-EHN could be used diagnostically, e.g., comparing how much ontology knowledge different training objectives preserve, rather than only as a ranking of final accuracies.","A testable extension is to build a control benchmark by randomly perturbing one concept node in the same definition graphs; if retrofitting improves accuracy on that control as much as on CA-EHN, then the benchmark is picking up generic graph-structure effects rather than commonsense specifically.","The paper's coverage of Chinese common words hints that a frequency-matched analogical benchmark for other languages could be derived from equivalent structured lexicons, allowing direct cross-lingual comparison of commonsense embedding quality."],"forward_implications":["CA-EHN can serve as an intrinsic evaluation for Chinese word embeddings that targets commonsense relations rather than morphology or named entities.","Because retrofitting E-HowNet taxonomy improves CA-EHN accuracy by up to 179 percent, the dataset can be used to test whether a representation method actually absorbs ontology-structured commonsense knowledge.","The 763 discovered relations offer a much finer inventory of commonsense relation types than the dozens of predefined relations in existing benchmarks.","Since E-HowNet senses carry English translations, the same analogy questions can be projected to English multi-word expressions, giving a path to cross-lingual commonsense analogy evaluation.","The extraction procedure, if applied to other ontologies with structured definitions, would produce comparable analogy benchmarks in other languages."],"supporting_citations":[{"why":"Introduces the Extended-HowNet representational framework that gives E-HowNet its structured sense definitions.","marker":"(Chen et al., 2005)"},{"why":"Supplies E-HowNet 2.0, the 88K-word lexicon and taxonomy from which all analogies are extracted.","marker":"(Ma and Shih, 2018)"},{"why":"Defines HowNet, the original ontology that E-HowNet extends.","marker":"(Dong and Dong, 2003)"},{"why":"Provides the CA8 Chinese analogy dataset, the morphological/semantic baseline CA-EHN is compared against.","marker":"(Li et al., 2018)"},{"why":"Provides CA-Google, the translated Chinese analogy baseline used in the comparison.","marker":"(Chen and Ma, 2018)"},{"why":"Supplies the GloVe embeddings used to benchmark CA-EHN and the other analogy sets.","marker":"(Pennington et al., 2014)"},{"why":"Supplies the SGNS embedding model and the vector-offset analogy evaluation setup that CA-EHN inherits.","marker":"(Mikolov et al., 2013)"},{"why":"Supplies the retrofitting algorithm used to inject E-HowNet taxonomy knowledge into embeddings.","marker":"(Faruqui et al., 2015)"},{"why":"Supplies ASBC 4.0, the frequency corpus whose five-occurrence threshold filters CA-EHN to common words.","marker":"(Ma et al., 2001)"}],"fun_headline_variants":["90K commonsense analogies from ontology definitions","CA-EHN: 90,505 analogies that probe embedding commonsense","E-HowNet yields 90,505 commonsense analogies","Mining commonsense analogies from E-HowNet ontology","New benchmark: 90,505 analogies for commonsense in embeddings"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"One load-bearing premise is that E-HowNet's structured definitions are accurate and complete enough that two definition graphs differing in exactly one concept node correspond to a human-meaningful commonsense analogy; if the ontology is noisy or the one-node-difference rule admits coincidental pairs, the dataset's relation labels are not valid commonsense relations.","fun_headline_variants_meta":{"raw":{"variants":["90K commonsense analogies from ontology definitions","CA-EHN: 90,505 analogies that probe embedding commonsense","E-HowNet yields 90,505 commonsense analogies","Mining commonsense analogies from E-HowNet ontology","New benchmark: 90,505 analogies for commonsense in embeddings"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000196,"raw_usage":{"total_tokens":1316,"prompt_tokens":859,"completion_tokens":457,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":475,"completion_tokens_details":{"reasoning_tokens":368}},"tokens_in":475,"tokens_out":457,"duration_ms":4228,"temperature":1.0,"reasoning_tokens":368,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T12:21:52.512422+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a random sample of CA-EHN analogies, remove the ontology definitions, and have native speakers judge whether the relation between the two word pairs is a meaningful commonsense analogy; if judged correctness falls far below the paper's reported annotation agreement for concept analogies, the extraction criterion is admitting spurious relations. A second check is to build a control benchmark from random single-node perturbations of the same definition graphs and see whether retrofitting improves accuracy on the control as much as it does on CA-EHN.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces the Extended-HowNet representational framework that gives E-HowNet its structured sense definitions."},{"cited_title":"and Dong, Q","cited_arxiv_id":null,"evidence_quote":"Defines HowNet, the original ontology that E-HowNet extends."},{"cited_title":"and Ma, W.-Y","cited_arxiv_id":null,"evidence_quote":"Provides CA-Google, the translated Chinese analogy baseline used in the comparison."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the GloVe embeddings used to benchmark CA-EHN and the other analogy sets."},{"cited_title":"K., Dyer, C., Hovy, E., and Smith, N","cited_arxiv_id":null,"evidence_quote":"Supplies the retrofitting algorithm used to inject E-HowNet taxonomy knowledge into embeddings."}],"review_version":1}