{"id":"f0306525-7e25-4a3a-adb5-c0efa7d0f7ab","arxiv_id":"2412.00856","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"The authors propose revisions to Korean Universal Dependencies guidelines, including last-noun heads and new relations, but provide no empirical validation.","lead":"This paper proposes a revised set of Korean Universal Dependencies annotation guidelines, changing the head of noun compounds to the last noun and adding new dependency relations such as obl:arg. A smart generalist might read it to understand how language-specific grammar annotation is argued and revised, though the paper offers no quantitative evaluation of the proposed changes.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The last-noun-head rule is asserted without evidence and its case-marker justification is circular; the compound/flat revisions built on it are therefore not validated.","rationale":"The paper is a guideline-revision proposal, and its value depends on whether the proposed relations are linguistically sound. The reader's conditional verdict already flags the unargued head-last premise, and my read agrees that this is the most load-bearing point. The case-marker rationale in Section 3.1 is not merely missing a citation; it mistakes a surface property of the whole NP for evidence about the internal head, which makes the argument circular. Since the flat and compound relations, and the associated argument relations, are revised on this basis, the entire proposal is vulnerable if the rule overgeneralizes. The absence of inter-annotator agreement statistics and the lack of any comparison with the existing scheme further prevent the reader from independently confirming the claim. These problems justify keeping the verdict conditional: the paper is a coherent proposal with concrete examples, but it has not demonstrated that its central rule is correct. A focused test on the Sejong-derived sentences or dictionary data would either validate the rule or expose the need for exceptions, and it is concrete enough to run without waiting for the authors to release the full 200-sentence annotation.","tokens_in":7859,"tokens_out":5442,"duration_ms":55526,"concrete_test":"Extract all noun-compound and multiword proper-name occurrences from the 200 Sejong sentences used in Section 4.1, and have two independent Korean-speaking annotators label the semantic head using a definitional hyponymy test (e.g., the compound denotes a subtype of its head noun). Compare these labels against the proposed last-noun rule. If the rule disagrees with the annotators on any attested case where the surface-final token receives the case but is not the semantic head, or on a non-negligible fraction of endocentric items, the categorical head-last rule is not supported and needs an exception mechanism.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that the revised guidelines better capture Korean syntax depends on the categorical head-last rule for noun compounds and multiword proper names introduced in Section 3.1. The only explicit justification for proper names is that 'the last proper noun would receive the case from the predicate.' This does not support the rule: Korean case markers attach to the surface-final element of the entire noun phrase, not necessarily to its semantic head, so the observation that the last token bears the case is a consequence of the annotation decision, not an independent fact about it. The argument is therefore circular. The paper also begins with 'usually the last noun' but then applies the rule unconditionally, and it offers no citation or discussion of left-headed, exocentric, or headless compounds. Because the revised compound and flat annotations, and the argument-structure relations assigned to their dependents, are built directly on this rule, an unsupported head rule undermines the paper's claim of improvement. The 200-sentence validation mentioned in Section 4.1 is not analyzed, so it cannot supply the missing evidence.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a set of revisions to the Korean Universal Dependencies (UD) guidelines, motivated by perceived inadequacies in the current Korean UD treebanks. The core proposals include annotating the last noun rather than the first noun as the head of noun compounds and multiword proper names, introducing the relation obl:arg for oblique arguments determined by Sejong dictionary subcategorization frames, adding dislocated:nsubj for topicalized nominals, and adding csubj:pass for clausal passive subjects. The paper also discusses annotation conventions for auxiliary verbs, copulas, and markers. To validate the revisions, the authors describe manually annotating 200 sentences from the Sejong project, but they do not report any quantitative results, agreement metrics, or comparisons with existing annotations. The paper concludes with plans to integrate the revised guidelines into UD documentation.","tokens_in":7996,"tokens_out":3703,"duration_ms":34695,"significance":"If the proposed revisions are linguistically sound and empirically validated, they could lead to more accurate Korean UD treebanks and better alignment with Sejong-style dependency annotation, potentially improving Korean dependency parsing and cross-linguistic comparability. The paper correctly identifies genuine weaknesses in the current Korean UD annotation, such as the treatment of noun-phrase heads and auxiliary verb constructions, and it offers concrete proposals for new relations (obl:arg, dislocated:nsubj, csubj:pass) that may be useful to the community. However, the paper offers no empirical evidence that the revised guidelines 'better capture Korean syntax' as claimed, and the manual annotation described in Section 4.1 is not analyzed or released. The significance is therefore conditional: the proposal is interesting, but the central claim of improvement is unsupported in the present version.","major_comments":[{"comment":"The head-last rule for noun compounds and multiword proper names is asserted without supporting evidence or discussion of competing analyses. The paper states that 'the head is typically the stem that determines the semantic category, usually the last noun in Korean,' but it then applies the rule unconditionally, without addressing left-headed, exocentric, or headless compounds. The justification for proper names—'the last proper noun would receive the case from the predicate'—is circular, because Korean case markers attach to the surface-final element of the whole noun phrase, which is a consequence of the annotation decision rather than an independent fact about headedness. This rule is load-bearing: the proposed compound and flat annotations, and the argument relations assigned to their dependents, all depend on it. The paper should either provide principled evidence for categorical head-last status in Korean compounds or restrict the rule to cases where it can be independently justified.","section":"Section 3.1"},{"comment":"The 200-sentence manual annotation study is described but never evaluated. The paper reports that the sentences were 'randomly selected from the Sejong project' and annotated according to the revised guidelines, but it gives no results: there are no inter-annotator agreement scores, no comparison with existing UD annotations or Sejong-style annotations, no error analysis, and no release of the annotated data. Consequently, the central claim that the revised guidelines 'better capture Korean syntax' is not supported by the evidence presented. To support the claim, the authors need to report quantitative measures of annotation validity and reliability, or explicitly reframe the contribution as a preliminary guideline proposal without an empirical validation claim.","section":"Section 4.1"},{"comment":"The use of Sejong dictionary subcategorization frames as the authority for distinguishing obl:arg from obl is not sufficiently specified. The paper relies on frames such as 'X=N0-i Y=N1-e|ege johda' to decide whether a postpositional phrase is an argument, but it does not describe the source edition of the Sejong dictionary, the selection criteria for frames, how frames are matched to predicates in actual sentences, or how cases not covered by the dictionary are handled. Without these details, the proposed obl:arg and dislocated:nsubj annotations are not reproducible, and the claim that the revised guidelines improve upon the current 'annotate everything as obl' practice cannot be assessed.","section":"Section 3.3"}],"minor_comments":[{"comment":"The example '싸움 ’ 이라는' includes a cop annotation, but it is not clear how the copular marker '-i' is tokenized when it is separated from the preceding lexeme; please clarify the tokenization rule.","section":"Section 3.4"},{"comment":"Figure 1 is difficult to read: the two annotation layers ('top' and 'bottom') are indicated solely by the position of the arrows, and the relations are not clearly distinguished. A side-by-side or color-coded layout would improve clarity.","section":"Figure 1"},{"comment":"The term 'catenative construction' is used without definition or a reference; please define it and explain how it differs from auxiliary and complement constructions in Korean, since the distinction is central to the proposed xcomp and aux annotations.","section":"Section 3.2"},{"comment":"The paper refers to the 'Sejong dictionary' and the 'Sejong project' but does not cite a specific edition or corpus; please provide a reference to the exact resource used for the subcategorization frames and for the 200-sentence sample.","section":"Section 4.1"},{"comment":"The description of the random selection of 200 sentences is underspecified; please state the source corpus, the sampling procedure, and whether the selections were stratified or simply a convenience sample.","section":"Section 4.1"},{"comment":"The claim that Korean is an 'end-focus language' and that the Sejong-style dependency structure 'consistently adheres to a right-to-left pattern' is asserted without quantitative support or a citation; please provide evidence or temper the claim.","section":"Section 4.2"},{"comment":"The conclusion mentions aligning with the KLUE benchmark consortium and the National Institute of Korean Language, but no details or letters of support are provided; either elaborate on these collaborations or omit them as unverifiable.","section":"Conclusion"}],"recommendation":"major_revision","confidential_remarks":"The paper is better characterized as a position paper or a proposal for guideline changes rather than a validated research contribution. The lack of any empirical analysis in Section 4.1 is a serious gap, and the head-last rule in Section 3.1 needs a firmer linguistic foundation. I would encourage the authors to either add a rigorous evaluation (including inter-annotator agreement and a comparison with existing UD annotations) or to explicitly present this as a proposal for community discussion rather than claiming that the revised guidelines 'better capture Korean syntax.' As it stands, the evidence is not sufficient for publication at a serious journal, but the proposal itself is plausible and worth developing."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a useful proposal for revising Korean UD guidelines, with a concrete set of changes and worked examples, but the central claim that the revisions are better is not backed by analysis. The head-last rule is asserted, not argued. Worth sending to reviewers, but only if they demand real evaluation.\n\nWhat is actually new: the paper compiles a specific set of revisions - last-noun heads for compounds and flat names, obl:arg, dislocated:nsubj, and csubj:pass - that go beyond the earlier Noh et al. (2018) critique. The authors ground argument-structure decisions in Sejong dictionary frames, which gives the proposal an external anchor. The examples are clear, and the discussion of auxiliary verbs versus catenative constructions is a legitimate issue in Korean UD.\n\nThe soft spot is the load-bearing justification. Section 3.1 says the head of a noun compound is \"usually the last noun in Korean,\" then applies the rule unconditionally, with no citation to Korean linguistics and no discussion of left-headed or exocentric compounds. The stated reason for last proper noun heads - \"the last proper noun would receive the case from the predicate\" - does not work as justification. Case markers attach to the surface-final element of the noun phrase, not necessarily to the semantic head, so that observation is a consequence of the annotation decision, not an independent fact supporting it. The stress-test note has this right. The flat and compound revisions, and the argument relations assigned to their dependents, rest on this unargued premise.\n\nEqually important, the 200-sentence annotation described in Section 4.1 is never analyzed. No inter-annotator agreement, no comparison with the current scheme, no released data. That means the paper's central claim - that the revised guidelines better capture Korean syntax - is plausible but unverified. I do not see this as a fatal flaw for a position paper, but it should be stated as a proposal, not a validated revision.\n\nWho is this for? People working on Korean dependency parsing and UD treebank design. It deserves a serious referee because it raises concrete questions about the current annotation and offers a specific alternative. But the reviewers should push for stronger linguistic grounding and empirical validation before any adoption.\n\nMy recommendation: accept for peer review with the expectation of major revision. The idea is worth engaging, but the evidence as presented is thin.","headline":"A concrete, clearly written proposal for revising Korean UD guidelines, but the head-last rule is asserted rather than argued and the validation is missing, so the revisions are not yet proven better.","tokens_in":8599,"tokens_out":1537,"would_cite":false,"duration_ms":15531,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that Korean Universal Dependencies guidelines should be revised so that the last noun heads noun compounds and proper-name chains, that oblique complements are labeled obl:arg via Sejong frame information, and that…","keywords":["Korean Universal Dependencies","dependency treebank annotation","noun compound head","obl:arg","subcategorization frame","Sejong treebank","catenative construction","clausal passive subject"],"falsifier":"Annotate every noun compound in the Korean_GSD treebank for whether the last noun determines the semantic category; if even a handful of compounds are left-headed or exocentric (the category matches the first noun or neither noun alone), then the blanket last-noun rule is falsified and must be relaxed into per-construction or per-item decisions.","tokens_in":7605,"feed_emoji":"🗂️","tokens_out":6906,"duration_ms":55543,"temperature":0.7,"pith_summary":"This paper proposes a set of revisions to the Korean Universal Dependencies (UD) annotation guidelines, targeting syntactic relations rather than POS tags. It argues that the current Korean treebanks systematically misannotate four construction types: noun compounds and multiword proper names (where the first noun is wrongly treated as head), auxiliary-versus-catenative verb constructions (where the first verb is always head), postpositional arguments (all labeled obl), and topic-marked subjects (labeled dislocated rather than nsubj). The proposed changes make the last noun the head, introduce obl:arg for arguments supported by Sejong dictionary subcategorization frames, add dislocated:nsubj and csubj:pass, and distinguish aux from xcomp. If adopted, these guidelines would require re-annotating the existing Korean GSD, Kaist, and Penn treebanks and would change the data used to train and evaluate Korean dependency parsers. The paper manually annotates 200 Sejong sentences as a first validation of the revised scheme.","feed_headline":"Last noun, not first, should head Korean compounds and names","feed_subtitle":"Revising the annotation scheme would reshape Korean dependency treebanks and the parsers trained on them.","key_machinery":"The load-bearing machinery is a head-selection rule plus a frame-based argument test. The head-selection rule states that the semantic head of a noun compound is the stem that determines the semantic category, usually the last noun in Korean, and that the last proper noun is chosen as the head of flat structures because it receives case from the predicate. The argument test consults the Sejong dictionary's subcategorization frames: if a predicate's frame lists a specific postposition phrase as an argument, the dependency is obl:arg; otherwise it is obl. These two devices jointly drive the new relations obl:arg, dislocated:nsubj, and csubj:pass, and distinguish aux from xcomp in catenative constructions.","core_discovery":"The central claim is that current Korean UD annotations follow a first-word convention that is wrong for Korean: in noun compounds such as jangdonggeon sajin-eul ('Jang Dong-gun picture'), the head is the last noun sajin ('picture'), which receives the accusative case from the predicate; in catenative constructions like bogo sipda ('see want'), the lexical head is the catenative verb sipda ('want'), not the preceding verb; adverbial postposition phrases that are complements of the predicate, e.g., jutaegsijang-e in jutaegsijang-e johda ('good for the housing market'), should be obl:arg rather than plain obl; and topicalized nominals such as kokkili-neun ('elephant.top') in double-subject sentences are subjects rather than generic dislocated elements, hence dislocated:nsubj. The revision also adds csubj:pass for clausal passive subjects, exemplified by ttwieonan dunoeim-i jeungmyeongdwaessda ('being an outstanding brain was proven'). Under the revised scheme, the head of an MWE or compound is the last word for flat and compound relations, and Sejong dictionary subcategorization frames are the authority for deciding which postpositional phrases count as arguments.","pith_inferences":["Because the same last-word-headedness logic is used for flat and compound relations, the proposal implicitly predicts that right-headedness holds across all Korean MWEs, including dates and names; a left-headed or exocentric counterexample would force a more per-construction rule.","The Sejong-frame test could be operationalized automatically: if frame information is available for all predicates in a treebank, the obl/obl:arg decision becomes lexically driven rather than annotation-dependent, which would improve annotation consistency but also make treebank quality sensitive to frame dictionary coverage.","If the revised guidelines were applied to Kaist and Penn as well, the three Korean UD treebanks would diverge less from the Sejong/KLUE dependency structure, easing the stated goal of a consensus model; this is a direction the paper says it is working toward but has not yet completed.","A testable extension would be inter-annotator agreement on a random sample of noun compounds; if agreement on the last-noun head is high, the rule is robust, and if not, the guideline needs exception classes."],"forward_implications":["If the last-noun head rule is correct, the existing Korean_GSD treebank misannotates every compound noun and multiword proper name whose head is not the first word, requiring bulk re-annotation.","Adopting obl:arg would change the counts of core versus non-core arguments in Korean treebanks, affecting parser evaluation and cross-lingual comparisons.","Using Sejong dictionary frames as the authority for argumenthood ties the UD annotations to a specific Korean lexical resource, so the revised treebank inherits that resource's coverage and decisions.","Adding csubj:pass and dislocated:nsubj to the relation inventory would make Korean treebanks more descriptively adequate for double-subject and passive clausal constructions.","The 200-sentence manual annotation demonstrates that the revised guidelines are applicable to unseen Sejong sentences, supporting the feasibility of full treebank conversion."],"supporting_citations":[{"why":"Defines the Universal Dependencies annotation framework and relation inventory that the paper revises.","marker":"(Nivre et al., 2016)"},{"why":"Documents UD v2 treebank collection and guidelines, the current standard being critiqued.","marker":"(Nivre et al., 2020)"},{"why":"Source of the Korean GSD treebank whose first-noun head annotation the paper argues against.","marker":"(McDonald et al., 2013)"},{"why":"Origin of the Kaist treebank, converted to dependency structure in Korean UD and covered by the proposed revisions.","marker":"(Choi et al., 1994)"},{"why":"Prior critique of Korean Universal Dependencies that motivates the revision effort.","marker":"(Noh et al., 2018)"},{"why":"Specifies the universal Stanford dependency relations (including obl and dislocated) that the revised scheme extends.","marker":"(de Marneffe et al., 2014)"},{"why":"KLUE benchmark follows Sejong-style dependency structure, the alternative the paper wants to align with.","marker":"(Park et al., 2021)"},{"why":"Established the Sejong constituency-to-dependency conversion, demonstrating right-to-left dependency patterns in Korean.","marker":"(Oh and Cha, 2013)"},{"why":"Reports morphological annotation error rates in Korean GSD, additional evidence that the treebank needs revision.","marker":"(Jo et al., 2023)"}],"fun_headline_variants":["Korean compound heads shift to last word in revamped UD","New Korean UD rules: last noun heads compounds and names","Flip the head: Korean UD prefers final word in phrases","Korean dependency fix: head right, not left, in compounds","Revising Korean UD: last word governs in compounds and names"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The revision rests on the premise that the last noun in a Korean noun compound is always the semantic head and that the Sejong dictionary's subcategorization frames are the right authority for whether a postpositional phrase is an argument; if either gives way, the new relations would misannotate rather than improve the treebank.","fun_headline_variants_meta":{"raw":{"variants":["Korean compound heads shift to last word in revamped UD","New Korean UD rules: last noun heads compounds and names","Flip the head: Korean UD prefers final word in phrases","Korean dependency fix: head right, not left, in compounds","Revising Korean UD: last word governs in compounds and names"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00016,"raw_usage":{"total_tokens":1192,"prompt_tokens":867,"completion_tokens":325,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":483,"completion_tokens_details":{"reasoning_tokens":238}},"tokens_in":483,"tokens_out":325,"duration_ms":3571,"temperature":1.0,"reasoning_tokens":238,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T04:54:47.152441+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Annotate every noun compound in the Korean_GSD treebank for whether the last noun determines the semantic category; if even a handful of compounds are left-headed or exocentric (the category matches the first noun or neither noun alone), then the blanket last-noun rule is falsified and must be relaxed into per-construction or per-item decisions.","supporting_citations":[{"cited_title":"Manning, Ryan McDonald, Slav Petrov, Sampo Pyysalo, Natalia Silveira, Reut Tsarfaty, and Daniel Zeman","cited_arxiv_id":null,"evidence_quote":"Defines the Universal Dependencies annotation framework and relation inventory that the paper revises."},{"cited_title":"Manning, Sampo Pyysalo, Sebastian Schuster, Francis Tyers, and Daniel Zeman","cited_arxiv_id":null,"evidence_quote":"Documents UD v2 treebank collection and guidelines, the current standard being critiqued."},{"cited_title":"a ckstr \\","cited_arxiv_id":null,"evidence_quote":"Source of the Korean GSD treebank whose first-noun head annotation the paper argues against."},{"cited_title":"Han, Young G","cited_arxiv_id":null,"evidence_quote":"Origin of the Kaist treebank, converted to dependency structure in Korean UD and covered by the proposed revisions."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Specifies the universal Stanford dependency relations (including obl and dislocated) that the revised scheme extends."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"KLUE benchmark follows Sejong-style dependency structure, the alternative the paper wants to align with."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Established the Sejong constituency-to-dependency conversion, demonstrating right-to-left dependency patterns in Korean."}],"review_version":1}