{"id":"9cdd7796-6a82-4a6b-8e31-4d0ca4367d6c","arxiv_id":"2608.12776","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"ViTOED provides 21,244 Vietnamese opinion quadruples for target-oriented emotion detection; the best baseline models still score only about 34 percent F1 on the full sentiment graph.","lead":"ViTOED is a new annotated Vietnamese dataset for target-oriented emotion detection, containing 21,244 quadruples that specify the source, target, emotional expression, and polarity in social media comments. The authors report that current language models reach only about 30 to 34 percent F1 on the full sentiment graph, showing the task remains challenging.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Annotation reliability is the load-bearing premise: reported Source F1=0.32 and Polarity=0.47, with no final-corpus IAA, do not support the 'strict guidelines' quadruple claim.","rationale":"ViTOED's value rests on the reliability of 21,244 quadruples; if source spans are noisy, the dataset does not support target-oriented emotion detection. The paper's strongest internal support for reliability is Table III, but it reports final training-round agreement on only 210 sentences, and Source=0.32 and Polarity=0.47 are far below what is normally called 'strict'. The contradictory label in Table I Example 2 is independent evidence that the guidelines were not consistently applied. The Table II versus Table IV count inconsistencies and the 52.10% missing-source sentence further reduce confidence in the reported statistics. None of these prove the dataset is unusable: cross-checking and adjudication may have repaired many errors, and the contradiction may be a typo. But as written, the central quality claim is not established. This matches the reader's weakest assumption, so agreement is 'agree'. The verdict stays CONDITIONAL rather than REJECT because the artifact may be salvageable if the authors provide data, fix the statistics, correct the example, and supply either final IAA or a re-annotation audit.","tokens_in":9565,"tokens_out":8128,"duration_ms":83028,"concrete_test":"Download the released ViTOED files from the GitHub link; select a random 300-comment subset from the test split, have two annotators who have not seen the gold labels re-annotate using the Section III.B guidelines, and compute pairwise span F1 and Polarity agreement with Equations 1-2. If Source F1 is near the reported 0.32 or Polarity near 0.47, low agreement is systematic and the quadruple-level benchmark claim is not supported; if agreement is high (e.g., Source >= 0.6, Polarity >= 0.7), the reported numbers are an artifact of the training rounds and the concern recedes.","verdict_should_be":"UNCHANGED","load_bearing_attack":"To support the central claim that ViTOED contains reliable opinion quadruples, the annotations must be trustworthy. Section III.C reports pairwise F1 after four training rounds: Source 0.32, Target 0.61, Expression 0.77, Polarity 0.47 (Table III). These numbers are measured on a 210-sentence sample only; no agreement is reported for the final corpus, and the assertion that 'the annotators understood the labeling guidelines' does not follow from Source/Polarity agreement at these levels. If Source spans are unreliable, the quadruple structure, and every model evaluation built on it, is compromised. This is not merely a low-number worry: Table I Example 2 labels 'These 7 guys are not human' as Positive, directly contradicting Section III.B's rule that rudeness or sensitive content is Negative. The claim of 'strict guidelines' therefore has an internal counterexample. The Table II versus Table IV source and target counts also disagree, and the '52.10% missing Source' sentence counts only the T-E subset, further suggesting the published statistics do not yet cohere.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces ViTOED, a new Vietnamese dataset for target-oriented emotion detection. It claims to contain 10,985 user comments and 21,244 manually annotated opinion quadruples, each consisting of source, target, expression, and polarity. The authors describe an annotation process with guidelines and four rounds of annotator training, analyze dataset statistics including opinion categories and word frequencies, and evaluate a structured sentiment graph baseline (following Barnes et al., 2021) with several Vietnamese pre-trained language models. The main contributions are stated as the dataset itself, an analysis of Vietnamese-specific phenomena such as implicit sources and targets, and baseline results that reveal challenges in span detection and relation extraction.","tokens_in":9670,"tokens_out":5195,"duration_ms":47574,"significance":"If the dataset is reliable, it would fill a real gap in Vietnamese target-oriented emotion detection resources and provide a benchmark that exposes how poorly current models connect emotions to specific targets in Vietnamese social media. The annotation of opinion quadruples is a new contribution, and the baseline experiments with multiple Vietnamese PLMs are a useful starting point. However, the central value of the paper depends entirely on the trustworthiness of the manual annotation, and the current evidence for that trustworthiness is weak.","major_comments":[{"comment":"The paper's central claim of a dataset with 'strict guidelines' and reliable opinion quadruples is not supported by the reported inter-annotator agreement. After four training rounds, Source F1 is 0.32 and Polarity agreement is 0.47; these are the final reported values, measured only on a 210-sentence sample. No IAA is reported for the final corpus, and the adjudication procedure is only described as 'cross-checking among annotators' without detail. The statement that 'the level of agreement indicates that the annotators understood the labeling guidelines and followed the instructions accurately' directly contradicts the low Source and Polarity figures. Please report IAA on the full corpus, describe how disagreements were resolved, and either justify why Source F1 of 0.32 is sufficient or revise the claims about annotation quality.","section":"III.C, Table III"},{"comment":"The annotation guidelines in Section III.B state that expressions containing profanity, rudeness, and sensitive aspects are classified as Negative, and that humor is Positive only if it does not contain hateful or offensive content. Example 2 in Table I, '7 thằng này không phải con người chúng mày ạ' ('These 7 guys are not human'), is labeled Positive even though calling people 'not human' is an offensive statement under any ordinary reading. This is an internal contradiction within the paper's own example. Either the example's label is wrong, or the guideline needs an explicit exception; as published, it undermines the claim of consistent annotation.","section":"III.B, Table I, Example 2"},{"comment":"The dataset statistics are internally inconsistent. Section III.A reports 6,000 sentences from UIT-VSMEC plus 5,010 collected comments, totaling 11,010, but Sections III.D and the abstract state 10,985 comments. Summing Table II gives source counts 2,359+385+572=3,316, while the source-bearing categories in Table IV (S-T-E + S-E) total 2,243+1,527=3,770; the corresponding target counts are 8,632+1,160+2,127=11,919 versus 2,243+11,069=13,312. These numbers must agree because they count the same opinion quadruples. Additionally, the sentence 'The T-E demonstrated that an opinion missing the Source accounts for 52.10% of the total opinions' is misleading: T-E is only one of the two categories without a Source (the other is E), and the true proportion of opinions missing Source is (11,069+6,405)/21,244 = 82.25%. Please reconcile all counts and correct the interpretation.","section":"III.A/III.D, Tables II and IV"}],"minor_comments":[{"comment":"There is a typo in 'sentencesin human language'; also, the notation 'opinion tuples O={o1, o2, ..., on}' should say 'quadruples' for consistency with the rest of the paper.","section":"I"},{"comment":"The phrase 'DSU nis' appears to be a typo or incorrect rendering of the dataset name; please check the official name of the dataset introduced in reference [6].","section":"II"},{"comment":"The label 'BERTologies' should likely be 'BERT' or 'BERT embeddings'.","section":"Figure 2"},{"comment":"The text cites 'Examples #3 and #4 in Table I' as cases lacking a source, but Table I contains only three examples and Example #3 has an explicit source ('tao'). Please correct the reference or extend the table.","section":"III.C"},{"comment":"The notation in Equation (2) uses the same symbols A and B for annotators and for token sets, which is confusing. Please rewrite with distinct notation, e.g., annotators 1 and 2 and sets X and Y.","section":"III.C, Eq. (2)"},{"comment":"The table title 'Results of graph annotation by opinion components by F1(%)' is unclear, and the column layout mixes entity types with polarity classes; please clarify what each column and row represents.","section":"V, Table VII"}],"recommendation":"major_revision","confidential_remarks":"The paper addresses a useful resource gap and the baseline experiments are reasonable, but the load-bearing issues are the low reported IAA with no full-corpus agreement, an internal contradiction in the annotation guidelines/example, and inconsistent dataset statistics. These are not mere presentation problems; they affect the validity of the dataset as a benchmark. I would encourage the authors to provide full IAA and adjudication details, correct the counts and the misleading 52.10% claim, and resolve the Table I Example 2 inconsistency. If those cannot be adequately addressed, the dataset's central claim would be substantially weakened."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"ViTOED is a genuinely new resource — the first Vietnamese dataset with (source, target, expression, polarity) quadruples — and it should find users once the data is actually released. The annotation effort is real: 21,244 quadruples over 10,985 comments, with four training rounds and steady improvement in agreement. The baseline follows Barnes et al.'s structured sentiment graph and gives a reasonable first picture of how Vietnamese PLMs handle target-oriented emotion detection. That part is fine. The citation pattern is fine too; the authors build on relevant prior work and are not hiding where the task definition comes from.\n\nThe soft spots are mainly around reliability claims and internal consistency. The reported IAA after round four is Source 0.32 and Polarity 0.47. Those are low, and the sentence that the annotators \"understood the labeling guidelines and followed the instructions accurately\" does not follow from those numbers. There is no agreement measure for the final corpus, only the 210-sentence sample. If source spans are unreliable, the quadruple structure — the whole point of the dataset — is shaky. Also, Table II and Table IV give inconsistent source and target totals (3,316 vs. 3,770 for source-bearing categories; 11,919 vs. 13,312 for targets). And Table I's second example labels \"These 7 guys are not human\" as Positive while the guidelines say profanity and sensitive content are Negative. That is a direct counterexample to the \"strict guidelines\" claim.\n\nThe T-E 52.10% statement is actually fine if read as the share of opinions in the T-E category; it is the most common type. But the surrounding text could be read as claiming that most opinions have no source, which is true, but it is the T-E category alone.\n\nThe GitHub link is unconfirmed and no snapshot or license is given, so nothing is verifiable from the preprint. Baseline scores are reported without variance or significance tests, which is minor for a dataset paper but worth noting.\n\nOverall: this is a serious dataset effort with a real gap to fill, but the paper oversells annotation reliability and contains fixable inconsistencies. For peer review, I'd send it out — with the condition that the authors release the data, correct the counts, and either present the IAA with more nuance or provide evidence that the final corpus is more reliable than the 210-sample round suggests.\n\nA reader who wants a Vietnamese target-oriented emotion detection benchmark should keep this on the radar. I would not cite it until the data is available and the issues are resolved.","headline":"Genuinely needed Vietnamese target-oriented emotion dataset, but the annotation reliability and internal consistency need serious work before it can be used.","tokens_in":10274,"tokens_out":3139,"would_cite":false,"duration_ms":28474,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper introduces ViTOED, a manually annotated Vietnamese dataset of 10,985 social media comments with 21,244 opinion quadruples, and shows that current language models struggle to attribute emotions to their targets.","keywords":["target-oriented emotion detection","Vietnamese social media","opinion quadruple annotation","structured sentiment graph","span detection","sentiment analysis","low-resource NLP"],"falsifier":"Have fresh annotators independently re-annotate a random sample of the test set using only the published guidelines, then compute pairwise F1 against the released source spans; if the agreement is no better than the original 0.32, or if training a model with source spans removed matches the full model's Targeted F1, then the source component is not carrying real signal and the quadruple claim is not supported.","tokens_in":9273,"feed_emoji":"🎯","tokens_out":7810,"duration_ms":74705,"temperature":0.7,"pith_summary":"ViTOED is a manually annotated Vietnamese dataset built from 10,985 social media comments and containing 21,244 opinion quadruples of the form (source, target, expression, polarity). The paper's central claim is that this resource captures how emotions attach to specific targets in Vietnamese, including phenomena sentence-level sentiment labels miss: over half of the opinions have no explicit source, and the same Vietnamese pronouns can serve as source, target, or both. It then evaluates a structured sentiment graph baseline on the dataset and reports that current language models extract expression spans fairly well but misalign source-target relations and confuse the two roles, leaving a wide performance gap. A sympathetic reader would care because this gives Vietnamese target-oriented emotion detection a reusable benchmark and a concrete list of error types that future models must address.","feed_headline":"New Vietnamese corpus tracks who feels what toward whom","feed_subtitle":"10,985 comments are tagged with 21,244 four-part records so models must name who feels what about whom.","key_machinery":"The load-bearing representation is the opinion quadruple (source, target, expression, polarity), where each component is annotated as a token span inside a Vietnamese comment. The baseline mechanism is the structured sentiment graph from [4], which converts each sentence into a dependency graph whose nodes are spans and whose edges carry source, target, and expression relations with polarity, using either head-first or head-final root choices. The graph is produced by a BiLSTM over word, POS-tag, character, and transformer-contextualized embeddings, scored with a bilinear transformation. This machinery turns emotion detection into a joint span-extraction and relation-extraction problem, and its failure modes become the paper's error taxonomy.","core_discovery":"ViTOED is a human-annotated Vietnamese dataset of 10,985 user comments collected from social networks, containing 21,244 opinion quadruples of the form (source, target, expression, polarity), with source, target, and expression recorded as token spans and polarity as positive, negative, or neutral. The paper claims that this resource captures Vietnamese-specific annotation phenomena that sentence-level sentiment labels miss: 52.10% of opinions have an implicit source (category T-E only), targets are mentioned less often than opinions, and common pronouns such as 'tao' and 'mày' appear on both sides of the quadruple, creating vocabulary ambiguity. Using a structured sentiment graph baseline from [4] over nine pretrained language models, the paper reports that current models succeed at expression span detection but struggle with source and target spans, graph root placement, and relation alignment. The error analysis attributes these failures to three causes: vocabulary complexity in span recognition, root and edge misprediction, and source-target confusion caused by Vietnamese morphology. The intended payoff is a benchmark on which future Vietnamese target-oriented emotion detection systems can be measured, plus an explicit list of error types to attack.","pith_inferences":["Since source-span agreement is only F1 = 0.32, part of the reported performance gap on sources may reflect annotation noise rather than model weakness; a natural next test is to score models only on sentences where annotators agreed on the source span.","The same quadruple annotation protocol could transfer to other subject-drop languages such as Japanese, Korean, or Thai, where implicit sources predominate, enabling cross-lingual comparison if those datasets adopt the same graph parser.","A simple and testable extension suggested by the data is to add a Vietnamese first-person and second-person pronoun feature to the graph parser, which might directly reduce the source-target confusion the paper identifies.","The paper reports per-component agreement but not exact quadruple-level agreement; computing the fraction of fully matching (source, target, expression, polarity) tuples would give a stricter, and possibly much lower, reliability estimate."],"forward_implications":["Vietnamese target-oriented emotion detection now has a public train/dev/test benchmark with 21,244 quadruples and fixed splits.","Because 52.10% of opinions in ViTOED lack an explicit source, any successful Vietnamese model must learn to infer the commenter as an implicit source rather than only detecting explicit mentions.","The structured sentiment graph baseline shows that both multilingual and monolingual Vietnamese language models leave substantial headroom on this task, particularly on source and target spans.","The head-final graph representation and the +inlabel extension offer concrete starting points for improving expression and span extraction on this dataset.","The paper's error taxonomy gives future work three defined targets: span misalignment, root and edge misprediction, and source-target confusion in Vietnamese morphology."],"supporting_citations":[{"why":"Supplies 6,000 of the 10,985 sentences, the existing Vietnamese emotion-annotated data that ViTOED extends.","marker":"[3]"},{"why":"Provides the structured sentiment graph parser, graph representations, and the four evaluation metrics used as the baseline.","marker":"[4]"},{"why":"Supplies the span-based F1 inter-annotator agreement method and the token-level aspect span detection setting this dataset builds on.","marker":"[11]"},{"why":"One of the multilingual pretrained encoders evaluated as a baseline in the experiments.","marker":"[12]"},{"why":"Provides the monolingual Vietnamese pretrained encoder used as a baseline language model.","marker":"[14]"},{"why":"A Vietnamese social-media pretrained model whose high targeted F1 supports the paper's monolingual advantage observation.","marker":"[15]"},{"why":"The encoder-decoder model that obtains the best parsing and sentiment graph results and anchors the error analysis.","marker":"[17]"}],"fun_headline_variants":["ViTOED: 21,244 quadruples map Vietnamese emotion targets","Models struggle with implicit sources in ViTOED sentiment set","10,985 Vietnamese comments yield 21K opinion quadruples","ViTOED benchmark: span and relation gaps limit Vietnamese sentiment"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The dataset's usefulness rests on the assumption that the annotators' opinion quadruples are reliable enough to learn from; in particular, if the low source-span agreement (F1 = 0.32) means source annotations are effectively arbitrary, the quadruple structure that defines ViTOED loses its validity.","fun_headline_variants_meta":{"raw":{"variants":["ViTOED: 21,244 quadruples map Vietnamese emotion targets","Models struggle with implicit sources in ViTOED sentiment set","10,985 Vietnamese comments yield 21K opinion quadruples","ViTOED benchmark: span and relation gaps limit Vietnamese sentiment"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000941,"raw_usage":{"total_tokens":3990,"prompt_tokens":882,"completion_tokens":3108,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":498,"completion_tokens_details":{"reasoning_tokens":3036}},"tokens_in":498,"tokens_out":3108,"duration_ms":21136,"temperature":1.0,"reasoning_tokens":3036,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T23:33:18.528813+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have fresh annotators independently re-annotate a random sample of the test set using only the published guidelines, then compute pairwise F1 against the released source spans; if the agreement is no better than the original 0.32, or if training a model with source spans removed matches the full model's Targeted F1, then the source component is not carrying real signal and the quadruple claim is not supported.","supporting_citations":[{"cited_title":"Emotion recognition for vietnamese social media text,","cited_arxiv_id":null,"evidence_quote":"Supplies 6,000 of the 10,985 sentences, the existing Vietnamese emotion-annotated data that ViTOED extends."},{"cited_title":"Structured sentiment analysis as dependency graph parsing,","cited_arxiv_id":null,"evidence_quote":"Provides the structured sentiment graph parser, graph representations, and the four evaluation metrics used as the baseline."},{"cited_title":"Span detection for aspect-based sentiment analysis in Vietnamese,","cited_arxiv_id":null,"evidence_quote":"Supplies the span-based F1 inter-annotator agreement method and the token-level aspect span detection setting this dataset builds on."},{"cited_title":"BERT: Pre- training of deep bidirectional transformers for language understanding,","cited_arxiv_id":null,"evidence_quote":"One of the multilingual pretrained encoders evaluated as a baseline in the experiments."},{"cited_title":"PhoBERT: Pre-trained language models for Vietnamese,","cited_arxiv_id":null,"evidence_quote":"Provides the monolingual Vietnamese pretrained encoder used as a baseline language model."},{"cited_title":"ViSoBERT: A pre- trained language model for Vietnamese social media text processing,","cited_arxiv_id":null,"evidence_quote":"A Vietnamese social-media pretrained model whose high targeted F1 supports the paper's monolingual advantage observation."},{"cited_title":"ViT5: Pretrained text-to- text transformer for Vietnamese language generation,","cited_arxiv_id":null,"evidence_quote":"The encoder-decoder model that obtains the best parsing and sentiment graph results and anchors the error analysis."}],"review_version":1}