{"id":"52193533-d441-40a2-8f32-26318e9a147a","arxiv_id":"2505.13483","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Introduces EmoMeta, a publicly available Chinese multimodal dataset of 5,000 metaphorical advertisements with fine-grained emotion labels across ten categories.","lead":"This paper releases EmoMeta, a dataset of 5,000 Chinese text-image advertisement pairs annotated for metaphors and ten fine-grained emotion categories. It is the first resource of its kind for multimodal metaphor emotion classification in Chinese, aimed at improving emotional AI systems.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Emotion-label reliability is the load-bearing risk: Fleiss κ = 0.58 is moderate, and the Section 3.2 rule of choosing 'the more intense emotion' when image and text conflict can systematically collapse mixed emotions, undermining the fine-grained ground truth.","rationale":"The reader identified emotion-label reliability as the weakest assumption, citing both the moderate Fleiss κ and the image-text conflict resolution rule. My reading of the full manuscript supports this: the paper explicitly reports α = 0.58 for emotion categories, calls it reliable, and describes a protocol that selects the more intense emotion in conflict cases without reporting conflict frequency or validating that this yields accurate fine-grained labels. This is not an external disagreement with consensus; it is an internal concern about whether the dataset's central labels can support the claimed fine-grained emotion classification. The dataset is still a legitimate resource, and the public release and clear annotation process are positive, so the conditional verdict remains appropriate rather than rejection. The proposed test is feasible and would directly distinguish reliable labels from protocol-driven artifacts.","tokens_in":6445,"tokens_out":3525,"duration_ms":35873,"concrete_test":"Release the per-emotion Fleiss κ values and the pairwise annotator confusion matrix for the three annotation groups. Then take a stratified random sample of 100–200 text-image pairs and re-annotate them with three separate labels: emotion conveyed by text, emotion conveyed by image, and final emotion per the paper's protocol. If image-text conflicts occur in more than about 20% of samples, and if final-label agreement on conflict cases is substantially lower than agreement on non-conflict cases, the 'stronger emotion' rule is systematically distorting the fine-grained emotion ground truth.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The dataset's advertised contribution is fine-grained emotion labels, yet those labels are the least secure component of the annotation pipeline. Section 3.4 reports inter-annotator agreement α = 0.58 for emotion categories and calls the annotations 'reliable'; by standard benchmarks this is moderate, not strong, and the paper reports no per-category kappa or pairwise confusion matrix. Section 3.2 states that when the image and text convey conflicting emotions, annotators use 'the more intense emotion from either source' as the final label. That rule assumes conflicts are rare or cleanly resolvable by intensity, but the paper gives no statistics on how often image-text conflicts occur. For a 10-way forced-choice schema with adjacent categories such as joy/love/trust and fear/sadness/disgust/anger, moderate agreement plus forced single-label resolution means the training signal for fine-grained distinctions may reflect annotator convention rather than stable properties of the stimuli. Because the central claim depends on the validity of emotion labels, this is a load-bearing concern, not a presentation issue.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript introduces EmoMeta, a dataset of 5,000 Chinese text-image advertisement pairs annotated for metaphor occurrence, source/target domains, and fine-grained emotion categories (joy, love, trust, fear, sadness, disgust, anger, surprise, anticipation, neutral). The authors describe collection from four sources, an annotation protocol with three groups and majority voting, and report Fleiss kappa values for metaphor identification, source domain, target domain, and emotion. The paper presents descriptive statistics of emotion distributions in public-service versus commercial advertisements and argues the resource fills a gap in Chinese multimodal metaphor emotion research.","tokens_in":6638,"tokens_out":3466,"duration_ms":31697,"significance":"If the emotion labels are reliable, EmoMeta is a useful first resource for Chinese multimodal metaphor emotion research and is the first of its kind in Chinese. The authors release the data publicly, describe manual annotation, and provide separate analyses for advertisement types. The main risk is label reliability, because the reported agreement for the emotion dimension is moderate and the conflict-resolution rule is not validated.","major_comments":[{"comment":"The reported Fleiss kappa of 0.58 for emotion categories is moderate by standard benchmarks (e.g., Landis and Koch), not the 'reliable' level the text claims. The paper provides no per-category kappa values, no pairwise confusion matrix, and no discussion of which emotion pairs are confused. Because the central contribution is fine-grained emotion labels, this evidence is insufficient to establish that the labels are stable ground truth.","section":"Section 3.4"},{"comment":"The rule for resolving image-text emotion conflicts by selecting 'the more intense emotion from either source' is arbitrary and unquantified. The paper gives no statistics on how often image and text convey conflicting emotions, nor any analysis of whether the intensity rule preserves or distorts mixed-emotion cases. This rule can systematically collapse genuine mixed emotions into a single label, undermining the fine-grained classification claim. The authors should report conflict frequency, show examples of resolved conflicts, and evaluate whether the rule introduces systematic bias.","section":"Section 3.2"},{"comment":"The dataset analysis is purely descriptive; no baseline experiments are reported. The claim that EmoMeta will 'facilitating further advancements' is unsupported without at least a simple text-only, image-only, and multimodal fusion baseline to show that the emotion labels are learnable and to calibrate expected performance. Adding such experiments would also provide an external check on label quality beyond inter-annotator agreement.","section":"Section 4"}],"minor_comments":[{"comment":"The subsection heading contains a typo: 'Metaphorcial or literal' should be 'Metaphorical or literal'.","section":"Section 3.3"},{"comment":"The sentence 'Current research on metaphorical emotions in typically utilizes broad emotion classifications' has a grammatical error; 'in' should be removed or the sentence restructured.","section":"Section 2"},{"comment":"The Kappa score is denoted with the symbol α, but the standard symbol for Fleiss' kappa is κ; using κ consistently would avoid confusion with Cronbach's alpha or Krippendorff's alpha.","section":"Section 3.4"},{"comment":"There is a stray apostrophe in \"It 's important to note\" that should be corrected.","section":"Section 3.3"},{"comment":"The phrase 'In cases where all data is metaphorical, we specifically note this condition' is vague; it is unclear what 'all data' refers to and how this condition is recorded in the annotation model.","section":"Section 3.2"},{"comment":"The sentence 'However, this part of the data is not accompanied by a URL in the dataset of the data source' is confusing; please clarify whether the images are redistributed as files and how this complies with the original dataset's terms.","section":"Section 3.1"},{"comment":"Figure 3 is referenced but its axes and units are not described in the text; please specify whether values are counts or percentages and define the emotion labels used in the figure.","section":"Section 4 / Figure 3"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a four-page companion paper with a potentially valuable resource, but the emotion-label reliability evidence is thin for the central claim. I would ask the authors to supply the additional reliability analyses and baseline experiments before publication, as outlined in the major comments."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know this is a resource paper, not a methods paper. The contribution is a public dataset of 5,000 Chinese text-image metaphorical advertisements annotated for metaphor presence, source/target domains, and ten fine-grained emotion categories. That combination is new as far as I can tell, and the release is real. The annotation taxonomy mixes Ekman plus Plutchik with a neutral class, which is sensible for advertising. The paper is short but honest about data sources, including reuse of the authors' own earlier datasets.\n\nWhat is good: the dataset fills a genuine gap. Existing multimodal metaphor datasets are mostly English or coarse-grained; EmoMeta adds fine-grained emotions in Chinese. The annotation process is described with three groups, majority voting, and Fleiss kappa. The distribution analysis of fear/anticipation in public service ads vs. surprise in commercial ads is plausible and useful. The GitHub release means the field can actually use it.\n\nNow the soft spots. The label that carries the whole point is the emotion category, and that is the weakest annotation. κ = 0.58 is moderate by any standard, and the paper calls it “reliable.” It does not report per-category kappa or a confusion matrix, which matters a lot for a 10-way scheme with close neighbors like joy/love/trust and fear/sadness/disgust. Worse, Section 3.2 says image-text conflicts are resolved by taking “the more intense emotion” from either source. That rule is arbitrary and can silently erase genuine mixed emotions. The paper gives no statistics on how often conflicts occur, so we cannot tell how much damage the rule does. There are also no baseline experiments, no task validation, no attempt to show the labels correlate with anything external. For a dataset paper, that is a real omission.\n\nThe 87/13 public/commercial split is worth noting but not fatal; the complementarity argument is reasonable.\n\nWho is this for? Researchers working on Chinese metaphor, multimodal emotion recognition, or advertising NLP. They will want the resource, but they should treat the emotion labels as noisy until the authors supply per-category agreement and a conflict-resolution analysis. The paper deserves a serious referee, but it should be sent back for major revision. I would not desk-reject it, and I would probably cite it if I worked in that niche, but I would not trust the fine-grained labels without independent inspection.\n\nRecommendation: accept for peer review with a request for validation experiments and a more careful treatment of label reliability.","headline":"A genuinely new Chinese multimodal metaphor-emotion dataset, but the emotion labels are the load-bearing part and their reliability (κ=0.58, arbitrary conflict resolution) needs real work before this is trustworthy as fine-grained ground truth.","tokens_in":7141,"tokens_out":1544,"would_cite":true,"duration_ms":15957,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A new Chinese-language dataset of 5,000 metaphorical advertisements, each annotated for metaphor occurrence, source and target domains, and one of ten fine-grained emotion categories, claims to be the first resource enabling fine-grained…","keywords":["multimodal metaphor","fine-grained emotion classification","Chinese dataset","metaphor annotation","text-image pairs","advertising","emotion taxonomy","public service advertisement"],"falsifier":"Re-annotate a random sample of the image–text conflict cases using a protocol that permits multiple emotion labels per pair; if more than a fifth of those pairs receive a label different from the published single label, or if a classifier trained on the published labels performs at chance against the multi-label annotations, the claim that the dataset supports fine-grained ground-truth emotion classification is undercut.","tokens_in":6267,"feed_emoji":"🎭","tokens_out":5599,"duration_ms":50817,"temperature":0.7,"pith_summary":"The paper introduces EmoMeta, a Chinese multimodal dataset of 5,000 text–image advertisement pairs, and claims it is the first resource of its kind: every pair is manually annotated for whether a metaphor occurs, what the source and target domains are, and which of ten fine-grained emotion categories the metaphor conveys. The motivation is that metaphors carry much of the emotional load in advertising, yet existing emotion resources are mostly English, mostly coarse-grained (positive/negative), and usually text-only or image-only. A sympathetic reader takes the central claim to be that this dataset makes fine-grained multimodal metaphorical emotion classification in Chinese possible and that its annotation model can be reused. The paper also reports distributional regularities, such as fear and anticipation dominating public-service advertisements while surprise dominates commercial advertisements.","feed_headline":"New dataset tags 5,000 Chinese metaphor ads with fine-grained emotions","feed_subtitle":"Image-text pairs carry metaphor, source-target domains, and ten emotion labels for Chinese advertising.","key_machinery":"The central object is the annotation model (Occurrence, source domain, target domain, emotion category) applied to multimodal advertisement pairs. The metaphor component relies on the source-to-target domain mapping, identified by judging an irreversible 'A is B' relationship across the text and image; the emotion component uses a ten-way taxonomy built from the six basic emotions plus trust and anticipation and a neutral class. The annotation process itself is the mechanism that carries the dataset: three groups of annotators independently label, disagreements are re-evaluated, and majority vote settles remaining conflicts, with inter-annotator agreement coefficients used as quality checks.","core_discovery":"The paper's central claim is that EmoMeta is the first Chinese multimodal dataset for fine-grained emotion classification in metaphorical advertisements. Each of its 5,000 text–image pairs carries an annotation of the form (Occurrence, source domain, target domain, emotion category), where the emotion categories are joy, love, trust, fear, sadness, disgust, anger, surprise, anticipation, and neutral. The emotion taxonomy combines the standard six basic emotions with two additional categories, trust and anticipation, plus a neutral class, and the metaphor annotation operationalizes the source-domain-to-target-domain mapping by detecting an irreversible 'A is B' relation across text, image, or both. When image and text convey conflicting emotions, the annotation takes the more intense emotion. The paper reports inter-annotator agreement coefficients from 0.58 for emotion categories to 0.68 for metaphor identification, and it documents a distributional pattern in which public-service advertisements mostly express fear and anticipation while commercial advertisements mostly express surprise, with the two types complementing each other across the whole dataset.","pith_inferences":[],"forward_implications":["If EmoMeta is a reliable resource, fine-grained emotion classification for Chinese multimodal metaphors becomes a concrete benchmark task rather than a data-scarce aspiration.","The documented emotion distributions give a testable expectation: public-service advertisements cluster around fear and anticipation, commercial advertisements around surprise, and the two types complement each other.","The annotation model (Occurrence, source domain, target domain, emotion category) can be transferred to other multimodal metaphor corpora, giving a common template for future annotation efforts.","The public release of the dataset provides a shared evaluation ground for models that must jointly reason about metaphor and emotion across text and image in Chinese.","Because 87% of the samples are public-service advertisements, models trained on EmoMeta will be especially sensitive to warning-driven fear rhetoric; applying them to other genres would be a separate evaluation.","An alternative annotation protocol that allows multiple emotion labels per pair, rather than forcing a single stronger emotion, could directly test whether the current labels erase mixed-emotion cases that fine-grained classification is supposed to capture.","The annotated source and target domains make EmoMeta usable for generation tasks, such as producing advertisement text or images conditioned on a chosen metaphor and emotion.","Chinese-specific metaphors, like the paper's example of 'dinosaur' meaning unattractive, suggest that cross-lingual transfer of metaphor-emotion models will be limited; EmoMeta can serve as a measurement ground for that gap."],"supporting_citations":[{"why":"Provides the source-domain to target-domain mapping definition that grounds the metaphor annotations.","marker":"[10]"},{"why":"Defines multimodal metaphor as cross-domain mapping across modes, which frames the text–image pair annotation.","marker":"[7]"},{"why":"Supplies the six basic emotions that form the core of the emotion taxonomy.","marker":"[3]"},{"why":"Adds trust and anticipation to the emotion set, yielding the finer-grained ten-way labeling scheme.","marker":"[18]"},{"why":"Contributes part of the advertisement data and the approach for determining whether a sample is metaphorical.","marker":"[22]"},{"why":"Contributes advertisement data and the metaphor-judgment methodology the annotation builds on.","marker":"[23]"},{"why":"Supplies the OCR model used to extract and then manually correct text from advertisement images.","marker":"[2]"},{"why":"Gives the inter-annotator agreement measurement used to report annotation reliability.","marker":"[5]"},{"why":"Inspires the fine-grained emotion label set that the paper adapts for metaphorical advertisements.","marker":"[1]"},{"why":"Provides the empirical link between metaphor and emotion that motivates building the dataset.","marker":"[14]"}],"fun_headline_variants":["EmoMeta: 5,000 Chinese metaphor ads tagged with 10 emotions","First Chinese multimodal dataset for fine-grained metaphor emotions","5k Chinese metaphor ads get emotion labels from joy to anticipation","New dataset adds trust and anticipation to metaphor emotion mapping","Image-text metaphor ads in Chinese now have 10 emotion labels"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The dataset's emotion labels are treated as reliable fine-grained ground truth even though annotators agreed only moderately on emotion categories ($\\kappa=0.58$), and the rule that the stronger emotion wins image–text conflicts can systematically erase genuinely mixed emotions.","fun_headline_variants_meta":{"raw":{"variants":["EmoMeta: 5,000 Chinese metaphor ads tagged with 10 emotions","First Chinese multimodal dataset for fine-grained metaphor emotions","5k Chinese metaphor ads get emotion labels from joy to anticipation","New dataset adds trust and anticipation to metaphor emotion mapping","Image-text metaphor ads in Chinese now have 10 emotion labels"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00047,"raw_usage":{"total_tokens":2320,"prompt_tokens":906,"completion_tokens":1414,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":522,"completion_tokens_details":{"reasoning_tokens":1329}},"tokens_in":522,"tokens_out":1414,"duration_ms":9339,"temperature":1.0,"reasoning_tokens":1329,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T22:16:50.155222+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-annotate a random sample of the image–text conflict cases using a protocol that permits multiple emotion labels per pair; if more than a fifth of those pairs receive a label different from the published single label, or if a classifier trained on the published labels performs at chance against the multi-label annotations, the claim that the dataset supports fine-grained ground-truth emotion classification is undercut.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the source-domain to target-domain mapping definition that grounds the metaphor annotations."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines multimodal metaphor as cross-domain mapping across modes, which frames the text–image pair annotation."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Adds trust and anticipation to the emotion set, yielding the finer-grained ten-way labeling scheme."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Contributes part of the advertisement data and the approach for determining whether a sample is metaphorical."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the empirical link between metaphor and emotion that motivates building the dataset."}],"review_version":1}