{"id":"cd9fa8ff-19b1-4b8f-a747-00e8c37c2f64","arxiv_id":"2411.08347","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"CMACD links 11,338 Weibo users' MBTI types to 566,900 posts annotated with six emotion intensity scores, validated with classification benchmarks.","lead":"This paper introduces CMACD, a Chinese dataset from Weibo that pairs each user's self-reported MBTI personality type with intensity-scored labels for six emotions. It addresses a real gap, the scarcity of Chinese affective computing resources that combine personality and fine-grained emotion, but the full dataset is not publicly released and the emotion labels come from the authors' own model.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The dataset's novel multi-label intensity annotations rest entirely on EQN's self-trained pseudo-labels; the manual check verifies only top-1/top-2 single emotions, not the multi-label intensity vectors, so the 'strong utility' claim is not yet established.","rationale":"The reader's weakest assumption is correct and is the load-bearing point. The paper's novelty is not just a collection of posts with MBTI tags; it is the multi-label intensity annotations. Those annotations come from an unpublished model whose training procedure includes a self-training loop, and the only human check is single-label top-emotion agreement. That check cannot validate the multi-label intensity vectors that define the dataset. The high BERT benchmark is not independent evidence because labels and classifier come from the same modeling family. This does not mean the dataset is worthless: the MBTI part has some independent support from manual screening and reasonable axis correlations, and the top-label agreement suggests signal. But the central 'strong utility' claim is conditional on external validation and data release. I therefore keep the reader's CONDITIONAL verdict rather than moving to REJECT or ACCEPT.","tokens_in":16818,"tokens_out":3649,"duration_ms":38343,"concrete_test":"Release a random sample of at least 500 posts with EQN labels withheld; have at least three trained annotators independently assign multi-label six-emotion labels with intensity ratings on the same 0-1 scale and thresholding rule; then compute per-post exact-match accuracy, per-emotion label F1, and intensity correlation (ICC or Pearson) between human and EQN labels. If exact-match agreement falls below approximately 50% or per-emotion ICC falls below 0.4, the multi-label intensity ground-truth claim is not supported. Additionally, report sensitivity of the downstream BERT benchmark to threshold values in [0, 0.1].","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim of 'strong utility' depends on the EQN-generated six-emotion intensity scores being accurate enough to serve as ground truth, but the paper provides no evidence that the multi-label intensity vectors are human-validated. In Methods, 'Labeling Emotions with Intensity Scores Using the EQN Framework', EQN is trained via self-training: starting from the single-label SMP2020-EWECT corpus, the model's own outputs are used as pseudo-labels for previously unlabeled emotions, with original labels set to 1; the retrained model then labels all posts. The Technical Validation manual spot check only compares the top one or top two EQN energy labels against a single human label on 1,000 posts (83.1% top-1, 92.3% top-2). It does not measure agreement on the remaining four labels, does not measure intensity calibration, and does not test the 0.05 threshold. The multi-label classification benchmark in Table 4 trains models on the same machine-generated labels; high E1_Acc and Ex_Acc, including BERT at 0.9284, demonstrate that the labels are learnable and self-consistent, not that they correspond to human judgments. Because the full CMACD is not released (only a small sample), the EQN labels cannot currently be independently audited. If EQN systematically emits or omits micro-emotion labels, every MBTI-emotion analysis and benchmark in the paper inherits that error.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces CMACD, a Chinese multi-label affective computing dataset built from Weibo posts. The dataset consists of 566,900 posts from 11,338 users who publicly self-identify with one of 16 MBTI personality types; each post is annotated by the authors' EQN framework with intensity scores for six emotions (anger, fear, happiness, neutrality, sadness, surprise). The authors report MBTI text-classification benchmarks on four binary axes, multi-label emotion classification results, manual spot-check accuracy for top emotion labels, and correlation analyses between MBTI axes and between emotions. The central claim is that CMACD is the first Chinese dataset integrating personality traits with multi-label emotion intensity labels and that it demonstrates strong utility for affective computing.","tokens_in":17113,"tokens_out":4270,"duration_ms":45172,"significance":"If the emotion labels were shown to be reliable, CMACD would fill a genuine gap: a large Chinese-language resource linking personality (MBTI) with multi-label, intensity-valued emotion annotations would be valuable for psychology, NLP, and computational social science. The scale (566,900 posts, 11,338 users) and the inclusion of public user personality tags are notable strengths, as is the authors' attention to privacy (removing URLs, names, and identifiers). The paper also makes a small public sample and the annotation code available, and it provides baseline benchmarks across several models. However, the central value of the dataset rests on the validity of machine-generated emotion intensity labels, and the current validation evidence is insufficient; this is the main factor limiting the paper's contribution.","major_comments":[{"comment":"The six-emotion intensity labels are generated by the authors' own EQN framework, which is trained on the single-label SMP2020-EWECT corpus using a self-training/regression procedure in which the model's own outputs serve as pseudo-labels for previously unlabeled emotions. The only human validation, reported in the Manual Spot Check, compares the top one or top two EQN energy labels against a single human label on 1,000 posts (83.1% top-1, 92.3% top-2). This check does not evaluate the correctness of the remaining four labels, does not compare intensity values, does not test the t=0.05 threshold, and reports no inter-annotator agreement. Since the paper's central contribution is precisely the multi-label intensity annotation, this validation gap is load-bearing and must be addressed with human evaluation of the full emotion vectors (e.g., exact-match accuracy, per-label precision/recall, and intensity agreement).","section":"Methods, 'Labeling Emotions with Intensity Scores Using the EQN Framework'; Technical Validation, 'Manual Spot Check…"},{"comment":"The multi-label classification experiments train and evaluate models on the same EQN-generated labels, so the high E1_Acc and Ex_Acc values (e.g., BERT E1_Acc 0.9284) demonstrate that the labels are learnable and internally consistent, but they do not show that the labels correspond to human judgments. This is a circularity concern: the benchmark only proves that a model can reproduce the annotation tool's outputs. The paper needs an external criterion—such as comparison with independent human multi-label annotations, exact-match agreement, or threshold-sensitivity analysis—before claiming the dataset has 'strong utility.'","section":"Technical Validation, 'Multi-label Classification Experiment to Assess Dataset Usability' and Table 4"},{"comment":"The full CMACD dataset is not actually released: Usage Notes state that access requires an email request, and Code Availability states that only a small sample is publicly available on GitHub. Because the emotion labels are machine-generated and cannot be independently audited without the full label vectors, the absence of public release (or an access protocol with clear terms) prevents independent verification of the central claims and limits the reproducibility of Table 4 and of all downstream analyses.","section":"Usage Notes and Data Records"},{"comment":"There are inconsistent post counts in the manuscript: the Abstract and Data Records state 566,900 posts, while Section 5 states '566,950 posts from 11,338 users'; note that 11,338 users × 50 posts = 566,900. In addition, Table 2's axis sums give 7,058 + 4,281 = 11,339, and the same inconsistency appears for all four axes. These arithmetic inconsistencies undermine confidence in the exact statistics and need to be corrected and reconciled.","section":"Data Records and 'Quantitative Emotional Data Analysis'"},{"comment":"The Pearson correlation analyses in Figs. 15 and 16 are internal sanity checks: a weak MBTI axis correlation or a 'reasonable' emotion correlation structure is consistent with many possible label-generation schemes, including biased or artifact-prone ones. These analyses do not provide quantitative evidence of annotation accuracy. The paper should either be more cautious in framing these results or replace them with direct human-validation metrics; in particular, the emotion correlation heatmap should not be presented as evidence that the annotation accuracy is 'relatively high.'","section":"Technical Validation, 'MBTI Axis Correlation Test' and 'Evaluation of the Reasonableness of the Overall Distribution…"}],"minor_comments":[{"comment":"The claims that CMACD is 'the first dataset to unify personality and emotion' and 'the first large-scale dataset of Chinese personality traits' should be substantially justified with a more systematic comparison to prior resources; otherwise these claims read as overstatements.","section":"Abstract and Background & Summary"},{"comment":"The description of the EQN self-training procedure is difficult to follow: the text says the model's outputs are used as pseudo-labels for unlabeled emotions and then a 'regression adjustment' sets original labels to 1, but it is not explained how the regression targets are formed or why this avoids circularity. More detail is needed, especially since reference 20 is an unpublished preprint.","section":"Methods, 'Labeling Emotions with Intensity Scores Using the EQN Framework' and Fig. 3"},{"comment":"There are notation and clarity issues in the formulas: Eq. (2) defines Pnum as a sum over F(i), which is not obviously the number of posts; the subscripts and summation ranges in Eqs. (3)-(10) are garbled; and the threshold t is written inconsistently. These formulas should be rewritten with clear definitions.","section":"Quantitative Emotional Data Analysis, Eq. (1)-(10)"},{"comment":"There are several typos in Table 3: KNN reports '0.62.90' and AdaBoost J/P reports '0,5701'. These should be corrected, and the table would benefit from a consistent number of decimal places.","section":"Table 3"},{"comment":"The definition of Ex_Acc is unclear: the text says it is 'Accuracy of labels with non-maximum energy scores,' but it likely means something like exact-match or partial-match accuracy over the non-top labels. Please clarify the metric and state how the label threshold is applied for evaluation.","section":"Technical Validation, 'Multi-label Classification Experiment'"},{"comment":"The correlation heatmaps are described in the text but the figures themselves are not shown in the provided manuscript; if they are included in the final version, they need color bars and value labels so that the reported correlation strengths can be verified.","section":"Figures 15 and 16"}],"recommendation":"major_revision","confidential_remarks":"This manuscript describes a potentially useful resource, but the current validation of the emotion labels is not sufficient for a data descriptor. The paper's central claim depends on the correctness of machine-generated multi-label intensity scores that are not independently verified. I recommend requiring (a) a proper human evaluation of full multi-label intensity vectors, including an exact-match metric and inter-annotator agreement; (b) public release of the full label set or a clear, reproducible access mechanism; and (c) correction of the numerical inconsistencies in the post and user counts. The paper is not beyond repair, but it needs substantial additional validation."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The dataset idea is genuinely new: pairing self-reported MBTI labels from Weibo users with fine-grained multi-emotion intensity annotations. Nothing else public does that for Chinese. The collection pipeline for MBTI users is careful (manual screening, filtering bots, excluding personality keywords in validation), and the reported benchmark numbers show the labels are at least learnable.\n\nThe soft spot is exactly where your reader put it. The six emotion intensity labels are machine outputs from the authors' own EQN framework, trained via self-training on a single-label corpus. The only human check is a 1,000-post spot check comparing top-1/top-2 predicted single labels against one human label per post. That does not validate the multi-label intensity vector, the intensity calibration, or the 0.05 threshold. So the 'strong utility' claim is not yet established. The classifier benchmarks in Table 4 mostly demonstrate self-consistency, not human agreement. There is also a numeric inconsistency (566,900 vs 566,950 posts) and the full dataset is gated behind an email request, so independent audit is not possible right now.\n\nThese are fixable. The authors should release the full data with a data statement, report exact-match and per-intensity accuracy against human multi-label annotations from multiple annotators, and test threshold sensitivity. If they can show that the intensity vectors track human judgments, this becomes a valuable resource. The MBTI side is relatively solid on its own.\n\nVerdict: deserving of peer review, not desk rejection. The gap is real, the work is presented with standard settings for reproducibility, and the authors are transparent about their method. But a serious referee should demand the validation before the utility claim is accepted. I wouldn't cite it in my own work yet.","headline":"A plausible gap-filling Chinese MBTI-emotion resource, but the emotion labels need human validation before the utility claim can be trusted.","tokens_in":17623,"tokens_out":2179,"would_cite":false,"duration_ms":22088,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The authors introduce CMACD, the first Chinese affective computing dataset that pairs each user's MBTI personality type with multi-label emotion intensity scores on 566,900 Weibo posts, and they validate it with benchmark classifiers.","keywords":["affective computing","multi-label emotion dataset","emotion intensity","MBTI personality","Chinese social media","Weibo","EQN framework","BERT"],"falsifier":"Randomly sample posts from CMACD and have trained annotators assign independent intensity scores for all six emotions, or at least label all present emotions, then compare their judgments to the EQN scores. If agreement on secondary emotions falls near chance, the multi-label intensity portion of the dataset is not validated.","tokens_in":16597,"feed_emoji":"🧠","tokens_out":4397,"duration_ms":41438,"temperature":0.7,"pith_summary":"This paper introduces CMACD, a large Chinese affective computing dataset built from 566,900 Weibo posts by 11,338 users who publicly self-identify with one of the 16 MBTI personality types. The paper's claim is that CMACD is the first dataset to combine personality labels with multi-label emotion annotations: each post carries intensity scores for six emotions (anger, fear, happiness, neutral, sadness, surprise) generated by the authors' EQN framework. The authors argue this fills a gap left by scarce Chinese emotion datasets and by English resources such as GoEmotions, which lack personality labels and intensity values. Validation with several classifiers is offered as evidence of the dataset's usability, with BERT reaching 0.9284 top-emotion accuracy and 0.74 to 0.80 accuracy on the four MBTI axes. If the claim holds, the field gains a reusable benchmark for studying how stable personality traits and fine-grained emotional expression relate in Chinese social media text.","feed_headline":"New Chinese dataset pairs MBTI tags with emotion intensity","feed_subtitle":"566,900 Weibo posts from 11,338 users carry personality labels and multi-emotion intensity scores for affective computing.","key_machinery":"The load-bearing machinery is the EQN (Extended Quantization Network) annotation framework, a BERT-based model trained on a manually annotated single-label Weibo emotion dataset, then adjusted by regression so that originally labeled emotions are set to 1 and EQN-generated values are retained for previously unlabeled emotions, then retrained to predict six emotion intensities per post. This framework is what turns raw posts into multi-label intensity annotations. The dataset itself is the other half of the machinery: user posts are organized into 16 MBTI folders with per-user CSV files, giving future researchers a ready-made structure for personality-aware emotion modeling.","core_discovery":"The central discovery is a first-of-its-kind resource: a Chinese multi-label affective computing dataset in which the same user's MBTI personality type and per-post emotion intensities are recorded together. Each of the 566,900 posts is labeled with scores in the range 0 to 1 for anger, fear, happiness, neutral, sadness, and surprise; scores below 0.05 are set to zero, so a post can carry several valid emotions at once. The personality labels come from Weibo users' self-identification, screened by manual review of phrases for each MBTI type, and the emotion intensities come from machine annotation with the EQN framework trained on the SMP2020 single-label Weibo emotion dataset. The authors demonstrate utility by reporting multi-label emotion classification (top-label accuracies from 0.6577 for SVM to 0.9284 for BERT), MBTI four-axis binary classification (BERT accuracies 0.7388 to 0.7963), a 1,000-post manual spot check with 83.1 percent agreement on the top emotion label, and Pearson correlation patterns among emotions that match everyday expectations.","pith_inferences":["The reliability of the dataset's micro-emotion labels is inherited from EQN, so a direct human multi-label agreement study would tell whether the non-top intensity scores are trustworthy or primarily noise; the paper's spot check only verifies the top label.","Because EQN was trained on a single-label dataset, the multi-label intensity annotations are likely shaped by the model's learned label correlations; users of CMACD should treat the intensity vectors as model predictions rather than human ground truth.","The skewed distribution across MBTI types, with INFP users outnumbering ESTP users by about 12 to 1, means personality classification benchmarks should be read with the class imbalance in mind; a balanced subsample would give a cleaner estimate of personality signal.","A natural extension would be to test whether BERT's high emotion accuracy on CMACD transfers to fully human-annotated multi-label Chinese emotion data, which would indicate how much of the reported performance comes from label patterns rather than text content."],"forward_implications":["CMACD gives Chinese-language affective computing a single benchmark that joins personality and emotion, so models can be trained and compared on tasks that require both at once.","Researchers can use the paired MBTI and emotion-intensity labels to test whether stable personality types show measurable differences in emotional expression, as the paper's statistics suggest.","The BERT results on both emotion and MBTI classification provide reference numbers that future systems can be measured against.","The dataset supports downstream work in psychology, education, marketing, finance, and politics by supplying Chinese social-media text with quantified emotional states and creator personality."],"supporting_citations":[{"why":"Supplies the EQN framework that produces the six emotion intensity labels; the paper applies this framework to the MBTI posts.","marker":"20"},{"why":"Provides the SMP2020 single-label Weibo emotion training set that EQN is fine-tuned on.","marker":"41"},{"why":"Identifies the BERT model class used both in EQN and in the validation experiments.","marker":"42"},{"why":"GoEmotions serves as the main comparison point for a fine-grained multi-label emotion dataset that lacks intensity and personality labels.","marker":"19"},{"why":"ChineseEmoBank is cited as the prior Chinese fine-grained emotion intensity resource that the new dataset extends toward multi-label personality-aware annotation.","marker":"18"}],"fun_headline_variants":["Chinese dataset links MBTI tags to emotion intensity","566K Weibo posts pair MBTI with multi-label emotions","MBTI-labeled Weibo posts gain emotion intensity scores","New Chinese corpus joins personality and per-emotion scores","Chinese affective dataset ties MBTI to micro-emotion levels"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The dataset's value depends on EQN's machine-generated intensity scores for all six emotions being accurate enough to serve as ground truth, even though the model was trained on single labels and the manual check only tested the top emotion.","fun_headline_variants_meta":{"raw":{"variants":["Chinese dataset links MBTI tags to emotion intensity","566K Weibo posts pair MBTI with multi-label emotions","MBTI-labeled Weibo posts gain emotion intensity scores","New Chinese corpus joins personality and per-emotion scores","Chinese affective dataset ties MBTI to micro-emotion levels"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000716,"raw_usage":{"total_tokens":3225,"prompt_tokens":962,"completion_tokens":2263,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":578,"completion_tokens_details":{"reasoning_tokens":2185}},"tokens_in":578,"tokens_out":2263,"duration_ms":17369,"temperature":1.0,"reasoning_tokens":2185,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T21:40:33.271479+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Randomly sample posts from CMACD and have trained annotators assign independent intensity scores for all six emotions, or at least label all present emotions, then compare their judgments to the EQN scores. If agreement on secondary emotions falls near chance, the multi-label intensity portion of the dataset is not validated.","supporting_citations":[{"cited_title":"International Mother Language Day","cited_arxiv_id":null,"evidence_quote":"Supplies the EQN framework that produces the six emotion intensity labels; the paper applies this framework to the MBTI posts."}],"review_version":1}