{"id":"61b581dd-ef27-4134-895d-e076e594904b","arxiv_id":"1908.03628","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":2.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A survey of deep learning methods for personality detection from text, audio, visual, and multimodal data, covering datasets, applications, and reported performance.","lead":"This paper reviews machine learning models for automatic personality detection, with a focus on deep learning and multimodal approaches. It surveys datasets, applications, and state-of-the-art models, but introduces no new experimental results.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Table 4 cannot support the state-of-the-art claim: it mixes classification accuracy, 1-MAE regression scores, and R-squared across different datasets and label protocols, and the text's ChaLearn winner is absent from the table.","rationale":"I read the paper as a survey whose value lies in organizing a broad literature, not in new experiments; its central claim is that it is the first comprehensive recent survey and that deep learning with multimodal fusion is state of the art. The first part is already questionable: Section 2 cites prior surveys [100], [52], [29], [1], [47] while asserting no recent overview exists. That could be defended if 'deep-learning-based' is the scope, but the claim as written is overstated. The second, more consequential part rests on Table 4. The table is internally inconsistent: no common metric, no common dataset, no common label protocol, and no baseline rows. The reader's weakest_assumption identified exactly this as the key risk. I agree, and would sharpen it: the problem is not just transcription errors but metric incommensurability and a concrete contradiction between Section 4.4's statement that DBR [111] achieved the highest ChaLearn accuracy and Table 4's assignment of the multimodal ChaLearn row to [38]. A survey can be corrected by adding methodology, reporting metric/splits, and tempering claims; the descriptive content is not fundamentally broken. Hence the CONDITIONAL verdict stands, and my read does not move it. I would additionally fix the Section 1.1 claim that Big Five are 'binary (yes/no) values'—standard Big-Five instruments are continuous Likert scales—because it signals the data descriptions need copy-editing.","tokens_in":20985,"tokens_out":7862,"duration_ms":84549,"concrete_test":"Reproduce the ChaLearn rows from [39], [38], and [111] using the official ChaLearn 2016 evaluation code and test partition. Record the metric (1-MAE vs accuracy), the partition, and exact scores; determine whether [111] or [38] is the best multimodal result. Then recover [64]'s Essays-I 57.99 and [98]'s AMI 64.84 from the original papers to verify they use the same metric type as their column header. If partitions or metric definitions differ, Table 4 cannot support the Section 5 SOTA conclusion.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The state-of-the-art sentence in Section 5 is supported only by Table 4, and Table 4 is not a comparative evidence base. The column 'Mean Best Accuracy' mixes classification accuracy (Essays, MBTI, FriendFeed, AMI, ELEA, Color FERET), the ChaLearn regression agreement score (1 - mean absolute error, so values around 0.91 are small errors, not 91% correct), and R-squared (YouTube Vlogs). Rows also use different datasets, trait sets, label sources (self-report vs perceived annotations vs judge ratings), and train/test partitions. The one place where two rows touch the same dataset—ChaLearn visual (90.94, [39]) vs multimodal (91.7, [38])—the text separately says DBR [111] won the ChaLearn 2016 challenge, yet [111] is absent from the table. Unless [38] and [39] used the official challenge partition and metric, the 0.76-point gap does not show multimodal fusion beats unimodal; if [111] beats [38], the table misattributes the best result. Outside ChaLearn no row is compared with a same-dataset alternative, so the column cannot justify the paper's central SOTA claim.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript is a survey of recent deep learning methods for automated personality detection, organized by input modality (text, audio, visual, bimodal, and trimodal). It reviews personality measures and applications, compiles popular datasets and feature-extraction tools, describes representative architectures, and presents a table of reported performance results. The paper's central claims are that it offers the first bird's-eye view of this area and that state-of-the-art personality detection is currently achieved by deep learning techniques combined with multimodal feature fusion.","tokens_in":21205,"tokens_out":3533,"duration_ms":37816,"significance":"If the survey's factual map is reliable, it would serve as a useful entry point for newcomers to the field, since it consolidates datasets, tools, and model families that are otherwise scattered across venue-specific papers. The paper also explicitly discusses fairness and ethics, which is a welcome dimension in this literature. The authors deserve credit for covering a broad range of papers across modalities and for making the survey's scope explicit. However, the central comparative claim about state-of-the-art methods rests on a performance table whose entries are not directly comparable, and the table is internally inconsistent with the text about the ChaLearn challenge winner. These issues are load-bearing because the survey draws its headline conclusion from that table.","major_comments":[{"comment":"The claim in Section 5 that 'the state of the art in personality detection has been achieved using deep learning techniques along with multimodal fusion of features' is not supported by Table 4. The column labeled 'Mean Best Accuracy' mixes incompatible quantities: classification accuracy percentages (Essays, MBTI, FriendFeed, AMI, ELEA, Color FERET), an R-squared value (YouTube Vlogs, 0.092), and ChaLearn first-impression scores around 0.91 that are 1 minus mean absolute error rather than classification accuracy. Since the rows also differ in dataset, trait set, label source, and train/test partition, the table cannot be used to compare methods across rows. The only two rows touching the same dataset (ChaLearn visual, 90.94 from [39]; ChaLearn multimodal, 91.7 from [38]) still require confirmation that both used the official challenge partition and evaluation metric. I recommend restructuring the table to state the metric explicitly for every row and to make comparisons only within the same dataset and metric.","section":"Section 5 and Table 4"},{"comment":"The text states that Deep Bimodal Regression (DBR) [111] 'achieved the highest accuracy in the ChaLearn Challenge 2016 for perceived personality analysis,' but [111] is absent from Table 4. Instead, the table reports ChaLearn results for [39] (visual only, 90.94) and [38] (multimodal, 91.7). If [111] is the challenge winner, its score should appear in the table with the official metric; otherwise the table may attribute the best result to [38] without justification. This inconsistency directly affects the paper's multimodal-fusion state-of-the-art claim, because the table must show whether the best reported ChaLearn number comes from a multimodal system and under which evaluation protocol.","section":"Section 4.4 and Table 4"},{"comment":"The Big-Five traits are described as 'binary (yes/no) values,' but standard Big-Five instruments such as the NEO-FFI and BFI-10 use continuous or multi-point Likert scales, and many of the papers discussed in this survey train regression models rather than binary classifiers. This is not merely a wording issue: it obscures the fact that Table 4 mixes classification and regression results. Please correct the definition and clarify how the different label representations used in the cited works map to classification versus regression.","section":"Section 1.1"}],"minor_comments":[{"comment":"The sentence 'The MBTI personality measure is the most popular personality measure used across the world right now' seems to conflict with Section 1.1, where the Big-Five is described as 'by far' the most popular measure in the automated personality detection literature. Please qualify the claim by domain (e.g., commercial use vs. academic research).","section":"Section 5"},{"comment":"The claim that the architecture of [81] 'performs better than the state of the art on IEMOCAP, MOUD and MOSI' refers to sentiment analysis datasets, not personality detection. Since the sentence appears in a personality-detection survey, please state explicitly that these are multimodal sentiment benchmarks and explain why the result is relevant to personality detection, or remove the sentence.","section":"Section 4.5"},{"comment":"The column header 'Mean Best Accuracy' is misleading for rows reporting R-squared or ChaLearn agreement scores. Please rename the column to something like 'Reported performance (metric)' and indicate the metric used in each row.","section":"Table 4"},{"comment":"The Aurora2 corpus and Columbia deception corpus are listed as audio datasets for personality detection, but their 'Personality Measure' entries are blank and they appear to be used for other tasks (speech recognition and deception detection). Please clarify their role in the surveyed personality-detection literature or remove them from the table.","section":"Table 2"},{"comment":"There are several typos and formatting issues, including 'Random Forrest' for 'Random Forest' (Section 4.2) and 'the the various image processing techniques' (Section 2).","section":"Various"},{"comment":"Some references are incomplete or lack venue details, for example [43] and [107]. Please supply full bibliographic information so readers can locate the cited works.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The manuscript includes several works by its own co-authors (e.g., refs. 10, 64, 81, 82, 83) as examples of approaches and in the performance table. This is not improper per se, but given that the paper claims to be a comprehensive 'bird's-eye view,' the authors should ensure the selection of representative works is balanced and that the self-citations do not inadvertently drive the state-of-the-art comparison. I would also note that the claim of being 'the first' survey of this scope is stronger than the related-work section itself suggests, since prior surveys focusing on visual or multimodal apparent personality are acknowledged."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Good survey, but don't trust Table 4.\n\nThe paper does something useful: it organizes the recent deep-learning personality detection literature by modality—text, audio, visual, multimodal—and catalogs datasets, feature tools, and applications. For a new graduate student or an engineer scouting the area, this is a convenient map. The multimodal emphasis is appropriate, and the paper correctly notes that most current work is bimodal audio-visual with little trimodal fusion. I'd give it credit for that organizational contribution.\n\nThe soft spots are mostly in the comparative claims. First, the 'first survey' claim is overstated: the paper itself cites Vinciarelli & Mohammadi (2014), Junior et al. (2018), and Escalera et al. (2018), which cover large parts of the same territory. Second—and more importantly—Table 4 cannot support the state-of-the-art conclusion in Section 5. The 'Mean Best Accuracy' column mixes classification accuracy, the ChaLearn regression agreement score (1 - MAE), and R-squared, across different datasets, trait sets, and label sources. The text calls DBR [111] the winner of ChaLearn 2016, but [111] never appears in the table. So the 0.76-point gap between the visual and multimodal ChaLearn rows doesn't establish that fusion beats unimodal, and no other row in the table is compared against a same-dataset baseline. The SOTA sentence would need to be cut or substantially qualified. Third, the Big-Five traits are introduced as binary yes/no values; they're continuous scales, even if some datasets binarize them for classification.\n\nThe descriptive core is sound enough, and the self-citations (Majumder et al., Poria et al.) are legitimate parts of the literature, not padding. But the paper needs a transparent selection methodology and a corrected Table 4 before I'd trust its comparative conclusions.\n\nWho's this for? A reader who wants a quick orientation to the field and doesn't need rigorous SOTA ranking. I'd send it to a serious referee, but with the expectation of major revision. I wouldn't cite it for any performance numbers.","headline":"Useful survey, but Table 4 can't support the SOTA claim.","tokens_in":21737,"tokens_out":2025,"would_cite":false,"duration_ms":20347,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Deep multimodal models now lead personality detection, survey finds","keywords":["personality detection","deep learning","multimodal fusion","Big Five personality traits","affective computing","survey","text modality","visual modality"],"falsifier":"Pick one dataset from Table 4, re-run the listed deep multimodal method and a strong unimodal or non-deep baseline under an identical train/test split and metric; if the unimodal or shallow model matches or beats the multimodal deep model, the paper's central state-of-the-art claim would not hold for that benchmark.","tokens_in":20775,"feed_emoji":"🧠","tokens_out":6094,"duration_ms":61806,"temperature":0.7,"pith_summary":"This survey tries to give the first up-to-date, bird's-eye view of machine-learning-based personality detection, with deep learning as the main focus. It organizes the field by input modality — text, audio, visual, and their combinations — and collects the datasets, feature-extraction tools, and published accuracies that a newcomer would need. The paper argues that the current state of the art comes from deep learning models that fuse features from more than one modality, most often audio and vision. If that map is accurate, it gives researchers a reliable starting point for choosing methods and benchmarks, and it points to trimodal fusion and better labelled data as the next open problems.","feed_headline":"Deep multimodal models now lead personality detection, survey finds","feed_subtitle":"A new review sorts the field by text, audio, and visual inputs and says the best results fuse them.","key_machinery":"The central object is the modality-based taxonomy of personality-detection systems: text, audio, visual, bimodal, and trimodal, with Table 4 ('Performance of the state-of-the-art methods on popular personality-detection datasets') as the comparative engine. The taxonomy carries the argument by showing where deep architectures have displaced shallow classifiers, and Table 4 supplies the empirical basis for saying that multimodal deep fusion sets current best results, with late fusion of audio and visual predictions being the most common winning recipe.","core_discovery":"The paper's central claim is that deep learning combined with multimodal feature fusion has become the dominant route to accurate automatic personality detection. It reports that visual features are the strongest single modality, that combining modalities usually beats any single one, that deep convolutional networks are the standard tool for visual personality inference, and that few systems yet exploit all three modalities together. On this basis it positions itself as the first review covering recent deep-learning-based and multimodal personality-detection work, and it uses a comparison table of 'Mean Best Accuracy' across popular datasets to support the state-of-the-art conclusion.","pith_inferences":["Beyond the paper: if the accuracy table is read at face value, the wide spread of architectures achieving similar scores suggests the binding constraint is labelled data, not model design.","Beyond the paper: the finding that visual features are most accurate in unimodal settings suggests perceived-personality benchmarks may reward appearance cues more than genuine behaviour; a fair test would compare systems on 'true personality' labels collected from self-reports.","Beyond the paper: the survey's own account implies an untested scalability claim — the end-to-end deep models it praises need large labelled datasets, yet most listed datasets are small; testing whether the same models hold up on a large newly collected corpus would be a direct check."],"forward_implications":["A newcomer can use the paper's dataset list and accuracy table to pick a benchmark and a baseline without redoing the literature search.","The reported pattern predicts that adding a text stream to audio-visual systems (trimodal fusion) is the most promising near-term direction, since few published systems try it.","If the field follows the paper's expectation, personality detection will move from social-media text toward audio, video, and multimodal inputs for applications such as assistants, job screening, and recommendation.","The comparison table implies that Big Five datasets dominate the field, so other measures such as MBTI and PEN remain under-resourced for deep learning."],"supporting_citations":[{"why":"It supplies the earlier 'personality computing' survey that this paper uses as the baseline it claims to update and extend.","marker":"[100]"},{"why":"It gives the earlier image-focused review whose coverage gap motivates the visual and multimodal parts of the survey.","marker":"[47]"},{"why":"It is the computer-vision survey of apparent personality that the paper says still leaves deep multimodal methods uncovered.","marker":"[52]"},{"why":"It is the compilation of apparent-personality work cited to show that no existing survey covers recent deep multimodal models.","marker":"[29]"},{"why":"It supplies the stream-of-consciousness essay dataset that anchors the text-modality accuracy row in the comparison table.","marker":"[74]"},{"why":"It supplies the ChaLearn First Impressions benchmark behind the top visual and multimodal accuracy entries.","marker":"[80]"},{"why":"It reports the deep bimodal regression framework that the paper identifies as the highest-accuracy entry in the ChaLearn 2016 challenge.","marker":"[111]"},{"why":"It reports the deep residual network whose 91.7 multimodal accuracy is the best multimodal number in the comparison table.","marker":"[38]"},{"why":"It reports the deep CNN text model whose 57.99 accuracy is the text-modality state-of-the-art entry in the comparison table.","marker":"[64]"}],"fun_headline_variants":["Deep learning fusion leads personality detection survey","Multimodal deep nets beat single cues for personality","Survey: best personality AI fuses text, audio, and video","First deep multimodal review: visual cues strongest alone","Personality detection: deep multimodal is new state of the art"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The survey's conclusions about what is state of the art assume that the 'Mean Best Accuracy' numbers in Table 4 are correctly transcribed and comparable across different datasets and metrics.","fun_headline_variants_meta":{"raw":{"variants":["Deep learning fusion leads personality detection survey","Multimodal deep nets beat single cues for personality","Survey: best personality AI fuses text, audio, and video","First deep multimodal review: visual cues strongest alone","Personality detection: deep multimodal is new state of the art"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000676,"raw_usage":{"total_tokens":2979,"prompt_tokens":750,"completion_tokens":2229,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":366,"completion_tokens_details":{"reasoning_tokens":2152}},"tokens_in":366,"tokens_out":2229,"duration_ms":15704,"temperature":1.0,"reasoning_tokens":2152,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T14:32:43.911103+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Pick one dataset from Table 4, re-run the listed deep multimodal method and a strong unimodal or non-deep baseline under an identical train/test split and metric; if the unimodal or shallow model matches or beats the multimodal deep model, the paper's central state-of-the-art claim would not hold for that benchmark.","supporting_citations":[{"cited_title":"IEEE Transactions on Affective Computing 5(3), 273–291 (2014)","cited_arxiv_id":null,"evidence_quote":"It supplies the earlier 'personality computing' survey that this paper uses as the baseline it claims to update and extend."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It gives the earlier image-focused review whose coverage gap motivates the visual and multimodal parts of the survey."},{"cited_title":"Journal of personality and social psychology 77(6), 1296 (1999)","cited_arxiv_id":null,"evidence_quote":"It supplies the stream-of-consciousness essay dataset that anchors the text-modality accuracy row in the comparison table."},{"cited_title":"In: European Conference on Computer Vision, pp","cited_arxiv_id":null,"evidence_quote":"It supplies the ChaLearn First Impressions benchmark behind the top visual and multimodal accuracy entries."},{"cited_title":"In: European Conference on Computer Vision, pp","cited_arxiv_id":null,"evidence_quote":"It reports the deep bimodal regression framework that the paper identifies as the highest-accuracy entry in the ChaLearn 2016 challenge."},{"cited_title":"IEEE Intelligent Systems 32(2), 74–79 (2017)","cited_arxiv_id":null,"evidence_quote":"It reports the deep CNN text model whose 57.99 accuracy is the text-modality state-of-the-art entry in the comparison table."}],"review_version":1}