{"id":"ab849753-1b51-4b2a-b3cb-79d51241ef77","arxiv_id":"1904.09728","paper_version":3,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"SocialIQA is the first large-scale benchmark with 38k crowdsourced questions testing commonsense about social interactions, where pretrained language models trail humans by over 20% but transfer to improve performance on Winograd Schemas and COPA.","lead":"The paper introduces SocialIQA, a new benchmark of 38,000 multiple-choice questions about why people behave in everyday social situations. A smart generalist might read it to understand current gaps in AI social intelligence and how new data can improve models on related reasoning tasks.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"Crowdsourcing mitigation may leave residual artifacts that models exploit instead of true social reasoning","rationale":"The reader's weakest assumption directly identifies the data-validity dependency that underpins both the challenge and transfer claims; no stronger internal inconsistency appears from the given description.","tokens_in":1687,"tokens_out":298,"duration_ms":27858,"concrete_test":"Re-collect a 2k-question subset using an alternative distractor method (e.g., independent workers write three plausible but incorrect answers without seeing a related question) and re-evaluate the same pretrained models; if the human-model gap shrinks by >10 points or the transfer improvements on Winograd/COPA disappear, the original framework's mitigation is insufficient.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline claims (20+% model-human gap plus SOTA transfer to Winograd/COPA) rest on the assumption that the 38k questions probe genuine emotional/social intelligence. The described framework (workers supply correct answer to a related question to generate distractors) reduces obvious stylistic cues, yet it can still embed generation-specific patterns: workers may systematically choose distractors that are plausible only relative to the paired question, or the resulting answer distributions may correlate with surface features of the original prompt. If models learn these meta-patterns rather than the intended commonsense, both the difficulty gap and the transfer gains become artifacts of the collection procedure rather than evidence of improved reasoning.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper introduces SocialIQA, a crowdsourced benchmark of 38,000 multiple-choice questions targeting commonsense reasoning about social and emotional situations. It reports that pretrained LM-based QA models lag human performance by more than 20% and demonstrates that fine-tuning on SocialIQA yields state-of-the-art transfer results on the Winograd Schema Challenge and COPA.","tokens_in":1804,"tokens_out":508,"duration_ms":30385,"significance":"If the questions genuinely probe social commonsense rather than collection artifacts, the benchmark would be a valuable addition for evaluating and improving AI social reasoning, with the transfer gains providing concrete evidence of utility. The scale and the explicit transfer experiments are strengths.","major_comments":[{"comment":"Data Collection section: The mitigation framework (workers supply correct answers to related questions to generate distractors) is described as reducing stylistic artifacts, yet no quantitative analysis is provided on whether residual patterns (e.g., answer distributions correlated with prompt surface features or generation-specific meta-patterns) remain exploitable by models. This directly affects the validity of both the >20% human-model gap and the transfer claims.","section":"Data Collection"},{"comment":"Experiments section (results tables): The reported model accuracies, human baseline, and transfer SOTA numbers lack details on statistical significance testing, variance across runs, or error analysis broken down by question type. Without these, the robustness of the central difficulty and transfer claims cannot be fully assessed.","section":"Experiments"},{"comment":"Transfer experiments: The SOTA results on Winograd Schemas and COPA are presented without ablations isolating the contribution of SocialIQA data versus other factors, and without comparison to more recent strong baselines available at the time of submission.","section":"Transfer Learning"}],"minor_comments":[{"comment":"Abstract: Specify example model families (e.g., BERT, GPT) when referring to 'pretrained language models' for immediate clarity.","section":"Abstract"},{"comment":"Related Work: Add explicit comparison to contemporaneous social reasoning datasets to sharpen the novelty claim.","section":"Related Work"}],"recommendation":"major_revision","confidential_remarks":"The crowdsourcing artifact concern is load-bearing and should be the focus of revision; if unaddressed, it could limit suitability for a top-tier NLP venue."},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive feedback on our SocialIQA benchmark paper. We address each major comment below with honest responses and indicate where revisions will be made to strengthen the manuscript.","responses":[{"response":"We agree that a quantitative analysis of residual artifacts would further validate the benchmark. The framework was specifically designed to reduce stylistic biases by requiring workers to answer a related question correctly before generating distractors, which we believe minimizes common patterns. However, we did not include such an analysis in the original submission. In revision, we will add a section quantifying answer distributions, correlations with surface features, and simple model exploitability tests (e.g., using bag-of-words baselines) to demonstrate that residual patterns do not explain the performance gap.","revision_made":"yes","referee_comment":"[Data Collection] Data Collection section: The mitigation framework (workers supply correct answers to related questions to generate distractors) is described as reducing stylistic artifacts, yet no quantitative analysis is provided on whether residual patterns (e.g., answer distributions correlated with prompt surface features or generation-specific meta-patterns) remain exploitable by models. This directly affects the validity of both the >20% human-model gap and the transfer claims."},{"response":"We acknowledge this limitation in the original presentation. The reported numbers reflect single-run results from standard fine-tuning procedures, but we agree that variance and significance testing are important for robustness. In the revision, we will rerun key experiments with multiple random seeds to report means and standard deviations, include statistical significance tests (e.g., McNemar's test for comparisons), and add an error analysis section breaking down performance by question categories such as emotional vs. social inference.","revision_made":"yes","referee_comment":"[Experiments] Experiments section (results tables): The reported model accuracies, human baseline, and transfer SOTA numbers lack details on statistical significance testing, variance across runs, or error analysis broken down by question type. Without these, the robustness of the central difficulty and transfer claims cannot be fully assessed."},{"response":"The transfer results compare models fine-tuned on SocialIQA against their non-fine-tuned counterparts and prior SOTA at submission time (e.g., BERT-based models). We did not include exhaustive ablations isolating every factor, which is a fair critique. For recent baselines, the paper was submitted in 2019 and used the strongest available methods then; we will update the transfer section with additional comparisons to contemporaneous strong models and add a simple ablation table showing performance with and without SocialIQA fine-tuning to better isolate its contribution.","revision_made":"partial","referee_comment":"[Transfer Learning] Transfer experiments: The SOTA results on Winograd Schemas and COPA are presented without ablations isolating the contribution of SocialIQA data versus other factors, and without comparison to more recent strong baselines available at the time of submission."}],"tokens_in":1339,"tokens_out":623,"duration_ms":24987,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"SocialIQA is the first large-scale dataset aimed at everyday social interactions, with 38k multiple-choice questions collected through a crowdsourcing setup that asks workers to generate distractors by correctly answering a related question. That approach is a real improvement over standard methods for cutting obvious stylistic cues in the wrong answers.","headline":"SocialIQA gives a practical new benchmark for social commonsense that models still struggle with and that transfers to other tasks, though the crowdsourcing method may not fully eliminate exploitable patterns.","tokens_in":2245,"tokens_out":144,"would_cite":true,"duration_ms":23670,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":{"model":"grok-4.3","evidence":[{"relation":"unclear","rs_module":"LogicAsFunctionalEquation","rs_theorem":null,"paper_passage":"We introduce Social IQa, the first largescale benchmark for commonsense reasoning about social situations. Social IQa contains 38,000 multiple choice questions for probing emotional and social intelligence in a variety of everyday situations"},{"relation":"unclear","rs_module":"Cost","rs_theorem":null,"paper_passage":"Empirical results show that our benchmark is challenging for existing question-answering models based on pretrained language models, compared to human performance (>20% gap). Notably, we further establish Social IQa as a resource for transfer learning of commonsense knowledge, achieving state-of-the-art performance on multiple commonsense reasoning tasks (Winograd Schemas, COPA)."}],"headline":"SocialIQA benchmark for social commonsense reasoning has no connection to RS framework","alignment":"orthogonal","rationale":"The paper's central machinery is a crowdsourced 38k-question QA dataset for probing emotional/social intelligence, using ATOMIC seeding and question-switching to reduce artifacts, plus transfer results to Winograd/COPA. No reference to J-cost, golden ratio, 8-tick periodicity, distinction-to-physics forcing, or any RS theorem. It operates in NLP benchmark/transfer-learning space while RS is a foundational logic/physics derivation; the two domains share no machinery or predictions.","tokens_in":271335,"confidence":"high","tokens_out":332,"duration_ms":27816,"cache_read_input_tokens":64,"cache_creation_input_tokens":0},"lean_confirmation":{"model":"grok-4.3","status":"out_of_scope","citations":[],"rationale":"The load-bearing premise is an empirical claim about model performance on a benchmark (not a mathematical theorem). Shape-of-logic contains no theorem establishing it; the status is out_of_scope per the guidelines for empirical papers.","tokens_in":271086,"confidence":"moderate","tokens_out":164,"duration_ms":20887,"inferential_bridge":"The paper's central results are empirical measurements of model accuracy on a new dataset and transfer gains; no mathematical/structural identity is claimed that could be machine-checked in Lean. The dataset construction, crowdsourcing, and evaluation are empirical and cannot be Lean-proved.","load_bearing_premise":"The SocialIQA benchmark (38k crowdsourced QA pairs probing social/emotional intelligence) is challenging for pretrained LM-based QA models (>20% gap to human performance) and enables SOTA transfer on Winograd/COPA.","cache_read_input_tokens":64,"cache_creation_input_tokens":0},"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Social IQa is a 38,000-question benchmark that exposes a greater than 20 percent performance gap between humans and pretrained language models on social commonsense reasoning.","keywords":["social commonsense","benchmark","question answering","emotional intelligence","transfer learning","Winograd schemas","COPA"],"falsifier":"A model that reaches human-level accuracy on Social IQa questions without any training on the dataset itself would show that the claimed gap and transfer benefit do not hold.","tokens_in":2597,"feed_emoji":"🧠","tokens_out":640,"duration_ms":41680,"temperature":0.7,"pith_summary":"The paper introduces Social IQa as the first large-scale multiple-choice dataset for testing how well systems understand everyday social interactions and the emotions behind them. Questions cover situations such as why someone might lean in to share a secret, with correct and incorrect answers collected through crowdsourcing. A new collection method asks workers to supply the right answer to a related question in order to generate plausible but wrong options, reducing superficial cues that models could exploit. Existing question-answering systems built on pretrained language models fall more than 20 percent behind human accuracy, yet fine-tuning on Social IQa raises state-of-the-art results on established commonsense tasks including Winograd Schemas and COPA.","feed_headline":"New benchmark reveals over 20% gap in AI social reasoning","feed_subtitle":"Social IQa dataset of 38,000 questions challenges language models and improves results on related commonsense tasks via transfer learning.","key_machinery":"The Social IQa benchmark, constructed via a crowdsourcing framework that generates incorrect answers by soliciting correct answers to related questions.","core_discovery":"Social IQa contains 38,000 multiple-choice questions that probe emotional and social intelligence across ordinary situations. The dataset is constructed by crowdsourcing both questions and answers while using a framework that mitigates stylistic artifacts in the incorrect options. Pretrained language-model-based question-answering systems show a performance gap exceeding 20 percent relative to humans. When used for transfer learning, the same resource produces state-of-the-art results on multiple other commonsense reasoning benchmarks such as Winograd Schemas and COPA.","pith_inferences":["Social commonsense may not emerge reliably from standard language modeling objectives alone.","The collection method could be adapted to create similar benchmarks for physical or temporal commonsense.","Models might benefit from pairing the dataset with explicit social knowledge representations."],"forward_implications":["Pretrained language models lack robust representations of social and emotional reasoning.","Fine-tuning on social interaction data can improve performance on other commonsense benchmarks.","Future systems will need explicit mechanisms for social intelligence to close the observed gap.","The benchmark supplies a concrete testbed for measuring progress in social reasoning."],"fun_headline_variants":["SocialIQa shows 20% gap in AI social reasoning","38k questions test AI on social commonsense situations","SocialIQa enables transfer on Winograd COPA benchmarks","Dataset probes emotional intelligence via 38k social questions"],"cache_read_input_tokens":64,"weakest_assumption_plain":"The crowdsourced questions and answers capture genuine social commonsense rather than new biases that models can exploit without true understanding.","fun_headline_variants_meta":{"raw":{"variants":["SocialIQa shows 20% gap in AI social reasoning","38k questions test AI on social commonsense situations","SocialIQa enables transfer on Winograd COPA benchmarks","Dataset probes emotional intelligence via 38k social questions"]},"model":"grok-4.3","cost_usd":0.009749,"raw_usage":{"total_tokens":4249,"prompt_tokens":645,"num_sources_used":0,"completion_tokens":64,"cost_in_usd_ticks":97490500,"prompt_tokens_details":{"text_tokens":645,"audio_tokens":0,"image_tokens":0,"cached_tokens":64},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":3540,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":645,"tokens_out":64,"duration_ms":30608,"temperature":1.0,"reasoning_tokens":3540,"cache_read_input_tokens":64,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-05-13T12:18:16.743823+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A model that reaches human-level accuracy on Social IQa questions without any training on the dataset itself would show that the claimed gap and transfer benefit do not hold.","supporting_citations":[],"review_version":1}