REVIEW 3 major objections 2 minor 110 cited by
SocialIQA: Commonsense Reasoning about Social Interactions
T0 review · 3 major / 2 minor · reviewed 2026-05-13 · grok-4.3
Pith's one-line read Social IQa is a 38,000-question benchmark that exposes a greater than 20 percent performance gap between humans and pretrained language models on social commonsense reasoning.
desk verdict SocialIQA gives a practical new benchmark for social commonsense that models still struggle with and that transfers to other tasks, though the crowdsourcing method may not fully eliminate exploitable patterns. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The Social IQa benchmark, constructed via a crowdsourcing framework that generates incorrect answers by soliciting correct answers to related questions.
What would settle it
A model that reaches human-level accuracy on Social IQa questions without any training on the dataset itself would show that the claimed gap and transfer benefit do not hold.
Extended reading notes
Core claim
Social IQa contains 38,000 multiple-choice questions that probe emotional and social intelligence across ordinary situations. The dataset is constructed by crowdsourcing both questions and answers while using a framework that mitigates stylistic artifacts in the incorrect options. Pretrained language-model-based question-answering systems show a performance gap exceeding 20 percent relative to humans. When used for transfer learning, the same resource produces state-of-the-art results on multiple other commonsense reasoning benchmarks such as Winograd Schemas and COPA.
Load-bearing premise
The crowdsourced questions and answers capture genuine social commonsense rather than new biases that models can exploit without true understanding.
Editorial extensions
If this is right
- Pretrained language models lack robust representations of social and emotional reasoning.
- Fine-tuning on social interaction data can improve performance on other commonsense benchmarks.
- Future systems will need explicit mechanisms for social intelligence to close the observed gap.
- The benchmark supplies a concrete testbed for measuring progress in social reasoning.
Reading between the lines
- Social commonsense may not emerge reliably from standard language modeling objectives alone.
- The collection method could be adapted to create similar benchmarks for physical or temporal commonsense.
- Models might benefit from pairing the dataset with explicit social knowledge representations.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces SocialIQA, a crowdsourced benchmark of 38,000 multiple-choice questions targeting commonsense reasoning about social and emotional situations. It reports that pretrained LM-based QA models lag human performance by more than 20% and demonstrates that fine-tuning on SocialIQA yields state-of-the-art transfer results on the Winograd Schema Challenge and COPA.
Significance. If the questions genuinely probe social commonsense rather than collection artifacts, the benchmark would be a valuable addition for evaluating and improving AI social reasoning, with the transfer gains providing concrete evidence of utility. The scale and the explicit transfer experiments are strengths.
major comments (3)
- [Data Collection] Data Collection section: The mitigation framework (workers supply correct answers to related questions to generate distractors) is described as reducing stylistic artifacts, yet no quantitative analysis is provided on whether residual patterns (e.g., answer distributions correlated with prompt surface features or generation-specific meta-patterns) remain exploitable by models. This directly affects the validity of both the >20% human-model gap and the transfer claims.
- [Experiments] Experiments section (results tables): The reported model accuracies, human baseline, and transfer SOTA numbers lack details on statistical significance testing, variance across runs, or error analysis broken down by question type. Without these, the robustness of the central difficulty and transfer claims cannot be fully assessed.
- [Transfer Learning] Transfer experiments: The SOTA results on Winograd Schemas and COPA are presented without ablations isolating the contribution of SocialIQA data versus other factors, and without comparison to more recent strong baselines available at the time of submission.
minor comments (2)
- [Abstract] Abstract: Specify example model families (e.g., BERT, GPT) when referring to 'pretrained language models' for immediate clarity.
- [Related Work] Related Work: Add explicit comparison to contemporaneous social reasoning datasets to sharpen the novelty claim.
Simulated Author's Rebuttal
We thank the referee for the constructive feedback on our SocialIQA benchmark paper. We address each major comment below with honest responses and indicate where revisions will be made to strengthen the manuscript.
read point-by-point responses
-
Referee: [Data Collection] Data Collection section: The mitigation framework (workers supply correct answers to related questions to generate distractors) is described as reducing stylistic artifacts, yet no quantitative analysis is provided on whether residual patterns (e.g., answer distributions correlated with prompt surface features or generation-specific meta-patterns) remain exploitable by models. This directly affects the validity of both the >20% human-model gap and the transfer claims.
Authors: We agree that a quantitative analysis of residual artifacts would further validate the benchmark. The framework was specifically designed to reduce stylistic biases by requiring workers to answer a related question correctly before generating distractors, which we believe minimizes common patterns. However, we did not include such an analysis in the original submission. In revision, we will add a section quantifying answer distributions, correlations with surface features, and simple model exploitability tests (e.g., using bag-of-words baselines) to demonstrate that residual patterns do not explain the performance gap. revision: yes
-
Referee: [Experiments] Experiments section (results tables): The reported model accuracies, human baseline, and transfer SOTA numbers lack details on statistical significance testing, variance across runs, or error analysis broken down by question type. Without these, the robustness of the central difficulty and transfer claims cannot be fully assessed.
Authors: We acknowledge this limitation in the original presentation. The reported numbers reflect single-run results from standard fine-tuning procedures, but we agree that variance and significance testing are important for robustness. In the revision, we will rerun key experiments with multiple random seeds to report means and standard deviations, include statistical significance tests (e.g., McNemar's test for comparisons), and add an error analysis section breaking down performance by question categories such as emotional vs. social inference. revision: yes
-
Referee: [Transfer Learning] Transfer experiments: The SOTA results on Winograd Schemas and COPA are presented without ablations isolating the contribution of SocialIQA data versus other factors, and without comparison to more recent strong baselines available at the time of submission.
Authors: The transfer results compare models fine-tuned on SocialIQA against their non-fine-tuned counterparts and prior SOTA at submission time (e.g., BERT-based models). We did not include exhaustive ablations isolating every factor, which is a fair critique. For recent baselines, the paper was submitted in 2019 and used the strongest available methods then; we will update the transfer section with additional comparisons to contemporaneous strong models and add a simple ablation table showing performance with and without SocialIQA fine-tuning to better isolate its contribution. revision: partial
Circularity Check
No circularity; benchmark and results rest on new crowdsourced data collection and direct empirical evaluation
full rationale
The paper introduces Social IQa via a described crowdsourcing framework that generates questions and answers about social situations, then reports direct model evaluations (pretrained QA models vs. humans) and transfer experiments on Winograd/COPA. No equations, fitted parameters, or derivations are present. Claims do not reduce to self-citations, prior fits, or self-definitions; they are independent empirical measurements on the newly collected 38k-question resource. Minor prior-work citations exist for context but are not load-bearing for the performance gap or transfer results.
Assumptions & free parameters
assumptions (1)
- domain assumption Crowdsourced annotations from the described framework accurately reflect genuine social commonsense without residual artifacts
Cite this review
Pith. "Pith review of SocialIQA: Commonsense Reasoning about Social Interactions." pith.science (2026). https://pith.science/paper/GZAASV5G
@misc{pith2026190409728,
author = {Pith},
title = {Pith review of: SocialIQA: Commonsense Reasoning about Social Interactions},
year = {2026},
howpublished = {\url{https://pith.science/paper/GZAASV5G}},
note = {Machine review of arXiv:1904.09728}
}
read the original abstract
We introduce Social IQa, the first largescale benchmark for commonsense reasoning about social situations. Social IQa contains 38,000 multiple choice questions for probing emotional and social intelligence in a variety of everyday situations (e.g., Q: "Jordan wanted to tell Tracy a secret, so Jordan leaned towards Tracy. Why did Jordan do this?" A: "Make sure no one else could hear"). Through crowdsourcing, we collect commonsense questions along with correct and incorrect answers about social interactions, using a new framework that mitigates stylistic artifacts in incorrect answers by asking workers to provide the right answer to a different but related question. Empirical results show that our benchmark is challenging for existing question-answering models based on pretrained language models, compared to human performance (>20% gap). Notably, we further establish Social IQa as a resource for transfer learning of commonsense knowledge, achieving state-of-the-art performance on multiple commonsense reasoning tasks (Winograd Schemas, COPA).
Forward citations
Showing 60 of 110 Pith papers that cite this
-
Path-Constrained Mixture-of-Experts
PathMoE constrains expert paths in MoE models by sharing router parameters across layer blocks, yielding more concentrated paths, better performance on perplexity and tasks, and no need for auxiliary losses.
-
Deep Delta Learning
Replacing additive residual connections with a gated rank-1 delta update that interpolates identity, projection, and reflection slightly improves language modeling and downstream averages in reported 124M/353M runs.
-
MesaNet: Sequence Modeling by Locally Optimal Test-Time Training
MesaNet uses conjugate-gradient-optimal test-time regression in a chunkwise-parallelizable recurrent layer, achieving strong language modeling and benchmark performance at up to 1B scale.
-
R^3-VQA: "Read the Room" by Video Social Reasoning
R3-VQA is a new real-world video benchmark on which the best tested model, GPT-4o, scores 83% on generated questions but only 54% on human-written ones, while humans score 91% and 80%.
-
Pushing the Limits of Low-Bit Optimizers: A Focus on EMA Dynamics
SOLO compresses Adam optimizer states to 2 to 3 effective bits using p-quantile-based logarithmic quantization for second moments and momentum reduction for first moments, preserving accuracy on most tested benchmarks.
-
Social Human Robot Embodied Conversation (SHREC) Dataset: Benchmarking Foundational Models' Social Reasoning
SHREC is a new benchmark dataset of embodied human-robot conversations that shows substantial performance gaps in state-of-the-art foundation models on tasks involving social error detection and rationale generation.
-
SpinQuant: LLM quantization with learned rotations
SpinQuant learns optimal rotations to enable accurate 4-bit quantization of LLM weights, activations, and KV cache, reducing the zero-shot gap to full precision to 2.9 points on LLaMA-2 7B.
-
Cosmos QA: Machine Reading Comprehension with Contextual Commonsense Reasoning
Cosmos QA is a new multiple-choice reading comprehension benchmark built from personal blogs, where correct answers require commonsense inference beyond the literal text and machines trail humans by about 25 points.
-
HSS-Synth: Humanities and Social Sciences Data Synthesis for LLMs
HSS-Synth generates 230k instruction-tuning samples for 14 humanities/social-science fields and reports state-of-the-art fine-tuning results on 16 benchmarks.
-
Scaling Native Multimodal Pre-Training From Scratch
In models trained from scratch on text plus images, the text-objective scaling law is data-mix-invariant while the image-conditioned objective shifts toward many more tokens relative to parameters as the multimodal sh...
-
ELSA3D: Elastic Semantic Anchoring for Unified 3D Understanding and Generation
ELSA3D introduces elastic semantic anchoring via sparse anchor tokens and a scale-aware octree tokenizer to unify 3D generation and captioning at reduced computational cost.
-
Building Social World Models with Large Language Models
SWM framework uses LLMs to model social belief dynamics from events via temporal pattern mining and ELBO optimization, outperforming time-series models on a new 12k-point benchmark from Kalshi and Polymarket predictio...
-
LiftQuant: Continuous Bit-Width LLM via Dimensional Lifting and Projection
LiftQuant uses dimensional lifting of weights to a higher-dimensional 1-bit lattice followed by projection to achieve tunable continuous bit-widths in LLM quantization while remaining hardware-friendly.
-
LLMs as Noisy Channels: A Shannon Perspective on Model Capacity and Scaling Laws
The Shannon Scaling Law treats LLM training as noisy-channel transmission and predicts U-shaped performance degradation when signal-to-noise ratio falls below a threshold, outperforming monotonic scaling laws on Pythi...
-
One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs
LLR uses heavy-tailed self-regularization theory to set per-layer learning rates in Transformers, yielding faster convergence and higher zero-shot accuracy than uniform rates across model scales.
-
UB-SMoE: Universally Balanced Sparse Mixture-of-Experts for Resource-adaptive Federated Fine-tuning of Foundation Models
UB-SMoE balances expert utilization in heterogeneous federated SMoE fine-tuning via Dynamic Modulated Routing and Universal Pseudo-Gradient, delivering up to 45% compute reduction and 8.7x performance gains for low-re...
-
Learning to Remember, Learn, and Forget in Attention-Based Models
Palimpsa adds a per-slot importance/precision state to gated linear attention, letting a fixed-size memory forget stale information and protect important information, and recovers Mamba2 as a high-forgetting limit.
-
Decouple Searching from Training: Scaling Data Mixing via Model Merging for Large Language Model Pre-training
Weighted averaging of component models trained on individual data sources can serve as a cheap, faithful proxy for training on arbitrary data mixtures, enabling cheaper data-mix search for LLM pre-training.
-
SpecQuant: Spectral Decomposition and Adaptive Truncation for Ultra-Low-Bit LLMs Quantization
SpecQuant uses outlier smoothing into weights followed by channel-wise low-frequency Fourier truncation to achieve 4-bit quantization of LLaMA-3 8B with only 1.5% zero-shot accuracy loss versus full precision.
-
ScaLoRA: Optimally Scaled Low-Rank Adaptation for Efficient High-Rank Fine-Tuning
ScaLoRA analytically derives per-update column scalings that let low-rank increments accumulate into high-rank weight updates, yielding faster convergence and higher accuracy than prior LoRA variants on LLMs up to 12B...
-
Short window attention enables long-term memorization
Short sliding windows in hybrid attention-xLSTM models boost long-context performance by encouraging long-term memory use, and stochastic window sizing improves both short and long tasks.
-
HyperAdapt: Simple High-Rank Adaptation
HyperAdapt performs parameter-efficient fine-tuning by row- and column-wise diagonal scaling to induce high-rank updates with only n+m trainable parameters.
-
ShinkaEvolve: Towards Open-Ended And Sample-Efficient Program Evolution
ShinkaEvolve improves sample efficiency in LLM-driven program evolution via parent sampling, code novelty rejection-sampling, and bandit LLM ensemble selection, achieving new SOTA circle packing with 150 samples and g...
-
Reinforcement Learning with Rubric Anchors
Rubric-based rewards extend reinforcement learning to open-ended text generation, yielding a 30B model that outperforms a 671B model on humanities-style benchmarks.
-
Towards Higher Effective Rank in Parameter-efficient Fine-tuning using Khatri--Rao Product
KRAdapter, a Khatri-Rao product adapter, produces full-rank high-effective-rank weight updates for parameter-efficient fine-tuning and reports improved out-of-distribution performance over LoRA and other full-rank PEF...
-
Language Models Improve When Pretraining Data Matches Target Tasks
Ranking pretraining documents by similarity to benchmark training examples (BETR) yields consistent benchmark gains and a 2.1x compute multiplier over DCLM-Baseline.
-
Does Math Reasoning Improve General LLM Capabilities? Understanding Transferability of LLM Reasoning
Math reasoning gains in LLMs rarely transfer to general domains; RL tuning generalizes while SFT causes forgetting and representation drift.
-
Revisiting LoRA through the Lens of Parameter Redundancy: Spectral Encoding Helps
SeLoRA reparameterizes LoRA updates as inverse Fourier or wavelet transforms of sparsely masked spectral coefficients, improving fine-tuning accuracy on LLaMA models with fewer trainable parameters.
-
LIFELONG SOTOPIA: Evaluating Social Intelligence of Language Agents Over Lifelong Social Interactions
Language agents' believability and goal achievement decline over multi-episode social interactions, and curated memory summaries only partially close the gap with humans.
-
Mixture-of-Experts Can Surpass Dense LLMs Under Strictly Equal Resource
MoE models with activation rates in an optimal region outperform dense LLMs of identical total parameter count, training compute, and data budget, with the optimal region consistent across scales.
-
Token Signature: Predicting Chain-of-Thought Gains with Token Decoding Feature in Large Language Models
The monotonicity of token probabilities during initial decoding predicts chain-of-thought gains, enabling dynamic selection between CoT and direct answers.
-
Come Together, But Not Right Now: A Progressive Strategy to Boost Low-Rank Adaptation
Gradually increasing the probability that LoRA adapters stay active during fine-tuning improves generalization, merging, and pruning robustness.
-
The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text
A new 8TB openly-licensed text corpus trains 7B LLMs that are competitive with Llama 1/2, showing that performant models need not depend on unlicensed web data.
-
Chameleon: A Flexible Data-mixing Framework for Language Model Pretraining and Finetuning
Chameleon uses kernel ridge leverage scores on domain embeddings to set LLM training-mixture weights, matching DoGE-level pretraining quality at roughly one fifth the compute and improving finetuning perplexity.
-
Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation
Under simulated leakage, n-gram-based detection beats permutation and truncation methods, and cleaning flag-prone MMLU instances changes model rankings only slightly.
-
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models
SocialMaze is a six-task benchmark that claims to evaluate LLM social reasoning along deep reasoning, dynamic interaction, and information uncertainty dimensions.
-
Herd Behavior: Investigating Peer Influence in LLM-based Multi-Agent Systems
LLM agents flip their answers more when their own confidence is low and their peer seems confident, and the format and order of peer information can amplify or dampen this herd behavior.
-
BASE-Q: Bias and Asymmetric Scaling Enhanced Rotational Quantization for Large Language Models
BASE-Q combines bias correction and asymmetric scaling under fixed rotations to improve 4-bit weight-activation quantization, narrowing the accuracy gap to full precision by up to 50.5% over prior rotation-based methods.
-
RefLoRA: Refactored Low-Rank Adaptation for Efficient Fine-Tuning of Large Models
RefLoRA picks a per-step optimal low-rank factorization (a matrix geometric mean) that balances LoRA's factors, improving fine-tuning convergence and accuracy.
-
Memory-Efficient LLM Training by Various-Grained Low-Rank Projection of Gradients
Adding a 'granularity' reshape to low-rank gradient projection improves memory efficiency and, in most tested settings, accuracy at a fixed memory cost.
-
Nemotron-CLIMB: CLustering-based Iterative Data Mixture Bootstrapping for Language Model Pre-training
CLIMB automatically discovers pre-training data mixtures by clustering text embeddings and iteratively refining mixture weights with a predictor, improving 1B-model reasoning accuracy over standard baselines.
-
FLIP Reasoning Challenge
The FLIP benchmark of 11,674 blockchain image-story puzzles shows best open and closed AI models reach 75.5% and 77.9% accuracy, below the 95.3% human consensus baseline.
-
LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws
Pretraining data determines loss-to-loss scaling laws in LLMs, while model size, optimization, tokenizer, and architecture have limited impact.
-
Emergent Response Planning in LLMs
Hidden representations of LLM prompts encode global attributes of the upcoming response, and simple probes can predict length, content choices, and answer confidence before generation begins.
-
A Lightweight Method to Disrupt Memorized Sequences in LLM
A decoding-time intervention that substitutes a small model's probabilities for common function words into a large model's output reduces exact training-data recall by up to 10x with minimal measured quality loss.
-
Fine, I'll Merge It Myself: A Multi-Fidelity Framework for Automated Model Merging
An automated multi-fidelity search framework discovers layer-wise and depth-wise model merging recipes that improve single- and multi-objective LLM reasoning performance without retraining.
-
CE-LoRA: Computation-Efficient LoRA Fine-Tuning for Language Models
CE-LoRA accelerates LoRA fine-tuning by approximating the dense activation-gradient matrix multiply with selected rows and columns and a frozen low-rank correction, reporting up to 3.39x faster backward passes with ne...
-
Language Models Prefer What They Know: Relative Confidence Estimation via Confidence Preferences
Relative pairwise confidence comparisons aggregated by rank aggregation produce more reliable confidence scores for language models than direct absolute confidence prompts.
-
RandLoRA: Full-rank parameter-efficient fine-tuning of large models
RandLoRA achieves full-rank weight updates in parameter-efficient fine-tuning by learning diagonal scalings over fixed random low-rank bases, outperforming LoRA across vision and language tasks.
-
Federated Sketching LoRA: A Flexible Framework for Heterogeneous Collaborative Fine-Tuning of LLMs
Federated Sketching LoRA (FSLoRA) uses random row/column sketching of LoRA modules so each client updates a low-cost submatrix, with a convergence rate that scales with the sketching ratio.
-
CLoQ: Enhancing Fine-Tuning of Quantized LLMs via Calibrated LoRA Initialization
CLoQ initializes LoRA adapters on quantized LLMs with a closed-form calibration-aware low-rank solution, improving 2-bit fine-tuning accuracy.
-
eaSEL: Promoting Social-Emotional Learning and Parent-Child Interaction through AI-Mediated Content Consumption
A system that generates social-emotional learning activities from children's videos increased emotion-word use in 5-8 year olds' story retellings, and parents saw it as helping family conversations.
-
OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting
OSTQuant quantizes LLM weights, activations, and KV cache to 4 bits using learnable orthogonal and scaling transformations plus a new KL-Top loss, reporting near-lossless W4-only and strong W4A4KV4 results on LLaMA models.
-
Recurrent Diffusion for Large-Scale Parameter Generation
RPG generates full weights for models up to 200M parameters, including ConvNeXt-L and LLaMA LoRA adapters, at accuracy comparable to trained checkpoints, using recurrent token prototypes to condition a 1D diffusion model.
-
S$^{2}$FT: Efficient, Scalable and Generalizable LLM Fine-tuning by Structured Sparsity
S2FT selects a few attention heads and FFN channels, permutes the neighboring weight matrices so the selected parts form dense blocks, and fine-tunes only those blocks, reporting better generalization and efficiency t...
-
Liquid: Language Models are Scalable and Unified Multi-modal Generators
Liquid extends existing LLMs with VQGAN image tokens and shows unified visual understanding and generation can scale, with the language-versus-image trade-off shrinking as model size grows.
-
CLOVER: Cross-Layer Orthogonal Vectors Pruning and Fine-Tuning
Attention pairs (Q-K and V-O) are SVD-decomposed so pruning or fine-tuning touches only a small singular-factor matrix, yielding better pruning tolerance and small PEFT gains.
-
LaMI: Augmenting Large Language Models via Late Multi-Image Fusion
LaMI augments LLMs with visual commonsense via late fusion of predictions from multiple text-generated images, outperforming prior augmented LLMs on visual tasks while matching VLMs and preserving or improving NLP per...
-
Chain-of-Verification Reduces Hallucination in Large Language Models
Chain-of-Verification reduces hallucinations in large language models by drafting responses, planning independent verification questions, answering them separately, and generating a final verified output.
-
KagNet: Knowledge-Aware Graph Networks for Commonsense Reasoning
KagNet grounds question-answer pairs in ConceptNet schema graphs and uses a GCN-LSTM-HPA module to improve CommonsenseQA accuracy over BERT baselines.
Reference graph
Works this paper leans on
-
[1]
Ian Apperly. 2010. Mindreaders: the cognitive basis of" theory of mind". Psychology Press
work page 2010
-
[2]
Simon Baron-Cohen, Alan M Leslie, and Uta Frith. 1985. Does the Autistic Child have a ``Theory of Mind''? Cognition, 21(1):37--46
work page 1985
-
[3]
Ernest Davis and Gary Marcus. 2015. Commonsense reasoning and commonsense knowledge in artificial intelligence. Commun. ACM, 58:92--103
work page 2015
-
[4]
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT : Pre-training of deep bidirectional transformers for language understanding. In NAACL
work page 2019
-
[5]
Jos \'e H. Espinosa and Henry Lieberman. 2005. Eventnet: Inferring temporal relations between commonsense events. In MICAI
work page 2005
-
[6]
MY Ganaie and Hafiz Mudasir. 2015. A Study of Social Intelligence & Academic Achievement of College Students of District Srinagar, J&K, India . Journal of American Science, 11(3):23--27
work page 2015
-
[7]
Travis Goodwin, Bryan Rink, Kirk Roberts, and Sanda M Harabagiu. 2012. UTDHLT : Copacetic system for choosing plausible alternatives. In NAACL workshop on SemEval, pages 461--466. Association for Computational Linguistics
work page 2012
-
[8]
Andrew S Gordon and Jerry R Hobbs. 2017. A Formal Theory of Commonsense Psychology: How People Think People Think. Cambridge University Press
work page 2017
Show all 140 references
-
[11]
Bowman, and Noah A
Suchin Gururangan, Swabha Swayamdipta, Omer Levy, Roy Schwartz, Samuel R. Bowman, and Noah A. Smith. 2018. Annotation artifacts in natural language inference data. In NAACL-HLT
2018
-
[12]
Vid Kocijan, Ana-Maria Cretu, Oana-Maria Camburu, Yordan Yordanov, and Thomas Lukasiewicz. 2019. A surprisingly robust trick for the winograd schema challenge. In ACL
2019
-
[13]
Baris Korkmaz. 2011. Theory of mind and neurodevelopmental disorders of childhood. Pediatr Res, 69(5 Pt 2):101R--8R
2011
-
[14]
Douglas B Lenat. 1995. Cyc: A large-scale investment in knowledge infrastructure. Communications of the ACM, 38(11):33--38
1995
-
[15]
Levesque
Hector J. Levesque. 2011. The winograd schema challenge. In AAAI Spring Symposium: Logical Formalizations of Commonsense Reasoning
2011
-
[16]
Li Lucy and Jon Gauthier. 2017. Are distributional representations ready for the real world? evaluating word vectors for grounded perceptual meaning. In RoboNLP@ACL
2017
-
[17]
Zhiyi Luo, Yuchen Sha, Kenny Q Zhu, Seung-won Hwang, and Zhongyuan Wang. 2016. Commonsense causal reasoning between short texts. In Fifteenth International Conference on the Principles of Knowledge Representation and Reasoning
2016
-
[18]
Gary Marcus. 2018. Deep learning: A critical appraisal. CoRR, abs/1801.00631
2018
-
[19]
Saif Mohammad. 2018. Obtaining reliable human ratings of valence, arousal, and dominance for 20,000 english words. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 174--184
2018
-
[20]
Chris Moore. 2013. The development of commonsense psychology. Psychology Press
2013
-
[21]
Griffiths
Aida Nematzadeh, Kaylee Burns, Erin Grant, Alison Gopnik, and Thomas L. Griffiths. 2018. Evaluating theory of mind in question answering. In EMNLP
2018
-
[22]
Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer. 2017. Automatic differentiation in pytorch. In NIPS-W
2017
-
[23]
Haoruo Peng, Daniel Khashabi, and Dan Roth. 2015. Solving hard coreference problems. In HLT-NAACL
2015
-
[24]
Jason Phang, Thibault F \'e vry, and Samuel R. Bowman. 2019. Sentence encoders on stilts: Supplementary training on intermediate labeled-data tasks. CoRR, abs/1811.01088
2019
-
[25]
Martha E. Pollack. 2005. Intelligent technology for an aging population: The use of ai to assist elders with cognitive impairment. AI Magazine, 26:9--24
2005
-
[26]
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. 2018. Improving language understanding by generative Pre-Training
2018
-
[27]
Altaf Rahman and Vincent Ng. 2012. Resolving complex cases of definite pronouns: The winograd schema challenge. In EMNLP , EMNLP-CoNLL '12, pages 777--789, Stroudsburg, PA, USA. Association for Computational Linguistics
2012
-
[29]
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy S. Liang. 2016. Squad: 100, 000+ questions for machine comprehension of text. In EMNLP
2016
-
[30]
Melissa Roemmele, Cosmin Adrian Bejan, and Andrew S. Gordon. 2011. Choice of plausible alternatives: An evaluation of commonsense causal reasoning. In AAAI Spring Symposium: Logical Formalizations of Commonsense Reasoning
2011
-
[31]
Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2019. Winogrande: An adversarial winograd schema challenge at scale. ArXiv, abs/1907.10641
2019
-
[32]
Maarten Sap, Ronan Le Bras , Emily Allaway, Chandra Bhagavatula, Nicholas Lourie, Hannah Rashkin, Brendan Roof, Noah A Smith, and Yejin Choi. 2019. Atomic: An atlas of machine commonsense for if-then reasoning. In AAAI
2019
-
[33]
Shota Sasaki, Sho Takase, Naoya Inoue, Naoaki Okazaki, and Kentaro Inui. 2017. Handling multiword expressions in causality estimation. In IWCS
2017
-
[34]
Sawilowsky
Shlomo S. Sawilowsky. 2009. New effect size rules of thumb. Journal of Modern Applied Statistical Methods, 8(2):597--599
2009
-
[35]
Roy Schwartz, Maarten Sap, Ioannis Konstas, Li Zilles, Yejin Choi, and Noah A Smith. 2017. The effect of different writing tasks on linguistic style: A case study of the ROC story cloze task. In CoNLL
2017
-
[36]
Rishi Kant Sharma, James Allen, Omid Bakhshandeh, and Nasrin Mostafazadeh. 2018. Tackling the story ending biases in the story cloze test. In ACL
2018
-
[37]
Robyn Speer and Catherine Havasi. 2012. Representing general relational knowledge in conceptnet 5. In LREC
2012
-
[38]
Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant. 2019. CommonsenseQA : A question answering challenge targeting commonsense knowledge. In NAACL
2019
-
[39]
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in neural information processing systems, pages 5998--6008
2017
-
[40]
Rowan Zellers, Yonatan Bisk, Ali Farhadi, and Yejin Choi. 2019 a . From recognition to cognition: Visual commonsense reasoning. In CVPR
2019
-
[41]
Rowan Zellers, Yonatan Bisk, Roy Schwartz, and Yejin Choi. 2018. SWAG : A large-scale adversarial dataset for grounded commonsense inference. In EMNLP
2018
-
[42]
Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. 2019 b . Hellaswag: Can a machine really finish your sentence? In ACL
2019
-
[43]
Sheng Zhang, Rachel Rudinger, Kevin Duh, and Benjamin Van Durme. 2017. Ordinal common-sense inference. Transactions of the Association of Computational Linguistics, 5(1):379--395
2017
-
[44]
Zemel, Ruslan R
Yukun Zhu, Ryan Kiros, Richard S. Zemel, Ruslan R. Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler. 2015. Aligning books and movies: Towards story-like visual explanations by watching movies and reading books. 2015 IEEE International Conference on Computer Vi...
2015
-
[45]
2012 , organization=
Goodwin, Travis and Rink, Bryan and Roberts, Kirk and Harabagiu, Sanda M , booktitle=. 2012 , organization=
2012
-
[46]
ArXiv , year=
Winogrande: An Adversarial Winograd Schema Challenge at Scale , author=. ArXiv , year=
-
[47]
Fifteenth International Conference on the Principles of Knowledge Representation and Reasoning , year=
Commonsense causal reasoning between short texts , author=. Fifteenth International Conference on the Principles of Knowledge Representation and Reasoning , year=
-
[48]
EMNLP , year=
QuAC: Question Answering in Context , author=. EMNLP , year=
-
[49]
, author=
Role of theory of mind and executive function in explaining social intelligence: a structural equation modeling approach. , author=. Aging & mental health , year=
-
[50]
Ganaie, MY and Mudasir, Hafiz , journal=
-
[51]
2019 , booktitle=
ATOMIC: An Atlas of Machine Commonsense for If-Then Reasoning , author=. 2019 , booktitle=
2019
-
[52]
Pediatr Res , volume=
Theory of Mind and Neurodevelopmental Disorders of Childhood , author=. Pediatr Res , volume=
-
[53]
Psychometric Properties of the ToM storybooks , author=
Measuring Theory of Mind in Children. Psychometric Properties of the ToM storybooks , author=. Journal of autism and Developmental Disorders , volume=. 2008 , publisher=
2008
-
[54]
ACL , year=
Tackling the Story Ending Biases in The Story Cloze Test , author=. ACL , year=
-
[55]
AI Magazine , year=
Intelligent Technology for an Aging Population: The Use of AI to Assist Elders with Cognitive Impairment , author=. AI Magazine , year=
-
[56]
EMNLP , year=
Evaluating Theory of Mind in Question Answering , author=. EMNLP , year=
-
[57]
1985 , publisher=
Baron-Cohen, Simon and Leslie, Alan M and Frith, Uta , journal=. 1985 , publisher=
1985
-
[58]
theory of mind
Mindreaders: the cognitive basis of" theory of mind" , author=. 2010 , publisher=
2010
-
[59]
arXiv preprint arXiv:1806.03822 , year=
Know What You Don't Know: Unanswerable Questions for SQuAD , author=. arXiv preprint arXiv:1806.03822 , year=
-
[60]
Advances in neural information processing systems , pages=
Attention is all you need , author=. Advances in neural information processing systems , pages=
-
[61]
Rowan Zellers and Yonatan Bisk and Roy Schwartz and Yejin Choi , booktitle=
-
[62]
Proceedings of the Second International Conference on Human Language Technology Research , series =
Schubert, Lenhart , title =. Proceedings of the Second International Conference on Human Language Technology Research , series =. 2002 , location =
2002
-
[63]
ACL , year=
WebChild 2.0 : Fine-Grained Commonsense Knowledge Distillation , author=. ACL , year=
-
[64]
CVPR , year=
From Recognition to Cognition: Visual Commonsense Reasoning , author=. CVPR , year=
-
[65]
WWW , year=
Distilling Task Knowledge from How-To Communities , author=. WWW , year=
-
[66]
WWW , year=
AMIE: association rule mining under incomplete evidence in ontological knowledge bases , author=. WWW , year=
-
[67]
ICLR , year =
Yang, Bishan and Yih, Scott Wen-tau and He, Xiaodong and Gao, Jianfeng and Deng, Li , title=. ICLR , year =
-
[68]
Proceedings of the 2013 Workshop on Automated Knowledge Base Construction , series =
Gordon, Jonathan and Van Durme, Benjamin , title =. Proceedings of the 2013 Workshop on Automated Knowledge Base Construction , series =. 2013 , isbn =. doi:10.1145/2509558.2509563 , acmid =
2013 doi
-
[69]
ACL , year=
Unsupervised Learning of Narrative Event Chains , author=. ACL , year=
-
[70]
EMNLP , year=
How not to evaluate your dialogue system: An empirical study of unsupervised evaluation metrics for dialogue response generation , author=. EMNLP , year=
-
[71]
1977 , publisher=
Scripts, Plans, Goals, and Understanding: An Inquiry Into Human Knowledge Structures , author=. 1977 , publisher=
1977
-
[72]
AAAI Spring Symposium: Logical Formalizations of Commonsense Reasoning , year=
Choice of Plausible Alternatives: An Evaluation of Commonsense Causal Reasoning , author=. AAAI Spring Symposium: Logical Formalizations of Commonsense Reasoning , year=
-
[73]
2017 , publisher=
A Formal Theory of Commonsense Psychology: How People Think People Think , author=. 2017 , publisher=
2017
-
[74]
2018 , booktitle=
Event2Mind: Commonsense Inference on Events, Intents, and Reactions , author=. 2018 , booktitle=
2018
-
[75]
2018 , booktitle=
Modeling Naive Psychology of Characters in Simple Commonsense Stories , author=. 2018 , booktitle=
2018
-
[76]
Did It Happen? The Pragmatic Complexity of Veridicality Assessment
de Marneffe, Marie-Catherine and Manning, Christopher D and Potts, Christopher. Did It Happen? The Pragmatic Complexity of Veridicality Assessment. Comput. Linguist
-
[77]
and Turney, Peter D
Mohammad, Saif M. and Turney, Peter D. , Booktitle =. Crowdsourcing a Word-Emotion Association Lexicon , Volume =
-
[78]
Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=
Obtaining reliable human ratings of valence, arousal, and dominance for 20,000 English words , author=. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=
-
[79]
CoRR , year=
Sentence Encoders on STILTs: Supplementary Training on Intermediate Labeled-data Tasks , author=. CoRR , year=
-
[80]
Journal of Modern Applied Statistical Methods , volume=
New Effect Size Rules of Thumb , author=. Journal of Modern Applied Statistical Methods , volume=
-
[81]
and Neumann, Mark and Iyyer, Mohit and Gardner, Matt and Clark, Christopher and Lee, Kenton and Zettlemoyer, Luke , title=
Peters, Matthew E. and Neumann, Mark and Iyyer, Mohit and Gardner, Matt and Clark, Christopher and Lee, Kenton and Zettlemoyer, Luke , title=. Proc. of NAACL , year=
-
[82]
SSST@EMNLP , year=
On the Properties of Neural Machine Translation: Encoder-Decoder Approaches , author=. SSST@EMNLP , year=
-
[83]
EMNLP , year=
Glove: Global Vectors for Word Representation , author=. EMNLP , year=
-
[84]
2015 IEEE International Conference on Computer Vision (ICCV) , year=
Aligning Books and Movies: Towards Story-Like Visual Explanations by Watching Movies and Reading Books , author=. 2015 IEEE International Conference on Computer Vision (ICCV) , year=
2015
-
[85]
HLT-NAACL , year=
Solving Hard Coreference Problems , author=. HLT-NAACL , year=
-
[86]
Frame-Semantic Parsing , year =
Dipanjan Das and Desai Chen and Andr\'. Frame-Semantic Parsing , year =
-
[87]
RoboNLP@ACL , year=
Are distributional representations ready for the real world? Evaluating word vectors for grounded perceptual meaning , author=. RoboNLP@ACL , year=
-
[88]
Language models are unsupervised multitask learners , author=
-
[89]
Improving Language Understanding by Generative Pre-Training
Radford, Alec and Narasimhan, Karthik and Salimans, Tim and Sutskever, Ilya. Improving Language Understanding by Generative Pre-Training
-
[90]
BERT : Pre-training of Deep Bidirectional Transformers for Language Understanding
Devlin, Jacob and Chang, Ming-Wei and Lee, Kenton and Toutanova, Kristina. BERT : Pre-training of Deep Bidirectional Transformers for Language Understanding. NAACL
-
[91]
NAACL-HLT , year=
Annotation Artifacts in Natural Language Inference Data , author=. NAACL-HLT , year=
-
[92]
The Effect of Different Writing Tasks on Linguistic Style: A Case Study of the ROC Story Cloze Task
Schwartz, Roy and Sap, Maarten and Konstas, Ioannis and Zilles, Li and Choi, Yejin and Smith, Noah A. The Effect of Different Writing Tasks on Linguistic Style: A Case Study of the ROC Story Cloze Task. CoNLL
-
[93]
COLING-ACL , year=
The Berkeley FrameNet Project , author=. COLING-ACL , year=
-
[94]
AAAI Spring Symposium: Logical Formalizations of Commonsense Reasoning , year=
The Winograd Schema Challenge , author=. AAAI Spring Symposium: Logical Formalizations of Commonsense Reasoning , year=
-
[95]
NAACL , year=
VerbNet overview, extensions, mappings and applications , author=. NAACL , year=
-
[96]
EMNLP , year=
The VerbCorner Project: Toward an Empirically-Based Semantic Decomposition of Verbs , author=. EMNLP , year=
-
[97]
TACL , year=
Semantic Proto-Roles , author=. TACL , year=
-
[98]
ACL , year=
HellaSwag: Can a Machine Really Finish Your Sentence? , author=. ACL , year=
-
[99]
EMNLP , year=
Zero-Shot Activity Recognition with Verb Attribute Induction , author=. EMNLP , year=
-
[100]
LREC , year=
Representing General Relational Knowledge in ConceptNet 5 , author=. LREC , year=
-
[101]
MICAI , year=
EventNet: Inferring Temporal Relations Between Commonsense Events , author=. MICAI , year=
-
[102]
Spin: Lexical Semantics, Transitivity, and the Identification of Implicit Sentiment , author=
-
[103]
EMNLP , year=
Universal Decompositional Semantics on Universal Dependencies , author=. EMNLP , year=
-
[104]
ACL , year=
Connotation Frames: A Data-Driven Investigation , author=. ACL , year=
-
[105]
EMNLP , year=
+/-EffectWordNet: Sense-level Lexicon Acquisition for Opinion Inference , author=. EMNLP , year=
-
[106]
AAAI , year=
Acquiring Knowledge of Affective Events from Blogs Using Label Propagation , author=. AAAI , year=
-
[107]
EACL , year=
Acquiring a Dictionary of Emotion-Provoking Events , author=. EACL , year=
-
[108]
EMNLP , year=
A Question Answering Approach for Emotion Cause Extraction , author=. EMNLP , year=
-
[109]
2017 , Eprint =
AllenNLP: A Deep Semantic Natural Language Processing Platform , author=. 2017 , Eprint =
2017
-
[110]
Adam: A Method for Stochastic Optimization
Kingma, Diederik P and Ba, Jimmy. Adam: A Method for Stochastic Optimization. ICLR
-
[111]
EMNLP , year=
Story Comprehension for Predicting What Happens Next , author=. EMNLP , year=
-
[112]
Mostafazadeh, Nasrin and Roth, Michael and Louis, Annie and Chambers, Nathanael and Allen, James , booktitle=
-
[113]
SemEval@NAACL-HLT , year=
SemEval-2015 Task 9: CLIPEval Implicit Polarity of Events , author=. SemEval@NAACL-HLT , year=
2015
-
[114]
Weakly Supervised Induction of Affective Events by Optimizing Semantic Consistency , booktitle=
Ding, Haibo and Riloff, Ellen , year=. Weakly Supervised Induction of Affective Events by Optimizing Semantic Consistency , booktitle=
-
[115]
ACL , year=
Learning Lexico-Functional Patterns for First-Person Affect , author=. ACL , year=
-
[116]
Why is an Event Affective?
Haibo Ding and Tianyu Jiang and Ellen Riloff , booktitle=. Why is an Event Affective?
-
[117]
IWCS , year=
Handling Multiword Expressions in Causality Estimation , author=. IWCS , year=
-
[118]
Resolving Complex Cases of Definite Pronouns: The Winograd Schema Challenge
Rahman, Altaf and Ng, Vincent. Resolving Complex Cases of Definite Pronouns: The Winograd Schema Challenge. EMNLP
-
[119]
NIPS-W , year=
Automatic differentiation in PyTorch , author=. NIPS-W , year=
-
[120]
EMNLP , year=
SQuAD: 100, 000+ Questions for Machine Comprehension of Text , author=. EMNLP , year=
-
[121]
EMNLP , year=
A large annotated corpus for learning natural language inference , author=. EMNLP , year=
-
[122]
NAACL-HLT , year=
A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference , author=. NAACL-HLT , year=
-
[123]
Event Representations for Automated Story Generation with Deep Neural Nets
Martin, Lara J and Ammanabrolu, Prithviraj and Hancock, William and Singh, Shruti and Harrison, Brent and Riedl, Mark O. Event Representations for Automated Story Generation with Deep Neural Nets. AAAI
-
[124]
Augmenting End-to-End Dialog Systems with Commonsense Knowledge
Young, Tom and Cambria, Erik and Chaturvedi, Iti and Huang, Minlie and Zhou, Hao and Biswas, Subham. Augmenting End-to-End Dialog Systems with Commonsense Knowledge. AAAI
-
[125]
Story Ending Generation with Incremental Encoding and Commonsense Knowledge
Guan, Jian and Wang, Yansen and Huang, Minlie. Story Ending Generation with Incremental Encoding and Commonsense Knowledge. arXiv:1808.10113
-
[126]
Knowledgeable Reader: Enhancing Cloze-Style Reading Comprehension with External Commonsense Knowledge
Mihaylov, Todor and Frank, Anette. Knowledgeable Reader: Enhancing Cloze-Style Reading Comprehension with External Commonsense Knowledge. ACL
-
[127]
2013 , publisher=
The development of commonsense psychology , author=. 2013 , publisher=
2013
-
[128]
ACL , year=
A Surprisingly Robust Trick for the Winograd Schema Challenge , author=. ACL , year=
-
[129]
A Corpus and Cloze Evaluation for Deeper Understanding of Commonsense Stories
Mostafazadeh, Nasrin and Chambers, Nathanael and He, Xiaodong and Parikh, Devi and Batra, Dhruv and Vanderwende, Lucy and Kohli, Pushmeet and Allen, James. A Corpus and Cloze Evaluation for Deeper Understanding of Commonsense Stories. NAACL
-
[130]
Machine Common Sense Concept Paper
Gunning, David. Machine Common Sense Concept Paper. arXiv:1810.07528
-
[131]
CoRR , year=
Deep Learning: A Critical Appraisal , author=. CoRR , year=
-
[132]
Commonsense reasoning and commonsense knowledge in artificial intelligence , author=. Commun. ACM , year=
-
[133]
Clarifying the Usage of Structural Models for Commonsense Causal Reasoning , author=
-
[134]
The Behavioral and brain sciences , year=
Building Machines That Learn and Think Like People , author=. The Behavioral and brain sciences , year=
-
[135]
CommonsenseQA : A Question Answering Challenge Targeting Commonsense Knowledge
Talmor, Alon and Herzig, Jonathan and Lourie, Nicholas and Berant, Jonathan. CommonsenseQA : A Question Answering Challenge Targeting Commonsense Knowledge. NAACL
-
[136]
Andrew S Gordon and Reid Swanson , booktitle=. Story
-
[137]
SEM2013 , year=
A Dataset of Syntactic-Ngrams Over Time from a Very Large Corpus of English Books , author=. SEM2013 , year=
-
[138]
, author=
ConceptNet 5.5: An Open Multilingual Graph of General Knowledge. , author=. AAAI , pages=
-
[139]
Proceedings of the Ninth Workshop on Statistical Machine Translation , pages=
A systematic comparison of smoothing techniques for sentence-level bleu , author=. Proceedings of the Ninth Workshop on Statistical Machine Translation , pages=
-
[140]
Transactions of the Association of Computational Linguistics , volume=
Ordinal Common-sense Inference , author=. Transactions of the Association of Computational Linguistics , volume=
-
[141]
arXiv preprint arXiv:1707.08852 , year=
Detecting and explaining causes from text for a time series event , author=. arXiv preprint arXiv:1707.08852 , year=
-
[142]
arXiv preprint arXiv:1708.06022 , year=
Learning to paraphrase for question answering , author=. arXiv preprint arXiv:1708.06022 , year=
-
[143]
Communications of the ACM , volume=
CYC: A large-scale investment in knowledge infrastructure , author=. Communications of the ACM , volume=. 1995 , publisher=
1995
Reviewed May 13, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.