Pith. sign in
Pith Number

pith:GZAASV5G

pith:2019:GZAASV5GL4HRU3KEYQRNEQUK5W
not attested not anchored not stored refs resolved

SocialIQA: Commonsense Reasoning about Social Interactions

Derek Chen, Hannah Rashkin, Maarten Sap, Ronan LeBras, Yejin Choi

Social IQa is a 38,000-question benchmark that exposes a greater than 20 percent performance gap between humans and pretrained language models on social commonsense reasoning.

arxiv:1904.09728 v3 · 2019-04-22 · cs.CL

Add to your LaTeX paper
\usepackage{pith}
\pithnumber{GZAASV5GL4HRU3KEYQRNEQUK5W}

Prints a linked badge after your title and injects PDF metadata. Compiles on arXiv. Learn more · Embed verified badge

Record completeness

1 Bitcoin timestamp
2 Internet Archive
3 Author claim open · sign in to claim
4 Citations open
5 Replications open
Portable graph bundle live · download bundle · merged state
The bundle contains the canonical record plus signed events. A mirror can host it anywhere and recompute the same current state with the deterministic merge algorithm.

Claims

C1strongest claim

Our benchmark is challenging for existing question-answering models based on pretrained language models, compared to human performance (>20% gap). Notably, we further establish Social IQa as a resource for transfer learning of commonsense knowledge, achieving state-of-the-art performance on multiple commonsense reasoning tasks (Winograd Schemas, COPA).

C2weakest assumption

That the crowdsourced questions and answers, even with the new framework to mitigate stylistic artifacts, accurately capture genuine social commonsense without introducing new biases or failing to probe true emotional intelligence.

C3one line summary

SocialIQA is the first large-scale benchmark with 38k crowdsourced questions testing commonsense about social interactions, where pretrained language models trail humans by over 20% but transfer to improve performance on Winograd Schemas and COPA.

References

140 extracted · 140 resolved · 4 Pith anchors

[1] theory of mind 2010
[2] Simon Baron-Cohen, Alan M Leslie, and Uta Frith. 1985. Does the Autistic Child have a ``Theory of Mind''? Cognition, 21(1):37--46 1985
[3] Ernest Davis and Gary Marcus. 2015. Commonsense reasoning and commonsense knowledge in artificial intelligence. Commun. ACM, 58:92--103 2015
[4] Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT : Pre-training of deep bidirectional transformers for language understanding. In NAACL 2019
[5] Espinosa and Henry Lieberman 2005

Formal links

2 machine-checked theorem links

Cited by

58 papers in Pith

Receipt and verification
First computed 2026-07-05T00:02:58.276093Z
Builder pith-number-builder-2026-05-17-v1
Signature Pith Ed25519 (pith-v1-2026-05) · public key
Schema pith-number/v1.0

Canonical hash

36400957a65f0f1a6d44c422d2428aed8e7bf85b4b5e6a05eeed01cfa54ebd89

Aliases

arxiv: 1904.09728 · arxiv_version: 1904.09728v3 · doi: 10.48550/arxiv.1904.09728 · pith_short_12: GZAASV5GL4HR · pith_short_16: GZAASV5GL4HRU3KE · pith_short_8: GZAASV5G
Agent API
Verify this Pith Number yourself
curl -sH 'Accept: application/ld+json' https://pith.science/pith/GZAASV5GL4HRU3KEYQRNEQUK5W \
  | jq -c '.canonical_record' \
  | python3 -c "import sys,json,hashlib; b=json.dumps(json.loads(sys.stdin.read()), sort_keys=True, separators=(',',':'), ensure_ascii=False).encode(); print(hashlib.sha256(b).hexdigest())"
# expect: 36400957a65f0f1a6d44c422d2428aed8e7bf85b4b5e6a05eeed01cfa54ebd89
Canonical record JSON
{
  "metadata": {
    "abstract_canon_sha256": "247cd33bc4d80ff2cecf07b29b452b3ca136b7853833216a493a2d416f48077a",
    "cross_cats_sorted": [],
    "license": "http://arxiv.org/licenses/nonexclusive-distrib/1.0/",
    "primary_cat": "cs.CL",
    "submitted_at": "2019-04-22T05:36:37Z",
    "title_canon_sha256": "18825dae92aef04eb6bd2f54934a367526413e883fc2b6c7041fb2380b9b7020"
  },
  "schema_version": "1.0",
  "source": {
    "id": "1904.09728",
    "kind": "arxiv",
    "version": 3
  }
}