Pith. sign in

REVIEW 1 cited by

UNIDECOR: A Unified Deception Corpus for Cross-Corpus Deception Detection

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.02827 v2 pith:LDMRADZS submitted 2023-06-05 cs.CL

UNIDECOR: A Unified Deception Corpus for Cross-Corpus Deception Detection

classification cs.CL
keywords deceptioncorpusdatasetsunidecorunifiedacrosscross-corpusdetection
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Verbal deception has been studied in psychology, forensics, and computational linguistics for a variety of reasons, like understanding behaviour patterns, identifying false testimonies, and detecting deception in online communication. Varying motivations across research fields lead to differences in the domain choices to study and in the conceptualization of deception, making it hard to compare models and build robust deception detection systems for a given language. With this paper, we improve this situation by surveying available English deception datasets which include domains like social media reviews, court testimonials, opinion statements on specific topics, and deceptive dialogues from online strategy games. We consolidate these datasets into a single unified corpus. Based on this resource, we conduct a correlation analysis of linguistic cues of deception across datasets to understand the differences and perform cross-corpus modeling experiments which show that a cross-domain generalization is challenging to achieve. The unified deception corpus (UNIDECOR) can be obtained from https://www.ims.uni-stuttgart.de/data/unidecor.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Semantics of Subterfuge: Benchmarking Legal Deception Detection Against General-domain State-of-the-Art

    cs.CL 2026-07 conditional novelty 5.0

    On deception detection benchmarks, fine-tuned transformers beat LLMs on data-rich datasets, few-shot GPT-4o wins the small legal corpus, and chain-of-thought prompting frequently reduces F1.