Pith. sign in

REVIEW 3 cited by

Character-level and Multi-channel Convolutional Neural Networks for Large-scale Authorship Attribution

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1609.06686 v1 pith:BGSQWJLN submitted 2016-09-21 cs.CL cs.LG

classification cs.CLcs.LG
keywords character-levelcnnsapproachesattributionauthorshipstate-of-the-artauthorconvolutional
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Convolutional neural networks (CNNs) have demonstrated superior capability for extracting information from raw signals in computer vision. Recently, character-level and multi-channel CNNs have exhibited excellent performance for sentence classification tasks. We apply CNNs to large-scale authorship attribution, which aims to determine an unknown text's author among many candidate authors, motivated by their ability to process character-level signals and to differentiate between a large number of classes, while making fast predictions in comparison to state-of-the-art approaches. We extensively evaluate CNN-based approaches that leverage word and character channels and compare them against state-of-the-art methods for a large range of author numbers, shedding new light on traditional approaches. We show that character-level CNNs outperform the state-of-the-art on four out of five datasets in different domains. Additionally, we present the first application of authorship attribution to reddit.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. AIDBench: A benchmark for evaluating the authorship identification capability of large language models

    cs.CL 2024-11 conditional novelty 5.0 of 10

    A new benchmark shows GPT-4 and several other LLMs can attribute anonymous texts to their authors at rates well above random chance, though performance drops sharply in harder cross-topic settings.

  2. Similarity Learning for Authorship Verification in Social Media

    cs.CL 2019-08 conditional novelty 5.0 of 10

    A hierarchical recurrent Siamese network with a two-threshold contrastive loss raises authorship verification accuracy on a social-media benchmark from about 71% to 85.3%.

  3. Learning Text Styles: A Study on Transfer, Attribution, and Verification

    cs.CL 2025-07 conditional novelty 3.0 of 10

    A thesis compiles published work claiming that lightweight adapters, contrastive disentanglement, and instruction tuning improve text style transfer, authorship attribution, and authorship verification.

Pith tools