Pith. sign in

REVIEW 3 cited by

Deriving Machine Attention from Human Rationales

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1808.09367 v1 pith:LSQ5SI6H submitted 2018-08-28 cs.CL

classification cs.CL
keywords attentionrationalesdomainshypothesislow-resourcemappingacrossamounts
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Attention-based models are successful when trained on large amounts of data. In this paper, we demonstrate that even in the low-resource scenario, attention can be learned effectively. To this end, we start with discrete human-annotated rationales and map them into continuous attention. Our central hypothesis is that this mapping is general across domains, and thus can be transferred from resource-rich domains to low-resource ones. Our model jointly learns a domain-invariant representation and induces the desired mapping between rationales and attention. Our empirical results validate this hypothesis and show that our approach delivers significant gains over state-of-the-art baselines, yielding over 15% average error reduction on benchmark datasets.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Efficient Text-Attributed Graph Learning through Selective Annotation and Graph Alignment

    cs.LG 2025-06 conditional novelty 6.0 of 10

    GAGA matches or exceeds state-of-the-art accuracy on several text-attributed graph benchmarks while requiring large language model annotations for only 1% of nodes or edges.

  2. Can human clinical rationales improve the performance and explainability of clinical text classification models?

    cs.CL 2025-07 conditional novelty 5.0 of 10

    Adding 96,679 human rationale highlights improves cancer-site classification less than adding the same number of full pathology reports, and the explainability gain is small.

  3. SFT-GO: Supervised Fine-Tuning with Group Optimization for Large Language Models

    cs.LG 2025-06 conditional novelty 4.0 of 10

    SFT-GO retrains LLMs by focusing on the worst-performing group of important or unimportant tokens, yielding modest average benchmark improvements over standard supervised fine-tuning.

Pith tools