Pith. sign in

REVIEW 1 cited by

Personalized Jargon Identification for Enhanced Interdisciplinary Communication

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.09481 v1 pith:5EWSAHMD submitted 2023-11-16 cs.CL

classification cs.CL
keywords jargonfamiliarityidentificationmethodsresearchersdatafeaturesindividual
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Scientific jargon can impede researchers when they read materials from other domains. Current methods of jargon identification mainly use corpus-level familiarity indicators (e.g., Simple Wikipedia represents plain language). However, researchers' familiarity of a term can vary greatly based on their own background. We collect a dataset of over 10K term familiarity annotations from 11 computer science researchers for terms drawn from 100 paper abstracts. Analysis of this data reveals that jargon familiarity and information needs vary widely across annotators, even within the same sub-domain (e.g., NLP). We investigate features representing individual, sub-domain, and domain knowledge to predict individual jargon familiarity. We compare supervised and prompt-based approaches, finding that prompt-based methods including personal publications yields the highest accuracy, though zero-shot prompting provides a strong baseline. This research offers insight into features and methods to integrate personal data into scientific jargon identification.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 3 citations worldwide. Full citation record

  1. Mapping Scientific Literature with Large Language Models and Topic Modeling

    cs.DL 2025-10 conditional novelty 4.0 of 10

    An LLM-based pipeline turns 20 years of PNAS engineering abstracts into sixteen interpretable topics and maps cross-topic links from full text, claiming to rediscover the journal's own dual-classification structure.

Pith tools