Pith. sign in

REVIEW 2 cited by

LLM-Measure: Generating Valid, Consistent, and Reproducible Text-Based Measures for Social Science Research

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.12722 v1 pith:Q5ZMGH2O submitted 2024-09-19 cs.CL

classification cs.CL
keywords conceptmeasuresresearchconsistentmethodreproduciblesciencesocial
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The increasing use of text as data in social science research necessitates the development of valid, consistent, reproducible, and efficient methods for generating text-based concept measures. This paper presents a novel method that leverages the internal hidden states of large language models (LLMs) to generate these concept measures. Specifically, the proposed method learns a concept vector that captures how the LLM internally represents the target concept, then estimates the concept value for text data by projecting the text's LLM hidden states onto the concept vector. Three replication studies demonstrate the method's effectiveness in producing highly valid, consistent, and reproducible text-based measures across various social science research contexts, highlighting its potential as a valuable tool for the research community.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. From Research Questions to Columns: Operationalization-Aware Data Discovery

    cs.DB 2026-08 conditional novelty 6.0 of 10

    Finding columns that operationalize a broad research question is harder than standard column search or schema linking; on the new OADD-Bench, the best agent achieves only 0.465 recall at a 5x budget.

  2. From Text to Discovery: How Large Language Models Are Reshaping Research Across Scientific and Humanistic Disciplines

    cs.DL 2026-06 unverdicted novelty 3.0 of 10

    LLMs accelerate research workflows from idea generation to writing but introduce challenges like hallucination, bias, opacity, and ten systemic risks requiring new governance frameworks.

Pith tools