Pith. sign in

super hub Mixed citations

A coefficient of agreement for nominal scales.Educational and Psychological Measurement, 20(1):37–46

Mixed citation behavior. Most common role is background (33%).

34 Pith papers citing it
30.9k external citations · Crossref
Background 33% of classified citations

hub tools

citation-role summary

background 3 method 2 dataset 1

citation-polarity summary

authors

co-cited works

representative citing papers

Causal state binding predicts action control in language agents

cs.AI · 2026-05-10 · unverdicted · novelty 7.0 · 3 refs

Causal state binding is introduced as a framework that predicts action control in language agents, validated across large benchmarks and SWE-bench Lite where adding the measure raised issue-to-file hit@3 AUC from 0.873 to 0.935.

ProactBench: Beyond What The User Asked For

cs.LG · 2026-05-09 · unverdicted · novelty 7.0

ProactBench measures LLM conversational proactivity in three phases using 198 multi-agent dialogues and finds recovery behavior hard to predict from existing benchmarks.

EO-Gym: A Multimodal, Interactive Environment for Earth Observation Agents

cs.AI · 2026-05-02 · unverdicted · novelty 7.0

EO-Gym supplies an executable multimodal environment and 9k-trajectory benchmark that turns Earth Observation into a tool-using, multi-step reasoning task, revealing that current VLMs struggle on temporal and cross-sensor workflows while fine-tuning lifts Pass@3 from 0.49 to 0.74.

Help! Need Advice on Identifying Advice

cs.CL · 2020-10-06 · unverdicted · novelty 6.0

Introduces a new English dataset from r/AskParents and r/needadvice annotated for advice sentences plus preliminary models showing pre-trained LMs outperform rule-based systems but the task remains challenging.

LLMs for automatic annotation of Mandarin narrative transcripts

cs.CL · 2026-05-17 · unverdicted · novelty 5.0

LLMs achieve near-human agreement (k=0.794 vs human-human k=0.872) on annotating Mandarin narrative macrostructure with the MAIN framework, reducing time by 65 percent but showing lower reliability on young adult narratives with greater lexical variation.

citing papers explorer

Showing 34 of 34 citing papers.