EvoPool evolves pools of programmatic annotators that outperform LLM annotation by 0.141 average macro-F1 on 7 of 8 specialized tasks while running thousands of times faster.
Get Your Vitamin
8 Pith papers cite this work, alongside 6 external citations. Polarity classification is still indexing.
representative citing papers
ClaimPKG generates entity-constrained pseudo-subgraphs from claims to retrieve KG evidence and then reasons with a general LLM, reporting state-of-the-art FactKG accuracy.
Synthetic context-utilisation benchmarks overstate phenomena like knowledge conflicts and context repulsion, while a new real-world dataset (DRUID) shows weaker, source-driven correlations with model behavior.
Across four fact-checking datasets, no automated fact-checking system generalizes reliably across domains, and retrieval quality, not verifier reasoning, is the main performance bottleneck.
Combining residual-stream motion with a coarse region code and a fine direction readout improves zero-shot selection of correct LLM answers across reasoning and factual benchmarks.
Final-token penultimate-block states of three 7–8B instruct models linearly decode Favors/Challenges/Wrong-Target for identical evidence under different causal targets, recovering 18–21 of 49 pairs.
Decontextualized, hyperlink-anchored evidence premises from fact-check articles improve retrieval and verdict prediction, released as the PrimeFacts resource.
A survey that organizes adversarial attacks on automated fact-checking into claim, evidence, and claim-evidence pair categories, and reports that current defenses address 13 of the 53 surveyed attacks.
citing papers explorer
No citing papers match the current filters.