Pith. sign in

The answer is (X)\

6 Pith papers cite this work. Polarity classification is still indexing.

6 Pith papers citing it

citation-role summary

background 1

citation-polarity summary

years

2026 5 2025 1

roles

background 1

polarities

background 1

representative citing papers

Aggregating LLM-Based Weak Verifiers for Spatial Layout Generation

cs.GR · 2026-06-03 · unverdicted · novelty 7.0

Aggregating many LLM-synthesized weak verifiers via weak learning from sparse labels yields stronger verifiers that improve F1 by up to 7X over direct LLM judges on 3D room and 2D poster tasks and boost generation quality by 66.2%.

FUSE: Ensembling Verifiers with Zero Labeled Data

stat.ML · 2026-04-20 · unverdicted · novelty 6.0

FUSE ensembles verifiers unsupervisedly by controlling their conditional dependencies to improve spectral ensembling algorithms, matching or exceeding semi-supervised baselines on benchmarks including GPQA Diamond and Humanity's Last Exam.

LLM-as-a-Verifier: A General-Purpose Verification Framework

cs.AI · 2026-07-06 · conditional · novelty 5.0

Expecting over scoring-token logits yields continuous, scalable verification that improves agent trajectory selection and dense RL rewards across coding, robotics, and medical benchmarks.

citing papers explorer

Showing 6 of 6 citing papers.

  • Aggregating LLM-Based Weak Verifiers for Spatial Layout Generation cs.GR · 2026-06-03 · unverdicted · none · ref 1 · internal anchor

    Aggregating many LLM-synthesized weak verifiers via weak learning from sparse labels yields stronger verifiers that improve F1 by up to 7X over direct LLM judges on 3D room and 2D poster tasks and boost generation quality by 66.2%.

  • Fine-Tuning Small Reasoning Models for Quantum Field Theory cs.LG · 2026-04-21 · unverdicted · none · ref 59 · internal anchor

    Small 7B reasoning models were fine-tuned on synthetic and curated QFT problems using RL and SFT, yielding performance gains, error analysis, and public release of data and traces.

  • FUSE: Ensembling Verifiers with Zero Labeled Data stat.ML · 2026-04-20 · unverdicted · none · ref 9 · internal anchor

    FUSE ensembles verifiers unsupervisedly by controlling their conditional dependencies to improve spectral ensembling algorithms, matching or exceeding semi-supervised baselines on benchmarks including GPQA Diamond and Humanity's Last Exam.

  • LLM-as-a-Verifier: A General-Purpose Verification Framework cs.AI · 2026-07-06 · conditional · none · ref 62 · internal anchor

    Expecting over scoring-token logits yields continuous, scalable verification that improves agent trajectory selection and dense RL rewards across coding, robotics, and medical benchmarks.

  • On the Generalization Gap in Self-Evolving Language Model Reasoning cs.CL · 2026-05-31 · unverdicted · none · ref 28 · internal anchor

    Closed-loop self-evolution on LLMs improves reasoning on Knights and Knaves tasks but plateaus short of oracle-supervised levels, with multi-turn revision nearly matching it for large models.

  • Variation in Verification: Understanding Verification Dynamics in Large Language Models cs.CL · 2025-09-22 · unverdicted · none · ref 1 · internal anchor

    Empirical tests across 12 benchmarks and 14 models show verifiers certify easy problems more reliably, catch errors from weak generators more easily, and correlate with their own solving skill in a difficulty-dependent way.