Pith. sign in

Mechanism design for llm fine- tuning with multiple reward models

6 Pith papers cite this work. Polarity classification is still indexing.

6 Pith papers citing it

years

2026 3 2025 3

verdicts

UNVERDICTED 6

representative citing papers

Incentivizing High-Quality Human Annotations with Golden Questions

cs.GT · 2025-05-25 · unverdicted · novelty 7.0

The paper derives a Θ(1/√(n log n)) hypothesis testing rate under strategic annotator behavior and shows that high-certainty, format-similar golden questions better reveal annotation quality than standard checks.

SynthFix: Adaptive Neuro-Symbolic Code Vulnerability Repair

cs.SE · 2026-04-19 · unverdicted · novelty 5.0

A router mixes supervised and reward fine-tuning with compiler/security feedback so small code LLMs produce more functionally correct and security-cleared vulnerability patches on three repair benchmarks.

citing papers explorer

Showing 6 of 6 citing papers.