Pith. sign in

REVIEW 2 cited by

The Capabilities and Limitations of Weak-to-Strong Generalization: Generalization and Calibration

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.01458 v3 pith:WABYMTGS submitted 2025-02-03 cs.LG stat.ML

classification cs.LGstat.ML
keywords generalizationstrongmodelweakboundserrorlimitationsapproach
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Weak-to-strong generalization, where weakly supervised strong models outperform their weaker teachers, offers a promising approach to aligning superhuman models with human values. To deepen the understanding of this approach, we provide theoretical insights into its capabilities and limitations. First, in the classification setting, we establish upper and lower generalization error bounds for the strong model, identifying the primary limitations as stemming from the weak model's generalization error and the optimization objective itself. Additionally, we derive lower and upper bounds on the calibration error of the strong model. These theoretical bounds reveal two critical insights: (1) the weak model should demonstrate strong generalization performance and maintain well-calibrated predictions, and (2) the strong model's training process must strike a careful balance, as excessive optimization could undermine its generalization capability by over-relying on the weak supervision signals. Finally, in the regression setting, we extend the work of Charikar et al. (2024) to a loss function based on Kullback-Leibler (KL) divergence, offering guarantees that the strong student can outperform its weak teacher by at least the magnitude of their disagreement. We conduct sufficient experiments to validate our theory.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. On Weak-to-Strong Generalization and f-Divergence

    cs.LG 2025-06 conditional novelty 6.0 of 10

    Replacing cross-entropy with f-divergence losses in weak-to-strong generalization gives modest accuracy gains and improved label-noise tolerance, though the paper's theoretical equivalence result is constructed after ...

  2. Contrastive Weak-to-strong Generalization

    cs.CL 2025-10 conditional novelty 5.0 of 10

    Contrastive decoding between pre- and post-alignment weak models generates better supervision samples, improving weak-to-strong generalization on AlpacaEval2 and Arena-Hard.

Pith tools