Pith. sign in

REVIEW 3 cited by

Super(ficial)-alignment: Strong Models May Deceive Weak Models in Weak-to-Strong Generalization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.11431 v3 pith:PMKMQPGC submitted 2024-06-17 cs.CL cs.AI

classification cs.CLcs.AI
keywords modelsweakstrongalignmentdeceptionweak-to-strongphenomenonareas
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Superalignment, where humans act as weak supervisors for superhuman models, has become a crucial problem with the rapid development of Large Language Models (LLMs). Recent work has preliminarily studied this problem by using weak models to supervise strong models, and discovered that weakly supervised strong students can consistently outperform weak teachers towards the alignment target, leading to a weak-to-strong generalization phenomenon. However, we are concerned that behind such a promising phenomenon, whether there exists an issue of weak-to-strong deception, where strong models deceive weak models by exhibiting well-aligned in areas known to weak models but producing misaligned behaviors in cases weak models do not know. We take an initial step towards exploring this security issue in a specific but realistic multi-objective alignment case, where there may be some alignment targets conflicting with each other (e.g., helpfulness v.s. harmlessness). We aim to explore whether, in such cases, strong models might deliberately make mistakes in areas known to them but unknown to weak models within one alignment dimension, in exchange for a higher reward in another dimension. Through extensive experiments in both the reward modeling and preference optimization scenarios, we find: (1) The weak-to-strong deception phenomenon exists across all settings. (2) The deception intensifies as the capability gap between weak and strong models increases. (3) Bootstrapping with an intermediate model can mitigate the deception to some extent, though its effectiveness remains limited. Our work highlights the urgent need to pay more attention to the true reliability of superalignment.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DeepCritic: Deliberate Critique with Large Language Models

    cs.CL 2025-05 conditional novelty 6.0 of 10

    A two-stage SFT and RL pipeline turns a 7B instruct model into a deliberate math critic that outperforms GPT-4o and same-size R1-distill models at locating the first erroneous step.

  2. Super Co-alignment of Human and AI for Sustainable Symbiotic Society

    cs.AI 2025-04 unverdicted novelty 4.0 of 10

    The authors propose 'Super Co-alignment', in which humans and superintelligent AI iteratively co-evolve shared values through external oversight and intrinsic empathy-based alignment.

  3. The Road to Artificial SuperIntelligence: A Comprehensive Survey of Superalignment

    cs.LG 2024-12 conditional novelty 3.0 of 10

    A survey of scalable oversight for superalignment, reviewing weak-to-strong generalization, debate, RLAIF, and sandwiching, and concluding that current methods are not yet sufficient for superintelligent AI.

Pith tools