Pith. sign in

REVIEW 1 cited by

Conditional Language Policy: A General Framework for Steerable Multi-Objective Finetuning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.15762 v2 pith:73527J7J submitted 2024-07-22 cs.LG cs.AIcs.CL

classification cs.LGcs.AIcs.CL
keywords languagefinetuningmodelsobjectivessteerablemultipleconditionalconflicting
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Reward-based finetuning is crucial for aligning language policies with intended behaviors (e.g., creativity and safety). A key challenge is to develop steerable language models that trade-off multiple (conflicting) objectives in a flexible and efficient manner. This paper presents Conditional Language Policy (CLP), a general framework for finetuning language models on multiple objectives. Building on techniques from multi-task training and parameter-efficient finetuning, CLP learn steerable models that effectively trade-off conflicting objectives at inference time. Notably, this does not require training or maintaining multiple models to achieve different trade-offs between the objectives. Through extensive experiments and ablations on two summarization datasets, we show that CLP learns steerable language models that outperform and Pareto-dominate the existing approaches for multi-objective finetuning.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Simple Optimizers for Convex Aligned Multi-Objective Optimization

    cs.LG 2025-09 reject novelty 6.0 of 10

    Convex AMOO is analyzed under Lipschitz and smooth assumptions with a maximum-gap metric, giving simple gradient methods with rates independent of the number of objectives, plus a flawed equal-weights lower bound.

Pith tools