Pith. sign in

hub

Efficient online data mixing for language model pre-training

11 Pith papers cite this work. Polarity classification is still indexing.

11 Pith papers citing it

hub tools

citation-role summary

background 1

citation-polarity summary

fields

cs.LG 9 cs.CL 2

roles

background 1

polarities

background 1

representative citing papers

Smooth Scaling Laws Hide Stepwise Token Learning

cs.CL · 2026-06-29 · conditional · novelty 7.0

Power-law LLM scaling laws are largely the aggregate of stepwise token learning events whose heavy-tailed learning-time spectrum reconstructs loss derivatives along step, data, and model axes.

Explaining Data Mixing Scaling Laws

cs.LG · 2026-06-06 · conditional · novelty 7.0

Under a shared-head/disjoint-tail assumption, multi-domain loss decomposes into a capacity-competition term c_i x_i^*(h)^{-b_i} plus a per-domain noise term A_i(Dh_i)^{-a_i}, and the fitted law extrapolates optimal mixtures to unseen scales.

DRIFT: Refining Instruction Data via On-Policy Data Attribution

cs.LG · 2026-06-16 · unverdicted · novelty 6.0

DRIFT applies on-policy influence functions with signed weighting and debiasing to attribute and refine SFT data, raising performance on 7B instruction and reasoning models over prior curation methods.

citing papers explorer

Showing 11 of 11 citing papers.