TRIE benchmarks stochastic PDE surrogates on two chaotic SPDEs, finding generative models best match long-term statistics and uncertainty while latent versions cut inference time by 12x.
Scheduled sampling for sequence prediction with recurrent neural networks.Advances in neural information processing systems, 28
9 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
years
2026 9verdicts
UNVERDICTED 9roles
background 2polarities
background 2representative citing papers
GTF-DEER augments the DEER framework with Generalized Teacher Forcing to allow effective parallel training of nonlinear recurrent models on extremely long sequences, improving dynamical systems reconstruction for data with long time scales.
SDFlow learns a global transport map via similarity-driven flow matching in VQ latent space, using low-rank manifold decomposition and a categorical posterior to handle discreteness, yielding SOTA long-horizon performance and inference speedups.
AsymTalker uses temporal reference encoding and asymmetric knowledge distillation to produce identity-consistent talking head videos up to 600 seconds long at 66 FPS.
AID-VAR attaches an adversarial discriminator and lightweight guidance injector to frozen VAR backbones to diagnose and correct fidelity gaps across scales, reporting 16% FID gains with 3% added parameters.
A generative navigation world model that uses sparse anchored rollout with epipolar constraints to reduce perceptual and geometric drift.
A state distribution view of post-training shows that on-policy supervision from the learner itself can outperform fixed-dataset SFT and preserve retention better than aggressive supervised updates.
f-OPD decomposes on-policy distillation drift into rollout and supervision components, then applies a sample-level freshness score to adaptively limit stale data influence and stabilize long-horizon agent training.
Prune-OPD detects prefix drift via top-k overlap and dynamically prunes unreliable teacher rewards in OPD, cutting training time 37.6-68% on AMC/AIME/HMMT while preserving performance.
citing papers explorer
-
TRIE: An Evaluation Framework for Stochastic PDE Surrogates
TRIE benchmarks stochastic PDE surrogates on two chaotic SPDEs, finding generative models best match long-term statistics and uncertainty while latent versions cut inference time by 12x.
-
Parallel-in-Time Training of Recurrent Neural Networks for Dynamical Systems Reconstruction
GTF-DEER augments the DEER framework with Generalized Teacher Forcing to allow effective parallel training of nonlinear recurrent models on extremely long sequences, improving dynamical systems reconstruction for data with long time scales.
-
SDFlow: Similarity-Driven Flow Matching for Time Series Generation
SDFlow learns a global transport map via similarity-driven flow matching in VQ latent space, using low-rank manifold decomposition and a categorical posterior to handle discreteness, yielding SOTA long-horizon performance and inference speedups.
-
AsymTalker: Identity-Consistent Long-Term Talking Head Generation via Asymmetric Distillation
AsymTalker uses temporal reference encoding and asymmetric knowledge distillation to produce identity-consistent talking head videos up to 600 seconds long at 66 FPS.
-
Adversarial Error Correction for Visual Autoregressive Generation
AID-VAR attaches an adversarial discriminator and lightweight guidance injector to frozen VAR backbones to diagnose and correct fidelity gaps across scales, reporting 16% FID gains with 3% added parameters.
-
Drift-Resistant Navigation World Model with Anchored Epipolar Guidance
A generative navigation world model that uses sparse anchored rollout with epipolar constraints to reduce perceptual and geometric drift.
-
Post-Training is About States, Not Tokens: A State Distribution View of SFT, RL, and On-Policy Distillation
A state distribution view of post-training shows that on-policy supervision from the learner itself can outperform fixed-dataset SFT and preserve retention better than aggressive supervised updates.
-
$\boldsymbol{f}$-OPD: Stabilizing Long-Horizon On-Policy Distillation with Freshness-Aware Control
f-OPD decomposes on-policy distillation drift into rollout and supervision components, then applies a sample-level freshness score to adaptively limit stale data influence and stabilize long-horizon agent training.
-
Prune-OPD: Efficient and Reliable On-Policy Distillation for Long-Horizon Reasoning
Prune-OPD detects prefix drift via top-k overlap and dynamically prunes unreliable teacher rewards in OPD, cutting training time 37.6-68% on AMC/AIME/HMMT while preserving performance.