SDGD uses cost-conditioned classifier-free guidance plus reward guidance with feasible trajectory relabeling to generate safe high-reward trajectories that adapt to changing safety budgets in offline RL.
Title resolution pending
3 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
years
2026 3verdicts
UNVERDICTED 3roles
method 1polarities
use method 1representative citing papers
Auto-CoT automatically generates and filters reasoning-enhanced demonstrations to improve in-context learning accuracy on complex reasoning tasks.
Non-linear transformers enable cross-domain generalization in in-context RL by representing value functions from different domains with shared weights inside a shared RKHS.
citing papers explorer
-
Decoupled Guidance Diffusion for Adaptive Offline Safe Reinforcement Learning
SDGD uses cost-conditioned classifier-free guidance plus reward guidance with feasible trajectory relabeling to generate safe high-reward trajectories that adapt to changing safety budgets in offline RL.
-
ACIL: Auto Chain of Thoughts for In-Context Learning
Auto-CoT automatically generates and filters reasoning-enhanced demonstrations to improve in-context learning accuracy on complex reasoning tasks.
-
One for All: A Non-Linear Transformer can Enable Cross-Domain Generalization for In-Context Reinforcement Learning
Non-linear transformers enable cross-domain generalization in in-context RL by representing value functions from different domains with shared weights inside a shared RKHS.