← back to paper
arxiv: 2608.04448 · 2 revisions
When does training on downscaled images yield the same gradients?