Pith. sign in

REVIEW 4 major objections 2 minor 52 references

Once the margin is large enough, classification becomes learnable in every metric space using only the triangle inequality.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-15 13:23 UTC pith:TLAJRXDS

load-bearing objection Clean, high-value learning-theory claims on margin without linear structure—but the package only gives the abstract, so the sharp threshold and non-embedding result are still unchecked. the 4 major comments →

arxiv 2603.07221 v2 pith:TLAJRXDS submitted 2026-03-07 cs.LG math.FA

Margin in Abstract Spaces

classification cs.LG math.FA MSC 68Q3268T05
keywords margin-based learningmetric spacestriangle inequalitygeneralizationBanach spacesdistance functionslearnability thresholdsover-parameterized learning
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Margin methods such as linear and kernel classifiers are classical examples where generalization does not grow with the number of parameters. This paper asks what minimal geometry makes that possible. It starts with a simple rule in an arbitrary metric space: pick a center and label points as positive if they are closer than r and negative if farther than R. Whenever R is more than three times r, that class is learnable in every metric space; the triangle inequality alone is enough. The main theorem lifts the same idea to concepts built from bounded linear combinations of distance functions and proves a sharp universal threshold: above a fixed constant the class is learnable everywhere, while below it there exist metric spaces in which it fails completely. The paper then shows that this phenomenon cannot always be reduced to ordinary linear margin classification after embedding into a Banach space.

Core claim

There exists a universal constant such that, for hypothesis classes defined by bounded linear combinations of distance functions, any margin larger than that constant yields learnability in every metric space, while any smaller margin admits metric spaces in which the same class is not learnable. Separately, there exist margin-learnable classes that cannot be realized by linear margin classification in any Banach space.

What carries the argument

The sharp margin threshold for linear combinations of distance functions: a single universal constant that separates “learnable in every metric space” from “fails in some metric spaces,” relying only on the triangle inequality.

Load-bearing premise

That the PAC-style notion of learnability used here is completely controlled by the margin relative to the triangle inequality, with no extra measurability or sampling restrictions that would limit the claim of working in every metric space.

What would settle it

Either exhibit a metric space in which the distance-combination class remains unlearnable for every margin larger than the claimed constant, or construct an embedding of the paper’s counter-example class into a Banach space where linear classification with positive margin is learnable.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Generalization guarantees that are independent of ambient dimension can arise from pure metric geometry once the margin is large enough.
  • Kernel-style embeddings into Banach spaces are not the only possible explanation for margin-based learnability.
  • Designing learning algorithms for abstract metric data can safely ignore linear structure provided the effective margin clears the universal threshold.
  • Below the threshold, hardness is metric-dependent: some spaces remain learnable while others become impossible.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The same threshold phenomenon may appear for other geometric primitives (e.g., geodesics or higher-order distance polynomials) once a suitable notion of margin is defined.
  • The negative embedding result suggests a hierarchy of “margin geometries” strictly richer than Banach-space linear margins.
  • Practical algorithms that only enforce large geometric margins on pairwise distances might inherit dimension-free sample bounds even when no kernel is available.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 2 minor

Summary. The abstract claims that margin-based learnability needs only the triangle inequality: (i) center-based concepts that label points by distance thresholds r vs R are PAC-learnable in every metric space whenever R>3r; (ii) for concepts given by bounded linear combinations of distance functions there is a universal margin constant C such that margin > C implies learnability in every metric space, while margin < C admits metric spaces where the class is not learnable; (iii) there exists a margin-learnable class that cannot be realized by embedding into any Banach space in which linear margin classification is learnable. The supplied full-text body, however, is an unrelated computer-vision paper (VINO, video SSL / masked distillation on Walking Tours), not a learning-theory manuscript on margins in metric spaces. Consequently no definitions of the hypothesis classes, no sample-complexity statements, no proofs, and no constructions of the claimed counterexample metric spaces or non-embeddable class are available for review.

Significance. If the abstract claims were established with full proofs, the work would be a substantial foundational contribution: it would isolate the triangle inequality as the minimal structure behind dimension-free margin generalization, give a sharp universal threshold for distance-combination classes, and separate margin learnability from Banach/kernel embeddings. Those results would be of clear interest to statistical learning theory and to the theory of over-parameterized models. As submitted, however, the mathematical content is absent, so the significance cannot be credited.

major comments (4)
  1. Manuscript integrity: the title/abstract (Margin in Abstract Spaces, arXiv:2603.07221, cs.LG) do not match the full text, which is the unrelated CV paper VINO (arXiv:2603.07222). No theorems, definitions, or proofs of the claimed margin results appear. The central claims are therefore unreviewable from the provided document.
  2. Abstract claim of a universal constant C for bounded linear combinations of distance functions: without a formal definition of the concept class, the precise PAC model (distribution-free vs fixed measure; finite vs infinite sample complexity), and the constructions of the bad metric spaces below C, the sharp-threshold statement cannot be checked. In particular it is unclear whether “every metric space” is unrestricted or tacitly assumes separability/completeness/Borel measurability.
  3. Abstract claim R>3r for center-based (r,R) concepts: the factor 3 is presented as relying only on the triangle inequality, but no proof or even a sketch is present in the supplied text, so correctness and tightness cannot be assessed.
  4. Negative embedding result (margin-learnable class not reducible to linear margin classification in any Banach space): no construction or non-embeddability argument is supplied, so the claim that margin learnability is strictly more general than kernel/Banach methods remains unsupported.
minor comments (2)
  1. Once the correct learning-theory manuscript is supplied, the abstract should state the precise PAC model, the form of the sample-complexity bound (or at least that it is finite and independent of ambient cardinality), and an explicit numerical value or characterization of the universal constant C if it is known.
  2. The supplied body (VINO) has its own presentation issues (garbled table entries, OCR-like character corruption in figures/captions) but those are irrelevant to the claimed paper and should not be treated as revisions of 2603.07221.

Circularity Check

0 steps flagged

No circularity identifiable: abstract states existence/threshold learnability claims; supplied full text is an unrelated paper, so no derivation chain can be reduced to its inputs.

full rationale

The target paper (Margin in Abstract Spaces, 2603.07221) is represented only by its abstract. That abstract asserts (i) R>3r ball-margin concepts are learnable in every metric space via the triangle inequality, (ii) a sharp universal margin threshold for bounded linear combinations of distance functions, and (iii) a negative embedding result into Banach spaces. These are pure existence/threshold statements about PAC-style learnability; they do not fit parameters to data and then re-label the fit as a prediction, nor do they define a quantity in terms of itself. The CACHEABLE full-manuscript block is an entirely different paper (VINO, 2603.07222, a self-supervised video SSL method). Consequently no equations, proofs, or self-citation chains from 2603.07221 are available to inspect. Under the hard rule that circularity may be claimed only when a specific reduction can be quoted, the only honest finding is no significant circularity. Residual dependence on standard PAC definitions is ordinary and not circular.

Axiom & Free-Parameter Ledger

0 free parameters · 3 axioms · 2 invented entities

Abstract-only review. Load-bearing background is standard metric-space and PAC-learning structure; the paper’s contribution is the threshold and non-embedding theorems built on those. No free parameters are visible in the abstract. Invented entities are the specific concept classes (center-based margin balls; bounded linear combinations of distance functions) introduced as the objects of study.

axioms (3)
  • standard math Metric spaces satisfy the triangle inequality (and non-negativity, symmetry, identity of indiscernibles).
    The R>3r argument and the claim that only the triangle inequality is needed rest on this.
  • domain assumption Standard PAC (or equivalent) learnability: finite sample complexity for agnostic/realizable classification with respect to a hypothesis class, independent of ambient “dimension.”
    “Learnable in every metric space” is meaningful only under a fixed learning model; the abstract assumes the classical margin/PAC setting.
  • domain assumption Linear classification with margins in Banach spaces is the right comparison class for “kernel-type” reductions.
    The negative embedding result is stated relative to this standard notion of linear margin learnability.
invented entities (2)
  • Center-based (r,R)-margin concepts in arbitrary metric spaces no independent evidence
    purpose: Minimal geometric hypothesis class used to show that large margins + triangle inequality alone yield learnability.
    Defined in the abstract as the starting object of study; not a physical entity but a new formal class for the paper.
  • Bounded linear combinations of distance functions as a concept class no independent evidence
    purpose: Intermediate class for which a sharp universal margin threshold is proved.
    The main positive/negative threshold theorem is stated for this class.

pith-pipeline@v1.1.0-grok45 · 18421 in / 2588 out tokens · 29090 ms · 2026-07-15T13:23:43.424543+00:00 · methodology

0 comments
read the original abstract

Margin-based learning, exemplified by linear and kernel methods, is one of the few classical settings where generalization guarantees are independent of the number of parameters. This makes it a central case study in modern highly over-parameterized learning. We ask what minimal mathematical structure underlies this phenomenon. We begin with a simple margin-based problem in arbitrary metric spaces: concepts are defined by a center point and classify points according to whether their distance lies below $r$ or above $R$. We show that whenever $R>3r$, this class is learnable in \emph{any} metric space. Thus, sufficiently large margins make learnability rely only on the triangle inequality, without any linear or analytic structure being necessary. Our first main result extends this phenomenon to concepts defined by bounded linear combinations of distance functions, and reveals a sharp threshold: there exists a universal constant such that whenever the margin is larger than this constant, the class is learnable in every metric space, while below it there exist metric spaces where it is not learnable at all. We then ask whether margin-based learnability can always be explained via an embedding into a linear space -- that is, reduced to linear classification in some Banach space through a kernel-type construction. We answer this negatively by demonstrating a margin learnable class that cannot be embedded into any Banach space in which linear classification with margins is learnable.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

52 extracted references · 7 linked inside Pith

  1. [1]

    In: International Conference on Learning Rep- resentations (2022)

    Bardes, A., Ponce, J., LeCun, Y.: Vicreg: Variance-invariance-covariance regular- ization for self-supervised learning. In: International Conference on Learning Rep- resentations (2022)

  2. [2]

    Carion, N., Gustafson, L., Hu, Y.T., Debnath, S., Hu, R., Suris, D., Ryali, C., Alwala, K.V., Khedr, H., Huang, A., Lei, J., Ma, T., Guo, B., Kalla, A., Marks, M., Greer, J., Wang, M., Sun, P., Rädle, R., Afouras, T., Mavroudi, E., Xu, K., Wu, T.H., Zhou, Y., Momeni, L., Hazra, R., Ding, S., Vaze, S., Porcher, F., Li, VINO15 F., Li, S., Kamath, A., Cheng,...

  3. [3]

    In: Advances in Neural Information Processing Systems 33 (2020)

    Caron, M., Misra, I., Mairal, J., Goyal, P., Bojanowski, P., Joulin, A.: Unsuper- vised learning of visual features by contrasting cluster assignments. In: Advances in Neural Information Processing Systems 33 (2020)

  4. [4]

    In: Proceedings of the IEEE/CVF International Conference on Computer Vision

    Caron, M., Touvron, H., Misra, I., Jégou, H., Mairal, J., Bojanowski, P., Joulin, A.: Emerging properties in self-supervised vision transformers. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. pp. 9630–9640 (2021)

  5. [5]

    In: International Conference on Machine Learning

    Chen, T., Kornblith, S., Norouzi, M., Hinton, G.E.: A simple framework for con- trastive learning of visual representations. In: International Conference on Machine Learning. pp. 1597–1607. PMLR (2020)

  6. [6]

    Chen,X.,He,K.:Exploringsimplesiameserepresentationlearning.In:Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 15750–15758 (2021)

  7. [7]

    In: European Conference on Computer Vision

    Damen, D., Doughty, H., Farinella, G.M., Fidler, S., Furnari, A., Kazakos, E., Moltisanti, D., Munro, J., Perrett, T., Price, W., Wray, M.: Scaling egocentric vision: The dataset. In: European Conference on Computer Vision. pp. 753–771. Springer (2018)

  8. [8]

    In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition

    Deng, J., Dong, W., Socher, R., Li, L., Li, K., Fei-Fei, L.: Imagenet: A large-scale hierarchical image database. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. pp. 248–255 (2009)

  9. [9]

    In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

    Ding, S., Li, M., Yang, T., Qian, R., Xu, H., Chen, Q., Wang, J., Xiong, H.: Motion-aware contrastive video representation learning via foreground-background merging. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 9716–9726 (2022)

  10. [10]

    In: International Conference on Learning Representations (2025)

    Dukic, N., Lebailly, T., Tuytelaars, T.: Object-centric pretraining via target en- coder bootstrapping. In: International Conference on Learning Representations (2025)

  11. [11]

    In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

    Dwibedi, D., Aytar, Y., Tompson, J., Sermanet, P., Zisserman, A.: Temporal cycle- consistency learning. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 1801–1810 (2019)

  12. [12]

    http://www.pascal-network.org/challenges/VOC/voc2012/workshop/index.html

    Everingham, M., Van Gool, L., Williams, C.K.I., Winn, J., Zisserman, A.: The PASCAL Visual Object Classes Challenge 2012 (VOC2012) Results. http://www.pascal-network.org/challenges/VOC/voc2012/workshop/index.html

  13. [13]

    In: Proceedings of the IEEE/CVF International Conference on Computer Vision

    Fan, D., Tong, S., Zhu, J., Sinha, K., Liu, Z., Chen, X., Rabbat, M., Ballas, N., LeCun, Y., Bar, A., Xie, S.: Scaling language-free visual representation learning. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. pp. 370–382 (2025)

  14. [14]

    Fu, Z., Zhao, T.Z., Finn, C.: Mobile ALOHA: learning bimanual mobile manipulation with low-cost whole-body teleoperation (2024), arXiv preprint arXiv:2401.02117

  15. [15]

    In: Advances in Neural Information Processing Systems 33 (2020)

    Grill, J., Strub, F., Altché, F., Tallec, C., Richemond, P.H., Buchatskaya, E., Do- ersch, C., Pires, B.Á., Guo, Z., Azar, M.G., Piot, B., Kavukcuoglu, K., Munos, R., Valko, M.: Bootstrap your own latent - A new approach to self-supervised learning. In: Advances in Neural Information Processing Systems 33 (2020)

  16. [16]

    In: International Con- ference on Learning Representations (2024) 16 Seul-Ki Yeom , Marcel Simon , Eunbin Lee , and Tae-Ho Kim

    Hamidieh, K., Zhang, H., Sankaranarayanan, S., Ghassemi, M.: Views can be de- ceiving: Improved SSL through feature space augmentation. In: International Con- ference on Learning Representations (2024) 16 Seul-Ki Yeom , Marcel Simon , Eunbin Lee , and Tae-Ho Kim

  17. [17]

    In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

    Han, T., Gokay, D., Heyward, J., Zhang, C., Zoran, D., Patraucean, V., Carreira, J., Damen, D., Zisserman, A.: Learning from streaming video with orthogonal gradients. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 13651–13660 (2025)

  18. [18]

    In: Proceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition

    He, K., Chen, X., Xie, S., Li, Y., Dollár, P., Girshick, R.B.: Masked autoencoders are scalable vision learners. In: Proceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition. pp. 15979–15988 (2022)

  19. [19]

    In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

    He, K., Fan, H., Wu, Y., Xie, S., Girshick, R.B.: Momentum contrast for unsuper- vised visual representation learning. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 9726–9735 (2020)

  20. [20]

    In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

    Kakogeorgiou, I., Gidaris, S., Karantzalos, K., Komodakis, N.: SPOT: self-training with patch-order permutation for object-centric learning with autoregressive trans- formers. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 22776–22786 (2024)

  21. [21]

    In: Avidan, S., Brostow, G.J., Cissé, M., Farinella, G.M., Hassner, T

    Kakogeorgiou, I., Gidaris, S., Psomas, B., Avrithis, Y., Bursuc, A., Karantzalos, K., Komodakis, N.: What to hide from your students: Attention-guided masked image modeling. In: Avidan, S., Brostow, G.J., Cissé, M., Farinella, G.M., Hassner, T. (eds.) European Conference on Computer Vision. pp. 300–318. Springer (2022)

  22. [22]

    Kay, W., Carreira, J., Simonyan, K., Zhang, B., Hillier, C., Vijayanarasimhan, S., Viola, F., Green, T., Back, T., Natsev, P., Suleyman, M., Zisserman, A.: The kinetics human action video dataset (2017), arXiv preprint arXiv:1705.06950

  23. [23]

    arXiv preprint arXiv:2406.09246 (2024)

    Kim, M., Pertsch, K., Karamcheti, S., Xiao, T., Balakrishna, A., Nair, S., Rafailov, R., Foster, E., Lam, G., Sanketi, P., Vuong, Q., Kollar, T., Burchfiel, B., Tedrake, R., Sadigh, D., Levine, S., Liang, P., Finn, C.: Openvla: An open-source vision- language-action model. arXiv preprint arXiv:2406.09246 (2024)

  24. [24]

    In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

    Kim, Y., Mo, S., Kim, M., Lee, K., Lee, J., Shin, J.: Discovering and mitigat- ing visual biases through keyword explanation. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 11082–11092 (2024)

  25. [25]

    In: Proceedings of the IEEE/CVF International Conference on Computer Vision

    Kirillov, A., Mintun, E., Ravi, N., Mao, H., Rolland, C., Gustafson, L., Xiao, T., Whitehead, S., Berg, A.C., Lo, W.Y., Dollar, P., Girshick, R.: Segment anything. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. pp. 4015–4026 (2023)

  26. [26]

    In: Interna- tional Conference on Learning Representations (2024)

    Lebailly, T., Stegmüller, T., Bozorgtabar, B., Thiran, J., Tuytelaars, T.: Cribo: Self-supervised learning via cross-image object-level bootstrapping. In: Interna- tional Conference on Learning Representations (2024)

  27. [27]

    In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

    Li, Z., Zhu, Y., Yang, F., Li, W., Zhao, C., Chen, Y., Chen, Z., Xie, J., Wu, L., Zhao, R., Tang, M., Wang, J.: Univip: A unified framework for self-supervised visual pre-training. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 14627–14636 (2022)

  28. [28]

    In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

    Liu, Y., Xu, Q., Wen, P., Dai, S., Huang, Q.: When the future becomes the past: Taming temporal correspondence for self-supervised video representation learning. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 24033–24044 (2025)

  29. [29]

    In: European Conference on Computer Vision

    Misra, I., Zitnick, C.L., Hebert, M.: Shuffle and learn: Unsupervised learning using temporal order verification. In: European Conference on Computer Vision. pp. 527–544. Springer (2016)

  30. [30]

    In: Ranzato, M., Beygelzimer, A., Dauphin, Y.N., Liang, P., Vaughan, J.W

    Mo, S., Kang, H., Sohn, K., Li, C., Shin, J.: Object-aware contrastive learning for debiased scene representation. In: Ranzato, M., Beygelzimer, A., Dauphin, Y.N., Liang, P., Vaughan, J.W. (eds.) Advances in Neural Information Processing Sys- tems 34. pp. 12251–12264 (2021) VINO17

  31. [31]

    NVIDIA: Cosmos world foundation model platform for physical ai (2025), arXiv preprint arXiv:2501.03575

  32. [32]

    Transactions on Ma- chine Learning Research2024(2024)

    Oquab, M., Darcet, T., Moutakanni, T., Vo, H.V., Szafraniec, M., Khalidov, V., Fernandez, P., Haziza, D., Massa, F., El-Nouby, A., Assran, M., Ballas, N., Galuba, W., Howes, R., Huang, P., Li, S., Misra, I., Rabbat, M., Sharma, V., Synnaeve, G., Xu, H., Jégou, H., Mairal, J., Labatut, P., Joulin, A., Bojanowski, P.: Di- nov2: Learning robust visual featur...

  33. [33]

    In: Proceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition

    Park, G.Y., Jung, C., Lee, S., Ye, J.C., Lee, S.W.: Self-supervised debiasing using low rank regularization. In: Proceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition. pp. 12395–12405 (2024)

  34. [34]

    In: International Conference on Learning Representations (2025)

    Ravi, N., Gabeur, V., Hu, Y.T., Hu, R., Ryali, C., Ma, T., Khedr, H., Rädle, R., Rolland, C., Gustafson, L., Mintun, E., Pan, J., Alwala, K.V., Carion, N., Wu, C.Y., Girshick, R., Dollar, P., Feichtenhofer, C.: SAM 2: Segment anything in images and videos. In: International Conference on Learning Representations (2025)

  35. [35]

    In: British Machine Vision Conference

    Siméoni, O., Puy, G., Vo, H.V., Roburin, S., Gidaris, S., Bursuc, A., Pérez, P., Marlet, R., Ponce, J.: Localizing objects with self-supervised transformers and no labels. In: British Machine Vision Conference. p. 310 (2021)

  36. [36]

    Siméoni, O., Vo, H.V., Seitzer, M., Baldassarre, F., Oquab, M., Jose, C., Khalidov, V., Szafraniec, M., Yi, S.E., Ramamonjisoa, M., Massa, F., Haziza, D., Wehrstedt, L., Wang, J., Darcet, T., Moutakanni, T., Sentana, L., Roberts, C., Vedaldi, A., Tolan, J., Brandt, J., Couprie, C., Mairal, J., Jégou, H., Labatut, P., Bojanowski, P.: Dinov3 (2025), arXiv p...

  37. [37]

    arXiv preprint arXiv:2101.02722 (2021)

    Stone, A., Ramirez, O., Konolige, K., Jonschkowski, R.: The distracting control suite – a challenging benchmark for reinforcement learning from pixels. arXiv preprint arXiv:2101.02722 (2021)

  38. [38]

    In: European Conference on Computer Vision

    Teed, Z., Deng, J.: RAFT: recurrent all-pairs field transforms for optical flow. In: European Conference on Computer Vision. pp. 402–419. Springer (2020)

  39. [39]

    In: International Conference on Learning Representations (2024)

    Venkataramanan, S., Rizve, M.N., Carreira, J., Asano, Y.M., Avrithis, Y.: Is ima- genet worth 1 video? learning strong image encoders from 1 long unlabelled video. In: International Conference on Learning Representations (2024)

  40. [40]

    In: International Conference on Learning Representations (2025)

    Wang, A.N., Hoang, C., Xiong, Y., LeCun, Y., Ren, M.: PooDLe: Pooled and dense self-supervised learning from naturalistic videos. In: International Conference on Learning Representations (2025)

  41. [41]

    In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

    Wang, J., Gao, Y., Li, K., Lin, Y., Ma, A.J., Cheng, H., Peng, P., Huang, F., Ji, R., Sun, X.: Removing the background by adding the background: Towards background robust self-supervised video representation learning. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 11804–11813 (2021)

  42. [42]

    In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

    Wang, X., Jabri, A., Efros, A.A.: Learning correspondence from the cycle- consistency of time. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 2566–2576 (2019)

  43. [43]

    In: IEEE Conference on Computer Vision and Pattern Recognition

    Wang, X., Zhang, R., Shen, C., Kong, T., Li, L.: Dense contrastive learning for self-supervised visual pre-training. In: IEEE Conference on Computer Vision and Pattern Recognition. pp. 3024–3033. Computer Vision Foundation / IEEE (2021)

  44. [44]

    In: Proceedings of the IEEE/CVF International Conference on Computer Vision

    Wei, Y., Church, S., Suciu, V., Lin, J., Wu, C.E., Morgado, P.: Trackverse: A large-scale object-centric video dataset for image-level representation learning. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. pp. 11153–11163 (2025) 18 Seul-Ki Yeom , Marcel Simon , Eunbin Lee , and Tae-Ho Kim

  45. [45]

    In: Advances in Neural Information Processing Systems (2022)

    Wen, X., Zhao, B., Zheng, A., Zhang, X., Qi, X.: Self-supervised visual repre- sentation learning with semantic grouping. In: Advances in Neural Information Processing Systems (2022)

  46. [46]

    In: Advances in Neural Information Processing Systems 34

    Xie, J., Zhan, X., Liu, Z., Ong, Y.S., Loy, C.C.: Unsupervised object-level represen- tation learning from scene images. In: Advances in Neural Information Processing Systems 34. pp. 28864–28876 (2021)

  47. [47]

    In: Proceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

    Xie, Z., Lin, Y., Zhang, Z., Cao, Y., Lin, S., Hu, H.: Propagate yourself: Exploring pixel-level consistency for unsupervised visual representation learning. In: Proceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 16684–16693 (2021)

  48. [48]

    In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

    Xu, D., Xiao, J., Zhao, Z., Shao, J., Xie, D., Zhuang, Y.: Self-supervised spatiotem- poral learning via video clip order prediction. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 10334–10343 (2019)

  49. [49]

    In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (2020)

    Yu, F., Chen, H., Wang, X., Xian, W., Chen, Y., Liu, F., Madhavan, V., Dar- rell, T.: Bdd100k: A diverse driving dataset for heterogeneous multitask learning. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (2020)

  50. [50]

    In: International Conference on Machine Learn- ing

    Zbontar, J., Jing, L., Misra, I., LeCun, Y., Deny, S.: Barlow twins: Self-supervised learning via redundancy reduction. In: International Conference on Machine Learn- ing. pp. 12310–12320. PMLR (2021)

  51. [51]

    In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

    Zhao, J., Li, T., Jiang, D., Wu, S., Ramirez, A., Lee, T.S.: Perceptual inductive bias is what you need before contrastive learning. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 9621–9630 (2025)

  52. [52]

    In: International Conference on Learning Representations (2022)

    Zhou, J., Wei, C., Wang, H., Shen, W., Xie, C., Yuille, A.L., Kong, T.: Image BERT pre-training with online tokenizer. In: International Conference on Learning Representations (2022)