Pith. sign in

Paper Citation Record · LEDGER

DINOv2: Learning Robust Visual Features without Supervision

As of 4 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 100 inbound Pith citation observations for arXiv:2304.07193.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2304.07193 v2

Coverage vector

measured 37 of 37 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-09T04:17:19.878360Z

measured 137 of 137 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 100 of 1292 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T21:31:08.235510Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-11T00:07:42.299741Z

Reference resolution

37 of 37 outbound references displayed

  • verified exact17
  • verified fuzzy11
  • unresolved5
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch4

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 55b016cc-551c-47ad-818b-92042c006700 · outbound

This paper cites an unresolved cited work.

DINOv2: Learning Robust Visual Features without Supervision Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-05-09T04:17:20.511139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T04:17:19.878360Z digest=sha256:cedfbaa7829c4b3896565dd11a30acb65b2238e38c298b93d7e73bffaddacffc

Observation c0eead4b-6bd2-4b7d-a913-e85abd8f0764 · outbound

This paper cites MultiGrain: a unified image embedding for classes and instances.

DINOv2: Learning Robust Visual Features without Supervision MultiGrain: a unified image embedding for classes and instances

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-09T04:17:20.459328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T04:17:19.878360Z digest=sha256:568f7ddc57dceb63548ca66862062fef11bca3ec7468e945fe5fd9c25befb75c

Observation af299ae4-93fd-4081-9dad-3e486c880ac6 · outbound

This paper cites Are we done with ImageNet?.

DINOv2: Learning Robust Visual Features without Supervision Are we done with ImageNet?

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-09T04:17:20.394194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T04:17:19.878360Z digest=sha256:c83d853ed4791f1b789fbfa66af2c846b0f5f027313c3a36b860dfb2cbb2764c

Observation e9595bf0-923e-469b-ac36-f611f6ad4762 · outbound

This paper cites On the Opportunities and Risks of Foundation Models.

DINOv2: Learning Robust Visual Features without Supervision On the Opportunities and Risks of Foundation Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-10T16:15:53.986238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T04:17:19.878360Z digest=sha256:3d7903dee3ff0685fb9cacc46a11ce412e1e35291541ea2b2f3428b8e91955ee

Observation 57ca71f7-d270-4429-9d0d-72c6731cf7dd · outbound

This paper cites Symbolic Discovery of Optimization Algorithms.

DINOv2: Learning Robust Visual Features without Supervision Symbolic Discovery of Optimization Algorithms

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-09T04:17:20.443603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T04:17:19.878360Z digest=sha256:cb38ccdbfc76e40f531749480b5d984de9d4c09ed4f4b2d3a6f06bf04e554d65

Observation 22f17709-8612-49f1-88e7-495ae501f91e · outbound

This paper cites An empirical study of training self-supervised vision transformers.

DINOv2: Learning Robust Visual Features without Supervision An empirical study of training self-supervised vision transformers

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-09T04:17:20.464274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T04:17:19.878360Z digest=sha256:a1ac2867ae71f8b8b321ea0b09f7e0a5f24822ffc12e770490560a72190c9630

Observation 2e1f5db9-d418-43e4-9e18-357bf242629b · outbound

This paper cites PaLM: Scaling Language Modeling with Pathways.

DINOv2: Learning Robust Visual Features without Supervision PaLM: Scaling Language Modeling with Pathways

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:45:07.766206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T04:17:19.878360Z digest=sha256:8b3f65557c077b4e8b9f80bede0af4a524b1f02bd51084c6385d1a419a93e49d

Observation 37267cd6-9275-444c-baf4-9082af89a338 · outbound

This paper cites A Simple Recipe for Competitive Low-compute Self supervised Vision Models.

DINOv2: Learning Robust Visual Features without Supervision A Simple Recipe for Competitive Low-compute Self supervised Vision Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-09T04:17:20.363804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T04:17:19.878360Z digest=sha256:f2ede91199f3f18b09cbc0d344afbcf20d9c4743cb5630b2b9ed1a17ac916300

Observation 9b809046-68ed-42d5-914a-67591100a026 · outbound

This paper cites Are Large-scale Datasets Necessary for Self-Supervised Pre-training?.

DINOv2: Learning Robust Visual Features without Supervision Are Large-scale Datasets Necessary for Self-Supervised Pre-training?

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-09T04:17:20.390653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T04:17:19.878360Z digest=sha256:a1963bdcbbb16836b66f8141a54a1248dafc86780c357687b63901c7f441b00c

Observation 9b24cbce-dfda-43fa-9423-a5e6a1a5e296 · outbound

This paper cites Eva: Exploring the limits of masked visual representation learning at scale.

DINOv2: Learning Robust Visual Features without Supervision Eva: Exploring the limits of masked visual representation learning at scale

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-09T04:17:20.468646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T04:17:19.878360Z digest=sha256:f25fce35f4e7b9be08f74472e402c86295673219f72dcdf79e9768a9b842357e

Observation 9976cf06-4437-4342-ba15-fce8aad4bc35 · outbound

This paper cites Self-supervised Pretraining of Visual Features in the Wild.

DINOv2: Learning Robust Visual Features without Supervision Self-supervised Pretraining of Visual Features in the Wild

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-09T04:17:20.403346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T04:17:19.878360Z digest=sha256:1cbb51a7820cbd8a79fc20640fa89db9613e98573b7b0b66d2c974765758d79f

Observation 8ba28ccd-9876-4e0d-9bab-498307162344 · outbound

This paper cites Vision Models Are More Robust And Fair When Pretrained On Uncurated Images Without Supervision.

DINOv2: Learning Robust Visual Features without Supervision Vision Models Are More Robust And Fair When Pretrained On Uncurated Images Without Supervision

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-09T04:17:20.412841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T04:17:19.878360Z digest=sha256:3755586fac2d3dbb13d95d43129fe74f6ed0ea5f00bd41333659f004cecc61ab

Observation 59a753fb-2ab0-4bd2-b0b0-4b73ef2635e2 · outbound

This paper cites The many faces of robustness: A critical analysis of out-of-distribution generalization.

DINOv2: Learning Robust Visual Features without Supervision The many faces of robustness: A critical analysis of out-of-distribution generalization

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-09T04:17:20.471080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T04:17:19.878360Z digest=sha256:4e8e9e853d7aeebb7cfdb3e32625ddb8be794eaf1b0dbc84f40a0b7c81a55f7a

Observation 6da0582a-384f-406c-b65b-f9bed07efe79 · outbound

This paper cites Training Compute-Optimal Large Language Models.

DINOv2: Learning Robust Visual Features without Supervision Training Compute-Optimal Large Language Models

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-10T16:01:32.444768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T04:17:19.878360Z digest=sha256:2b4e0d93ea064627f91d69d11efcaa7b0e8dcc8a7c96721cd9d9c96844b4add3

Observation 1d52adaa-7998-4af3-a14b-c171b5807a2e · outbound

This paper cites The Kinetics Human Action Video Dataset.

DINOv2: Learning Robust Visual Features without Supervision The Kinetics Human Action Video Dataset

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:13:45.682066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T04:17:19.878360Z digest=sha256:6a3d5ed102f62af283e350beac0f368c351e1e461d129e0bd2b92f7459a3e5b7

Observation 171fa669-b057-4b32-98a4-f60e80f231fc · outbound

This paper cites BinsFormer: Revisiting Adaptive Bins for Monocular Depth Estimation.

DINOv2: Learning Robust Visual Features without Supervision BinsFormer: Revisiting Adaptive Bins for Monocular Depth Estimation

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-09T04:17:20.434323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T04:17:19.878360Z digest=sha256:866a6c4eb61a7d7c581cea3ca9f79db4e8fbc14d2d00ee6e9f61a95134f386b2

Observation 62e0ff4e-9338-4509-bf01-155d3599c779 · outbound

This paper cites Polarized Self-Attention: Towards High-quality Pixel-wise Regression.

DINOv2: Learning Robust Visual Features without Supervision Polarized Self-Attention: Towards High-quality Pixel-wise Regression

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T04:17:20.440163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T04:17:19.878360Z digest=sha256:4bdf9072bbcfdc032c98451eebfead4f060ab55817b67d2fba8c8d021db5cfb3

Observation 124182fc-b3af-4ada-bf24-26a8e07d4439 · outbound

This paper cites an unresolved cited work.

DINOv2: Learning Robust Visual Features without Supervision Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-05-09T04:17:20.473207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T04:17:19.878360Z digest=sha256:813727e3ff0167beb2e7babcf6c7bbd375f1277b3f881fca8c701d340b8db52d

Observation 73cc2f39-7ce9-48e1-b823-1660fe4dd218 · outbound

This paper cites Carbon Emissions and Large Neural Network Training.

DINOv2: Learning Robust Visual Features without Supervision Carbon Emissions and Large Neural Network Training

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:48:25.196363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T04:17:19.878360Z digest=sha256:7e6d1bd3d72e0991c68e60ff979d1f6868f2427eaa6325ba1537ad7e92382493

Observation 013e819f-8aa5-41c6-a2be-3dbd69246de3 · outbound

This paper cites Learning to Generate Reviews and Discovering Sentiment.

DINOv2: Learning Robust Visual Features without Supervision Learning to Generate Reviews and Discovering Sentiment

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T04:17:20.454668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T04:17:19.878360Z digest=sha256:34fb6e31013a3fe9682d392dad5830cdc7b8a3f7dd247c1a52ee90b4c87d88ce

Observation db2f8fc9-39e0-4edf-8cea-71d977674c59 · outbound

This paper cites Imagenet large scale visual recognition challenge.IJCV.

DINOv2: Learning Robust Visual Features without Supervision Imagenet large scale visual recognition challenge.IJCV

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-09T04:17:20.475822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T04:17:19.878360Z digest=sha256:60fb6204aec44b753f7d367446d58b0d02ab536d11390ee41b17776f1af8e11b

Observation fc68bd92-588a-4e1b-b843-40313b096556 · outbound

This paper cites GLU Variants Improve Transformer.

DINOv2: Learning Robust Visual Features without Supervision GLU Variants Improve Transformer

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-10T17:05:01.922266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T04:17:19.878360Z digest=sha256:1e5ec2251d13f9b7ff17112a2f67ed4fa570ccbd8e3eb6955b8c4793acfffe53

Observation db7e4380-5da3-4887-8f2c-f7a3599c91c2 · outbound

This paper cites UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild.

DINOv2: Learning Robust Visual Features without Supervision UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:25:00.378223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T04:17:19.878360Z digest=sha256:0c04974eb5ea610cf772bee2da8ac28385b7f70069dcb029526285f8fa9bbed7

Observation 6ed76909-00b3-40a8-8b28-c91d7db6b4b7 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

DINOv2: Learning Robust Visual Features without Supervision LLaMA: Open and Efficient Foundation Language Models

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-09T04:17:20.350896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T04:17:19.878360Z digest=sha256:c1ff6f988c4dadacc95bff91337e0162c70e37aa871b802d7847ae55b79e5b2a

Observation fc3561b8-0b4e-412c-afce-e9676ed9742c · outbound

This paper cites Bench- marking representation learning for natural world image collections.

DINOv2: Learning Robust Visual Features without Supervision Bench- marking representation learning for natural world image collections

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-09T04:17:20.484765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T04:17:19.878360Z digest=sha256:421ae65245c534e2551611e2761fa6815db1c9b6397d52ff7b36e5e6645ae018

Observation 4efad419-ca9d-4f8a-817e-a438decb6382 · outbound

This paper cites an unresolved cited work.

DINOv2: Learning Robust Visual Features without Supervision Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-05-09T04:17:20.489551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T04:17:19.878360Z digest=sha256:502acc96717afac2cd334afbd69dfe3ec700247d066cb920c82af86fd908dd7e

Observation d5ad5530-cbb2-4705-9211-1bfe7ba1c95d · outbound

This paper cites Masked Autoencoders that Listen.

DINOv2: Learning Robust Visual Features without Supervision Masked Autoencoders that Listen

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T04:17:20.373783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T04:17:19.878360Z digest=sha256:07cf4d7ef5ebd10c38111d6b14c681bd49de493b126696a6172572e36a28b43a

Observation 2b6c4006-bbb4-477a-9f9b-89b82d51594e · outbound

This paper cites Billion-scale semi-supervised learning for image classification.

DINOv2: Learning Robust Visual Features without Supervision Billion-scale semi-supervised learning for image classification

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-09T04:17:20.377918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T04:17:19.878360Z digest=sha256:e9f346c91534235faf78caf522f181ec1110d4573450b467975afb53d7c933c3

Observation 5368219e-fe1e-45ec-9f52-1946bd270291 · outbound

This paper cites Mugs: A Multi-Granular Self-Supervised Learning Framework.

DINOv2: Learning Robust Visual Features without Supervision Mugs: A Multi-Granular Self-Supervised Learning Framework

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T04:17:20.383765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T04:17:19.878360Z digest=sha256:aa7a88897dfd0f9a57996d5ceb2b8b9638e147e2140e3fcd657cb99cf98d4d6c

Observation 51f78af0-e787-4f12-8270-1277c70523bd · outbound

This paper cites an unresolved cited work.

DINOv2: Learning Robust Visual Features without Supervision Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-05-09T04:17:20.491851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T04:17:19.878360Z digest=sha256:75459974dcb29c93c06cbe2b41cdc1d21bcdc3d9850cd04b91c69e4623d742f4

Observation 7af36918-2a23-4a77-8ca4-e411a2ca391e · outbound

This paper cites We apply the KoLeo regularizer with a weight of 0.1 between the class tokens of the first global crop, for all samples within a GPU without cross-communication for this step.

DINOv2: Learning Robust Visual Features without Supervision We apply the KoLeo regularizer with a weight of 0.1 between the class tokens of the first global crop, for all samples within a GPU without cross-communication for this step

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-09T04:17:20.494541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T04:17:19.878360Z digest=sha256:77f932692f7e510ac93eaf254c5258cf4e0c25db6081cc0e106021bafff5fbb0

Observation 9d2a2454-8422-42a0-8e64-f0e9ac780900 · outbound

This paper cites an unresolved cited work.

DINOv2: Learning Robust Visual Features without Supervision Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-05-09T04:17:20.499666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T04:17:19.878360Z digest=sha256:9c971fafe3ec1d6f31c01d5b3e878bff389bcca5d72a9ee08b5f6187e7587c64

Observation f91d300c-1f50-42d5-9752-77788ab4c79c · outbound

This paper cites EMA update for the teacher.

DINOv2: Learning Robust Visual Features without Supervision EMA update for the teacher

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-09T04:17:20.501888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T04:17:19.878360Z digest=sha256:fb78f221bdc803906678702a65833988df1170aad5325fb453bd164866114adb

Observation da1bec99-2754-4351-b94f-13602ac2a4fd · outbound

This paper cites (Everingham et al.

DINOv2: Learning Robust Visual Features without Supervision (Everingham et al

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-09T04:17:20.504022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T04:17:19.878360Z digest=sha256:ea8e81bbe9c4f273f038b1f368dc404fc2083430f664648a762615e1aa2ba556

Observation 8f6af0b6-6c74-4f2d-9268-cf84126ab147 · outbound

This paper cites (Van Horn et al.

DINOv2: Learning Robust Visual Features without Supervision (Van Horn et al

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-09T04:17:20.507188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T04:17:19.878360Z digest=sha256:b76d6c2b4772dd74314d5eb55b907d5e6ed85e459d827876d87ab0f6ee05cbaa

Observation 8b529756-49db-4c09-8226-a3bd04f896f6 · outbound

This paper cites (Van Horn et al.

DINOv2: Learning Robust Visual Features without Supervision (Van Horn et al

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-09T04:17:20.509241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T04:17:19.878360Z digest=sha256:d0b236ca99131d67ccd6f372357aafeb6008af0645f14abca05327f1cfc01fe8

Observation e912152b-8116-4a04-a0a0-f19fd0189464 · outbound

This paper cites (Everingham et al.

DINOv2: Learning Robust Visual Features without Supervision (Everingham et al

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-09T04:17:20.461799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T04:17:19.878360Z digest=sha256:fca945dea436180f78b397425eef2c3ff72837ee1d92a1f2e0cfb76e6d789ff6

Pith citing papers

Observation 577d983d-58a0-4a35-a2ee-2e27ae4f9a2f · inbound

RoMa: Robust Dense Feature Matching cites this paper.

RoMa: Robust Dense Feature Matching DINOv2: Learning Robust Visual Features without Supervision

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-24T08:59:14.891821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-24T08:56:39.120899Z digest=sha256:a3bbe37e366de88f28fedf58905291bd708e0e3d117bb6147c71211a81954b36

Observation 2e7b95bf-ef3d-4ca8-b2ba-ab0734c7791a · inbound

A Survey on Multimodal Large Language Models cites this paper.

A Survey on Multimodal Large Language Models DINOv2: Learning Robust Visual Features without Supervision

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-16T02:56:42.293758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T02:56:41.658658Z digest=sha256:0f5d15de5ff406f6638331a8a8beb54c061b20be6de54427751741600d808afd

Observation c758da9f-4ae8-4c2a-829c-5088e4d69c04 · inbound

Project Aria: A New Tool for Egocentric Multi-Modal AI Research cites this paper.

Project Aria: A New Tool for Egocentric Multi-Modal AI Research DINOv2: Learning Robust Visual Features without Supervision

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-15T08:25:20.895920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T08:25:20.868735Z digest=sha256:5f29bc0d69c220b368431e88b1502eed305e97fd3a9c23d5d98c6e3e3278b856

Observation a1b0f07f-ab34-41b3-9bb8-417fb5194bdc · inbound

Vision Transformers Need Registers cites this paper.

Vision Transformers Need Registers DINOv2: Learning Robust Visual Features without Supervision

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T09:41:38.090469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-13T09:41:37.937046Z digest=sha256:0cbb4566ef30513106d8dd4b8c809df48fb90b953bdf4adf6d8658c8596a80bc

Observation c18b963d-d5e6-4add-aa51-6e8965ce4b82 · inbound

Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution cites this paper.

Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution DINOv2: Learning Robust Visual Features without Supervision

Reference 249

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T08:12:31.100422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-16T08:12:30.984870Z digest=sha256:3a65d5e4f08d7a642286ab3c15f7b1bfa7275948b56bba168f092d63eb891336

Observation a28bd355-7217-4e35-879d-d1ba5c41ef21 · inbound

Causal Unsupervised Semantic Segmentation cites this paper.

Causal Unsupervised Semantic Segmentation DINOv2: Learning Robust Visual Features without Supervision

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-24T05:53:57.285384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-24T05:53:15.996646Z digest=sha256:0140803f40cd781f0594168a3757dd1c56d556fe01204c64f3e6b1c980e8c929

Observation c4a0bc69-d527-426e-bb7a-b1c3b5d43e85 · inbound

InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks cites this paper.

InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks DINOv2: Learning Robust Visual Features without Supervision

Reference 112

Resolution
verified exact
local_arxiv, observed 2026-05-13T22:46:10.078185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T22:46:09.693156Z digest=sha256:b49e5bfd5d57ef24b3796dffd7eac9e18e929e826da50668100651e89578000c

Observation 1d8194b2-b8a3-4c50-922f-23b075ee3a32 · inbound

Data-Centric Foundation Models in Computational Healthcare: A Survey cites this paper.

Data-Centric Foundation Models in Computational Healthcare: A Survey DINOv2: Learning Robust Visual Features without Supervision

Reference 215

Resolution
verified exact
local_arxiv, observed 2026-05-24T04:13:52.777696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-24T04:13:05.328492Z digest=sha256:cf9b0b811f5451b6711da4bfdea26bd39e2747bd21b6a81598427987f15313c8

Observation ba32c141-fd1b-4623-b3e4-291510747b5d · inbound

Massive Activations in Large Language Models cites this paper.

Massive Activations in Large Language Models DINOv2: Learning Robust Visual Features without Supervision

Reference 140

Resolution
verified exact
local_arxiv, observed 2026-05-16T07:02:54.029334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:6221d49f7ab2e9296dd20193471e17264d0cca86141e5cbaf068d38bde007cdc

Observation 3cb14679-d2ff-4fa7-8471-f157b95c1c26 · inbound

TempCompass: Do Video LLMs Really Understand Videos? cites this paper.

TempCompass: Do Video LLMs Really Understand Videos? DINOv2: Learning Robust Visual Features without Supervision

Reference 112

Resolution
verified exact
local_arxiv, observed 2026-05-17T02:46:16.788683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-17T02:46:16.632743Z digest=sha256:e7de69aa636ef510ce97289eaae2d6cf3a57c8bbfc5fa4a7a1a69ec705f2b332

Observation 6827c88e-c9f4-4974-9d28-a097dea10b28 · inbound

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training cites this paper.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training DINOv2: Learning Robust Visual Features without Supervision

Reference 88

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T04:09:36.151501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:bf520d5db103c28f1a7bff6d430e1dc00e15bee608f224c82ab072445706819c

Observation e49771d3-488e-4d21-b591-bf848bf5d706 · inbound

Revisiting Feature Prediction for Learning Visual Representations from Video cites this paper.

Revisiting Feature Prediction for Learning Visual Representations from Video DINOv2: Learning Robust Visual Features without Supervision

Reference 263

Resolution
verified exact
local_arxiv, observed 2026-05-12T12:40:23.983455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T12:40:23.709098Z digest=sha256:30ffa97588137796b6357dcc4d12b2aca7be1aebd01d77abe1d02def665cfc8f

Observation 80e3b255-1a0f-4d1d-be93-5dab2b44f614 · inbound

Leveraging Medical Foundation Model Features in Graph Neural Network-Based Retrieval of Breast Histopathology Images cites this paper.

Leveraging Medical Foundation Model Features in Graph Neural Network-Based Retrieval of Breast Histopathology Images DINOv2: Learning Robust Visual Features without Supervision

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-05-24T01:13:42.982734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-24T01:09:07.277220Z digest=sha256:91a6218668f537bb0e5c2bc63eea5be992fa48a39498b3e76474b69f70209b84

Observation 2e95bcef-086c-4e18-9468-a39b72a2bae7 · inbound

The Platonic Representation Hypothesis cites this paper.

The Platonic Representation Hypothesis DINOv2: Learning Robust Visual Features without Supervision

Reference 197

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T06:03:56.614486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-15T06:03:56.328012Z digest=sha256:07400cc0f64d2e82a62a99f0de67202ccc99dbfefb1bc8642c18b74670045507

Observation 9d98c612-6b4b-46ef-80d9-e0a50e3b47bb · inbound

OpenVLA: An Open-Source Vision-Language-Action Model cites this paper.

OpenVLA: An Open-Source Vision-Language-Action Model DINOv2: Learning Robust Visual Features without Supervision

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-10T14:46:36.324591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T14:46:35.942338Z digest=sha256:fc72d3dd6e1446b6f770fd9495ff07d13fe96c7ef3e569c49468d3a0c935e83a

Observation bcc57d77-533f-44e5-ba13-1e1ea51e6e6d · inbound

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models cites this paper.

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models DINOv2: Learning Robust Visual Features without Supervision

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-11T06:01:53.869766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T06:01:53.730356Z digest=sha256:248ad17b7434b5a9de45c793b61fd73c32f71d7ebb35a1571e29b14bc75bae49

Observation 507ab40c-7d9b-4c79-bb9b-06b34cea6ec6 · inbound

ReKep: Spatio-Temporal Reasoning of Relational Keypoint Constraints for Robotic Manipulation cites this paper.

ReKep: Spatio-Temporal Reasoning of Relational Keypoint Constraints for Robotic Manipulation DINOv2: Learning Robust Visual Features without Supervision

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-16T08:25:18.015536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:25:17.847571Z digest=sha256:56306be42c8a2ce19b10ad1fa2ed6fcadfe724069c6426537622da6629500caf

Observation cc5d14ed-fe89-45d6-b204-2febe308a956 · inbound

MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark cites this paper.

MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark DINOv2: Learning Robust Visual Features without Supervision

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-05-14T00:51:48.393113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-14T00:51:48.163349Z digest=sha256:fe22091f5a024cbbf33680eeb7052d8487a7b8851ca9004519eec664360f7e9c

Observation e13f6784-27fc-4f31-a098-5fce1a425133 · inbound

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models cites this paper.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models DINOv2: Learning Robust Visual Features without Supervision

Reference 91

Resolution
verified exact
local_arxiv, observed 2026-05-15T01:55:12.679057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:a1f74b98e30f8d17973ff1913318da12d13cec4dc8ad10a685ef7f15565f7a7a

Observation 61c9ea69-b5ee-4590-bad6-dba33e6bffd7 · inbound

TOAST: Transformer Optimization using Adaptive and Simple Transformations cites this paper.

TOAST: Transformer Optimization using Adaptive and Simple Transformations DINOv2: Learning Robust Visual Features without Supervision

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-23T19:48:23.026543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-23T19:46:07.124996Z digest=sha256:eb17899eb9be0ccb70b77aad2c0a3e3a9a5fd3da14674c970ecf53f0c202bbaa

Observation a07877d8-63bb-4be8-9e92-280e74695b8a · inbound

UNCOM: Zero-shot Context-Aware Command Understanding for Tabletop Scenarios cites this paper.

UNCOM: Zero-shot Context-Aware Command Understanding for Tabletop Scenarios DINOv2: Learning Robust Visual Features without Supervision

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-23T19:18:20.622008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T19:17:12.954354Z digest=sha256:5b812ed039fd9e96350d7a921c15d23ca491eb4b2c9b1cc137431cc1be6d9702

Observation 2493597c-661c-4c95-aa04-690c68546a8b · inbound

LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding cites this paper.

LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding DINOv2: Learning Robust Visual Features without Supervision

Reference 20

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T13:53:33.682274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T13:53:33.585035Z digest=sha256:4880178283c3d3234cc44810443b07eac9b9c8591782410b021f006520a99b8b

Observation 954f3313-166b-48fd-9fad-01a0ebc22257 · inbound

DINO-WM: World Models on Pre-trained Visual Features enable Zero-shot Planning cites this paper.

DINO-WM: World Models on Pre-trained Visual Features enable Zero-shot Planning DINOv2: Learning Robust Visual Features without Supervision

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-05-17T16:06:09.565500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-17T16:06:09.448517Z digest=sha256:0419525ce0797871a76768ec2d9a44e272d20ea2019c1bb14ff82637adc0013c

Observation 939fee17-02ae-477d-a9f4-6f81433fc6f7 · inbound

DissolveStereo: Coarse Depth Injection for Zero-Shot Stereo Video Generation cites this paper.

DissolveStereo: Coarse Depth Injection for Zero-Shot Stereo Video Generation DINOv2: Learning Robust Visual Features without Supervision

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-23T17:18:13.915063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T17:17:48.961672Z digest=sha256:004fb4096d7bada8f14a4ba113cfb3269648d13064ded37d56d49b8daa27379f

Observation 3c98600f-acdc-4a4f-854f-28d7deed86b4 · inbound

Large Language Model-Brained GUI Agents: A Survey cites this paper.

Large Language Model-Brained GUI Agents: A Survey DINOv2: Learning Robust Visual Features without Supervision

Reference 229

Resolution
verified exact
local_arxiv, observed 2026-05-19T11:08:27.800783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T11:08:27.472508Z digest=sha256:96e6a69550e7d136d0dc8b9d484c54360beb36e701b62d7aaa6f191d12b807d8

Observation a3cb008c-22b0-44a4-af34-5c67146fb598 · inbound

Multimodal Contextualized Support for Enhancing Video Retrieval System cites this paper.

Multimodal Contextualized Support for Enhancing Video Retrieval System DINOv2: Learning Robust Visual Features without Supervision

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-23T07:15:28.477625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T07:14:34.843867Z digest=sha256:af40aaa7214e200455abc46180f931d8a27777f2f8536b22a939048b4cc6592e

Observation 3171fb7b-80e1-4b08-81e8-5d4a6fa242a5 · inbound

TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies cites this paper.

TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies DINOv2: Learning Robust Visual Features without Supervision

Reference 90

Resolution
verified exact
local_arxiv, observed 2026-05-15T18:27:22.986858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-15T18:27:22.760982Z digest=sha256:dde7963a937fabe4a503019c584830c6b1f0d447126ee6273a89e867971b02ec

Observation da509556-579c-4bb6-a1d7-df292ecc4451 · inbound

DepthMaster: Taming Diffusion Models for Monocular Depth Estimation cites this paper.

DepthMaster: Taming Diffusion Models for Monocular Depth Estimation DINOv2: Learning Robust Visual Features without Supervision

Reference 66

Resolution
verified exact
local_arxiv, observed 2026-05-23T06:12:38.899373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T06:08:14.988386Z digest=sha256:88383b32e64ac5ac39f4fa3cf361d7422880ac0a7be0a3bf49c5396e93b4fc0f

Observation 90e48ea4-efbe-412f-9243-777fcc6fd9c2 · inbound

Benchmarking Vision Foundation Models for Input Monitoring in Autonomous Driving cites this paper.

Benchmarking Vision Foundation Models for Input Monitoring in Autonomous Driving DINOv2: Learning Robust Visual Features without Supervision

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-05-23T05:17:35.604131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T05:16:36.255597Z digest=sha256:7f7ab4b13bb73b7f0312aa36fffc88a8cccc199375e73e831473530af8c4366e

Observation f217d42f-e05c-4275-a0ac-8787f3afc269 · inbound

Inference-Time Scaling for Diffusion Models beyond Scaling Denoising Steps cites this paper.

Inference-Time Scaling for Diffusion Models beyond Scaling Denoising Steps DINOv2: Learning Robust Visual Features without Supervision

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-05-20T11:45:17.642392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T11:45:17.473970Z digest=sha256:64d2650d97cf1e94f3b2ca8476fb67e569140b77fabb7dc8480d3f8d8a3cfe67

Observation 4ae3679b-685b-40b1-8724-444754e2495c · inbound

Personalization Toolkit: Training Free Personalization of Large Vision Language Models cites this paper.

Personalization Toolkit: Training Free Personalization of Large Vision Language Models DINOv2: Learning Robust Visual Features without Supervision

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-23T03:35:21.096950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T03:33:13.549114Z digest=sha256:408e17e7d6626dde693b67b476c94f94e9aa54757e373c7d71b8b99a95fbaa9d

Observation ac2ce2e8-9d79-460e-baff-044e19feca99 · inbound

TripoSG: High-Fidelity 3D Shape Synthesis using Large-Scale Rectified Flow Models cites this paper.

TripoSG: High-Fidelity 3D Shape Synthesis using Large-Scale Rectified Flow Models DINOv2: Learning Robust Visual Features without Supervision

Reference 134

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T21:51:18.439894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-16T21:51:18.323840Z digest=sha256:690d5d7ff7e5113832e0cee7b584019e585526a1ac27f48ec4874ed0529f7c36

Observation 228ae164-ff72-4721-b5c7-b8646e294103 · inbound

Geometry-aided Vision-based Localization of Future Mars Helicopters in Challenging Illumination Conditions cites this paper.

Geometry-aided Vision-based Localization of Future Mars Helicopters in Challenging Illumination Conditions DINOv2: Learning Robust Visual Features without Supervision

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-23T03:05:20.223418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T03:04:06.940154Z digest=sha256:866360588e77bad8019761c52246b264a96064d06701720ec7d137544cbf6680

Observation 7d4a3c5e-9deb-412a-b49d-140949da65fd · inbound

Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success cites this paper.

Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success DINOv2: Learning Robust Visual Features without Supervision

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-11T04:35:32.741990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T04:35:31.914360Z digest=sha256:2d7baf11d6765adb90acb8544be096f61c8e89af384088c5bf6a0cf733994073

Observation 9de94152-88ce-4e88-9fce-ae6004b8eb52 · inbound

UniDepthV2: Universal Monocular Metric Depth Estimation Made Simpler cites this paper.

UniDepthV2: Universal Monocular Metric Depth Estimation Made Simpler DINOv2: Learning Robust Visual Features without Supervision

Reference 90

Resolution
verified exact
local_arxiv, observed 2026-05-17T09:11:04.301077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T09:11:04.176769Z digest=sha256:be50ce3104c3d278ea8f9e7d4c809e37f7de7e768af28a71565fabe42c4bcf0a

Observation 434c806a-a7d3-4c2e-905c-20df8b78e3c6 · inbound

Primus: Enforcing Attention Usage for 3D Medical Image Segmentation cites this paper.

Primus: Enforcing Attention Usage for 3D Medical Image Segmentation DINOv2: Learning Robust Visual Features without Supervision

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-05-23T01:25:16.526151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T01:22:29.901174Z digest=sha256:32af9374d6cb23ec1780e49f4631d38458ad11fceac51c38a909f582c6e9b46f

Observation 41ba1e14-2956-46f7-b53f-3424dd92297e · inbound

Adaptive Camera Sensor for Vision Models cites this paper.

Adaptive Camera Sensor for Vision Models DINOv2: Learning Robust Visual Features without Supervision

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-23T02:17:24.484593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T02:17:02.071721Z digest=sha256:6dea1132a854c75d1edd43df9835ddfd36f21983e2fd8b4b35b5102394032cd2

Observation c48a6835-bca2-433b-a9e2-86a5c833d0a4 · inbound

Seedream 2.0: A Native Chinese-English Bilingual Image Generation Foundation Model cites this paper.

Seedream 2.0: A Native Chinese-English Bilingual Image Generation Foundation Model DINOv2: Learning Robust Visual Features without Supervision

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-17T08:27:36.327182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T08:27:36.242416Z digest=sha256:81d8685374dcc4628684e665cbcf67e6e537208d4ed0dc994899d01cd78d69e8

Observation 547e41d2-bed1-405e-ad5d-21b1c3ab55fc · inbound

HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model cites this paper.

HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model DINOv2: Learning Robust Visual Features without Supervision

Reference 72

Resolution
verified exact
local_arxiv, observed 2026-05-15T22:00:48.779743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T22:00:48.667428Z digest=sha256:087f80d5b987c5dbd401dbdf41a460deb6986152722ad2052acad6ba0d4fc52e

Observation b53f411a-70d2-43a5-b099-6a64821ed13e · inbound

GAIR: Location-Aware Self-Supervised Contrastive Pre-Training with Geo-Aligned Implicit Representations cites this paper.

GAIR: Location-Aware Self-Supervised Contrastive Pre-Training with Geo-Aligned Implicit Representations DINOv2: Learning Robust Visual Features without Supervision

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-05-22T22:42:13.394805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T22:41:00.027266Z digest=sha256:b45045fae8a039f8978cb6f9988edf6da97d70f6716791788c709141c511c712

Observation 2de09c84-0f0f-4d7d-9a02-a7ad63efd874 · inbound

Toward Generalizable Forgery Detection and Reasoning cites this paper.

Toward Generalizable Forgery Detection and Reasoning DINOv2: Learning Robust Visual Features without Supervision

Reference 78

Resolution
verified exact
local_arxiv, observed 2026-05-22T22:27:12.408489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T22:26:49.062978Z digest=sha256:ba4fa0bd2625afc736983bce5683c97e32708c9c6d64e4b56889da3047779005

Observation 44fcb28a-8b3a-482d-a361-35daef47c51a · inbound

Seedream 3.0 Technical Report cites this paper.

Seedream 3.0 Technical Report DINOv2: Learning Robust Visual Features without Supervision

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-13T07:55:38.751884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T07:55:38.690569Z digest=sha256:b9a605fb5f5883841637813121db698584fa9b9c7532abab17823f7d287da781

Observation 395041c0-217e-4114-8a32-235e63c4abf7 · inbound

Benchmarking Large Vision-Language Models on Fine-Grained Image Tasks: A Comprehensive Evaluation cites this paper.

Benchmarking Large Vision-Language Models on Fine-Grained Image Tasks: A Comprehensive Evaluation DINOv2: Learning Robust Visual Features without Supervision

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-22T18:26:55.273681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T18:26:12.597756Z digest=sha256:4bc057f08dc4b0639365ddc99f8cb51b9a337931181b28a7258b961984fa2114

Observation 64e7dc58-8ace-4e75-a10d-0716daf9c294 · inbound

FreeGraftor: Training-Free Cross-Image Feature Grafting for Subject-Driven Text-to-Image Generation cites this paper.

FreeGraftor: Training-Free Cross-Image Feature Grafting for Subject-Driven Text-to-Image Generation DINOv2: Learning Robust Visual Features without Supervision

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-22T18:11:54.536944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T18:09:50.633826Z digest=sha256:b38a8d898cd470902dcab23fe859be12a0f5c0499264bcbcf5b9600f538664cd

Observation 6170fdc2-8c73-4355-9335-c2b64d8b021f · inbound

NORA: A Small Open-Sourced Generalist Vision Language Action Model for Embodied Tasks cites this paper.

NORA: A Small Open-Sourced Generalist Vision Language Action Model for Embodied Tasks DINOv2: Learning Robust Visual Features without Supervision

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-16T15:53:29.471677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T15:53:29.412890Z digest=sha256:545cb96a2825e3d7b4929a7e92819c5566c52abcfde2951c1d04805d3f997e65

Observation 9252b305-6182-4cad-bb40-52ad42356e16 · inbound

In-Context Edit: Enabling Instructional Image Editing with In-Context Generation in Large Scale Diffusion Transformer cites this paper.

In-Context Edit: Enabling Instructional Image Editing with In-Context Generation in Large Scale Diffusion Transformer DINOv2: Learning Robust Visual Features without Supervision

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-05-16T16:07:53.127322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T16:07:53.054355Z digest=sha256:fbb4d0c7972e344e8def5822489814e7bba6e431573450916f4a98b3b8dd2372

Observation e348e3f9-83b2-42d7-85a6-d4f4880692db · inbound

GraspVLA: a Grasping Foundation Model Pre-trained on Billion-scale Synthetic Action Data cites this paper.

GraspVLA: a Grasping Foundation Model Pre-trained on Billion-scale Synthetic Action Data DINOv2: Learning Robust Visual Features without Supervision

Reference 71

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:55:52.269304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T20:55:52.109166Z digest=sha256:500597cf3b648e3346e81d1d4d7ee6fd085e8fa82aef1b7d6b3adaf079c4d802

Observation c1be779b-1e02-44fb-9007-29bec9901230 · inbound

Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation cites this paper.

Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation DINOv2: Learning Robust Visual Features without Supervision

Reference 59

Resolution
verified exact
local_arxiv, observed 2026-05-17T07:24:04.589166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T07:24:04.460276Z digest=sha256:eb9df825eeffec2409b8796e5dd23d04723d76f5e8593cdc30791ea00b3994d6

Observation 6f2c3e87-a354-4abf-bb94-4c044ab34a4e · inbound

Seed1.5-VL Technical Report cites this paper.

Seed1.5-VL Technical Report DINOv2: Learning Robust Visual Features without Supervision

Reference 100

Resolution
verified exact
local_arxiv, observed 2026-05-11T05:26:06.146024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:c48e1d91a2ea561524c1351711546f26c85c88dcb6b38543c35ed34537502e2f

Observation c060c368-9ac5-48ff-9cb7-a943d60c43bc · inbound

VGGT-SLAM: Dense RGB SLAM Optimized on the SL(4) Manifold cites this paper.

VGGT-SLAM: Dense RGB SLAM Optimized on the SL(4) Manifold DINOv2: Learning Robust Visual Features without Supervision

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-05-20T21:03:14.380167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T21:03:14.217231Z digest=sha256:ecbfa4ac0a0ecdce5782fd71fb8f683484280e767b9f46992ceca509da182dc1

Observation dd52bf2f-0608-4efe-abea-051583d32cbe · inbound

MAGI-1: Autoregressive Video Generation at Scale cites this paper.

MAGI-1: Autoregressive Video Generation at Scale DINOv2: Learning Robust Visual Features without Supervision

Reference 35

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T20:31:15.838885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T20:31:15.700943Z digest=sha256:ea93762ad5880d89a1969ccc96ce8a160e689a626340cd3239c7ee149e535b14

Observation 7556c271-7c8a-4958-8cc7-d183c8d72026 · inbound

Policy Contrastive Decoding for Robotic Foundation Models cites this paper.

Policy Contrastive Decoding for Robotic Foundation Models DINOv2: Learning Robust Visual Features without Supervision

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-22T14:11:38.551674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T14:09:48.762737Z digest=sha256:5106c1e51ddc0ed115a9158b7cc65e65965baa7f64ab877fd68333042ae568f1

Observation 34967fdb-6d6f-45bc-9e83-b086bbcb7a24 · inbound

FractalMamba++: Scaling Vision Mamba Across Resolutions via Hilbert Fractal Geometry cites this paper.

FractalMamba++: Scaling Vision Mamba Across Resolutions via Hilbert Fractal Geometry DINOv2: Learning Robust Visual Features without Supervision

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-22T14:01:38.196240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T13:58:28.333444Z digest=sha256:b2664dba5d6ad8cfb12360f3e923e7f2fc53330582b879af667a477ccff2661f

Observation d16b66a9-182f-40ab-bfcd-13469f92fcde · inbound

Interactive Post-Training for Vision-Language-Action Models cites this paper.

Interactive Post-Training for Vision-Language-Action Models DINOv2: Learning Robust Visual Features without Supervision

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-21T14:25:47.240959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T14:25:47.178714Z digest=sha256:601c36399c413a5346a61aed6b9aecd682fa251e66fe4cd6c7440ae492abfa68

Observation a65220d4-64fe-4246-b377-f6a8418344d8 · inbound

Slot-MLLM: Object-Centric Visual Tokenization for Multimodal LLM cites this paper.

Slot-MLLM: Object-Centric Visual Tokenization for Multimodal LLM DINOv2: Learning Robust Visual Features without Supervision

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-05-22T02:10:56.101559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T02:06:35.204166Z digest=sha256:909ac276b8baf6e52aeee3649a371f40f61289d96b9934920e8f6087dfeb7208

Observation d50dc313-3588-4bc8-ba8c-80acc68977f3 · inbound

VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning cites this paper.

VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning DINOv2: Learning Robust Visual Features without Supervision

Reference 55

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T12:55:40.490754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T12:55:40.245908Z digest=sha256:18986e2e90782da50d1e85b924809828b0df1bcf6f2eba1f154e199712a059a4

Observation 3a85a4d7-8d77-4fbb-99b8-a7a6480fdf5f · inbound

ImgEdit: A Unified Image Editing Dataset and Benchmark cites this paper.

ImgEdit: A Unified Image Editing Dataset and Benchmark DINOv2: Learning Robust Visual Features without Supervision

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-05-12T18:17:45.332209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T18:17:45.123690Z digest=sha256:33d4523aa7b7b7fa1148d625ead88be06d457b21eec43f1132019276f2dfcff1

Observation df892426-61d1-41f0-a5ea-88a4d230b471 · inbound

Geometry-Editable and Appearance-Preserving Object Compositon cites this paper.

Geometry-Editable and Appearance-Preserving Object Compositon DINOv2: Learning Robust Visual Features without Supervision

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-22T02:14:31.264963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T02:13:56.792400Z digest=sha256:0eabc95869733e177becb8b6516253d22c3192cf02698ce0765a4f80e36e7a81

Observation e56d951b-5ba2-40db-8de8-f45ba64782f0 · inbound

A European Multi-Center Breast Cancer MRI Dataset cites this paper.

A European Multi-Center Breast Cancer MRI Dataset DINOv2: Learning Robust Visual Features without Supervision

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-25T08:25:33.859184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-25T08:20:47.972916Z digest=sha256:fb06e07b6e0122ec55f7019e6d9181909ebca28ca5a1652faeaaf0c86d85d578

Observation 6ba41af7-6b24-486f-93a4-085ee0cbc71a · inbound

AuralSAM2: Enabling SAM2 Hear Through Pyramid Audio-Visual Feature Prompting cites this paper.

AuralSAM2: Enabling SAM2 Hear Through Pyramid Audio-Visual Feature Prompting DINOv2: Learning Robust Visual Features without Supervision

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-05-19T11:17:15.503446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T11:14:29.394752Z digest=sha256:28eb3912a9605e162e9360d82ce81748753681075a5518168ca585022d6d6357

Observation 276b12ab-a255-4faf-b8fa-e263f013080b · inbound

Beyond Interpretability: When, Why, and How Sparse Autoencoders Enable Label-Free Visual Steering cites this paper.

Beyond Interpretability: When, Why, and How Sparse Autoencoders Enable Label-Free Visual Steering DINOv2: Learning Robust Visual Features without Supervision

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-05-19T11:37:15.838022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T11:35:01.836096Z digest=sha256:e3e715939ffc912750cc0c0141b21c71783142002329e4d3a49d0329ee46a925

Observation 396c38b5-9580-44d3-b0d8-77bc02935d8e · inbound

UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation cites this paper.

UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation DINOv2: Learning Robust Visual Features without Supervision

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-12T17:34:27.176396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T17:34:26.951644Z digest=sha256:52267512228277c1d05309f95e573f196fa39a61cd49694aced99b027044e1e6

Observation 7664a232-90a5-4b13-a6f7-17530da8d901 · inbound

Adapting Vision-Language Foundation Model for Next Generation Medical Ultrasound Image Analysis cites this paper.

Adapting Vision-Language Foundation Model for Next Generation Medical Ultrasound Image Analysis DINOv2: Learning Robust Visual Features without Supervision

Reference 24

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T10:37:15.132106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T10:34:28.875334Z digest=sha256:fd5e607cf00ead51594e3745cfe6f58c58876fc8d4ae81e8bb5a030057bef886

Observation 2bda31ff-adde-4223-ae84-e3c5287a7ff7 · inbound

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models cites this paper.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models DINOv2: Learning Robust Visual Features without Supervision

Reference 71

Resolution
verified exact
local_arxiv, observed 2026-05-19T11:13:02.898494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:874ec7b86b5ce1a675bc6097c5f0e78582b383ec3e2b92d660438c708af2176f

Observation 433bc99e-583f-4e76-9b53-8f9b6dc37c0c · inbound

The Less You Depend, The More You Learn: Synthesizing Novel Views from Sparse, Unposed Images with Minimal 3D Knowledge cites this paper.

The Less You Depend, The More You Learn: Synthesizing Novel Views from Sparse, Unposed Images with Minimal 3D Knowledge DINOv2: Learning Robust Visual Features without Supervision

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-19T09:33:02.591664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T09:32:48.787211Z digest=sha256:dbca6423a338ff706632609bd97147d2b0b21a0a7173ac8a06fc9b17c3d0ebab

Observation 57365b87-d7e6-4990-aca8-c82a929eeb0a · inbound

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning cites this paper.

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning DINOv2: Learning Robust Visual Features without Supervision

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-05-11T00:33:51.136489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T00:33:50.471804Z digest=sha256:b820fd5c81e9ee1c5ba921d8ae410bcf8a8ee97332a9e78fbf9d4d5c47b1fc61

Observation 72851fcb-8c27-4421-a253-0c85d09b6649 · inbound

CLIP the Landscape: Automated Tagging of Crowdsourced Landscape Images cites this paper.

CLIP the Landscape: Automated Tagging of Crowdsourced Landscape Images DINOv2: Learning Robust Visual Features without Supervision

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-19T09:03:02.229474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T09:02:18.576084Z digest=sha256:f609fad97a9b2e466d396ad19f77eff8ff38a96f039a05cc55d4afcc94d42f02

Observation efb1dbd7-41fa-4565-9923-076a759c2bbb · inbound

Hunyuan3D 2.1: From Images to High-Fidelity 3D Assets with Production-Ready PBR Material cites this paper.

Hunyuan3D 2.1: From Images to High-Fidelity 3D Assets with Production-Ready PBR Material DINOv2: Learning Robust Visual Features without Supervision

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-17T23:10:17.727980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T23:10:17.659621Z digest=sha256:6d689c7797deaca960b02e01e6ea760f198aecabdb2c84ccf9bba38240125c4e

Observation 78d241c5-bccb-4740-99ee-1162d93747af · inbound

OmniGen2: Towards Instruction-Aligned Multimodal Generation cites this paper.

OmniGen2: Towards Instruction-Aligned Multimodal Generation DINOv2: Learning Robust Visual Features without Supervision

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-05-19T07:52:10.794403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:795e2e27f3430b425a7aae8c48047d2e1970232518011766022cb3baf85352fd

Observation c4e4a262-85db-4ee2-8300-f1c4f750e3ca · inbound

GenHSI: Controllable Generation of Human-Scene Interaction Videos cites this paper.

GenHSI: Controllable Generation of Human-Scene Interaction Videos DINOv2: Learning Robust Visual Features without Supervision

Reference 66

Resolution
verified exact
local_arxiv, observed 2026-05-19T07:27:08.873264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:3397834efae6f9e1d01d9ab3fb7973e9b29b09a400736a33de2662bcfc5f9181

Observation a30cfea2-24dd-4de9-b4c3-379cc83011a6 · inbound

MoGe-2: Accurate Monocular Geometry with Metric Scale and Sharp Details cites this paper.

MoGe-2: Accurate Monocular Geometry with Metric Scale and Sharp Details DINOv2: Learning Robust Visual Features without Supervision

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-05-14T21:19:44.196479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T21:19:44.144633Z digest=sha256:bdd854a6b034c5450efca4e852d616ef765285c2b1fd9425d45b6de6da070f1e

Observation d8de1717-616a-4f19-986b-1fd583b7780c · inbound

Perception-Aware Policy Optimization for Multimodal Reasoning cites this paper.

Perception-Aware Policy Optimization for Multimodal Reasoning DINOv2: Learning Robust Visual Features without Supervision

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-19T05:12:04.866205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-19T05:11:54.685897Z digest=sha256:198e2b7352c9255fba5d63875b822285d5c19702cfcde2571ea112809ff7156f

Observation 6c6d406f-35f3-480c-946f-bd2dde97eb95 · inbound

SCOOTER: A Human Evaluation Framework for Unrestricted Adversarial Examples cites this paper.

SCOOTER: A Human Evaluation Framework for Unrestricted Adversarial Examples DINOv2: Learning Robust Visual Features without Supervision

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-05-19T05:37:05.574267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-19T05:35:16.138603Z digest=sha256:1d064430b2d046703209f9f5abe6110d477b694bd731b24501177212f0300080

Observation bfd871df-9dab-4102-98ff-45d0d4354cc0 · inbound

Geometry Forcing: Marrying Video Diffusion and 3D Representation for Consistent World Modeling cites this paper.

Geometry Forcing: Marrying Video Diffusion and 3D Representation for Consistent World Modeling DINOv2: Learning Robust Visual Features without Supervision

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-05-19T05:17:06.706444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-19T05:13:28.767788Z digest=sha256:2cf1cb41b4c71d417c70a68cc0f34f968777295b6bad3b783a85eba91fa4c091

Observation a4350520-e548-4337-82dc-7a2d069ce964 · inbound

Navigating the Challenges of AI-Generated Image Detection in the Wild: What Truly Matters? cites this paper.

Navigating the Challenges of AI-Generated Image Detection in the Wild: What Truly Matters? DINOv2: Learning Robust Visual Features without Supervision

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-21T23:34:26.493966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T23:31:40.691896Z digest=sha256:1211b85e08c9cc48b9c73857333e2d27a0b163e9f96036032ecf02cc0e8acfa1

Observation ddd65d56-4e3d-437a-af53-33287ad36be9 · inbound

Streaming 4D Visual Geometry Transformer cites this paper.

Streaming 4D Visual Geometry Transformer DINOv2: Learning Robust Visual Features without Supervision

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-15T22:58:18.995541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T22:58:18.929748Z digest=sha256:672116dbc1040e3f0b20fce295ec8b580ee2bce4946ed17666ac636731832cf2

Observation 66191822-d84f-4aad-baa4-d435340c8d80 · inbound

$\pi^3$: Permutation-Equivariant Visual Geometry Learning cites this paper.

$\pi^3$: Permutation-Equivariant Visual Geometry Learning DINOv2: Learning Robust Visual Features without Supervision

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-12T17:14:22.803718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T17:14:22.678806Z digest=sha256:222eb7f66d63eee3d7fcf0c23c77d9abae69de5dfa5a910b76a1249da09565da

Observation 3ba6a66d-6633-4545-bcdc-788539450d06 · inbound

Frozen Forecasting: A Unified Evaluation cites this paper.

Frozen Forecasting: A Unified Evaluation DINOv2: Learning Robust Visual Features without Supervision

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-19T03:42:57.290288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T03:42:54.620069Z digest=sha256:146861866a29f89426dd301d4c482c0b074faaa978a843efb1b9cd5057edec12

Observation 42ee7a51-17f5-44c2-a573-e9ac9fc8f509 · inbound

ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent Planning cites this paper.

ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent Planning DINOv2: Learning Robust Visual Features without Supervision

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-19T03:22:01.025538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T03:18:14.655384Z digest=sha256:6a0e35f1693719b80deec0a5f7dde9328d88102121cc1336007b4cd1211f37e4

Observation fa37519e-3991-459a-b01e-5b462ab3d50d · inbound

AI in Agriculture: A Survey of Deep Learning Techniques for Crops, Fisheries and Livestock cites this paper.

AI in Agriculture: A Survey of Deep Learning Techniques for Crops, Fisheries and Livestock DINOv2: Learning Robust Visual Features without Supervision

Reference 151

Resolution
verified exact
local_arxiv, observed 2026-05-19T02:06:58.726240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T02:03:46.331803Z digest=sha256:0eae9e27df58812e594376c0a13da11afde0288178c6a41c95290fc2ffbaadd5

Observation 92f37ce5-ae18-4595-b74b-214ac051d98c · inbound

Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation cites this paper.

Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation DINOv2: Learning Robust Visual Features without Supervision

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-15T21:28:42.049871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T21:28:41.904725Z digest=sha256:3c9763dddfa7224b47035eef46a465a83cc0a1cca351776c64140e8d71466176

Observation 80a9f126-b12e-4b48-a13a-770070cccaf7 · inbound

IntrinsicWeather: Controllable Weather Editing in Intrinsic Space cites this paper.

IntrinsicWeather: Controllable Weather Editing in Intrinsic Space DINOv2: Learning Robust Visual Features without Supervision

Reference 13

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T00:01:56.029563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T23:58:56.189814Z digest=sha256:614ce73245919b0183bbf72922f5f2de9fc25512439b4c9215cf38d92b0025eb

Observation 3853af0a-42cc-4146-b354-2c67a8ed117d · inbound

MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation cites this paper.

MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation DINOv2: Learning Robust Visual Features without Supervision

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-15T20:43:24.517101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T20:43:24.417901Z digest=sha256:cd25cf9e8ee49ae18fc60eadbfb8b9c0d2d6618659a9fe4de9369b8f9823a414

Observation 370fa0db-f45b-47ec-a1bf-8106d6fd89f3 · inbound

Dino U-Net: Exploiting High-Fidelity Dense Features from Foundation Models for Medical Image Segmentation cites this paper.

Dino U-Net: Exploiting High-Fidelity Dense Features from Foundation Models for Medical Image Segmentation DINOv2: Learning Robust Visual Features without Supervision

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-18T20:22:50.679850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T20:22:41.555806Z digest=sha256:8cd273a603a9aa1c7a55d99c5699b5ca5d24bde0a20b983437332c005009d56e

Observation 3382f296-6469-4f09-a436-c37ae4d106a7 · inbound

FastVGGT: Training-Free Acceleration of Visual Geometry Transformer cites this paper.

FastVGGT: Training-Free Acceleration of Visual Geometry Transformer DINOv2: Learning Robust Visual Features without Supervision

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-15T23:36:06.954412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T23:36:06.870759Z digest=sha256:1fc414a8d8c56aacce031b1b94ffc7241693ca46a3389986db838d50f7c2bf41

Observation 16c39280-a7d5-4a04-8cf1-a25cab76ea5b · inbound

LUIVITON: Learned Universal Interoperable VIrtual Try-ON cites this paper.

LUIVITON: Learned Universal Interoperable VIrtual Try-ON DINOv2: Learning Robust Visual Features without Supervision

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-21T22:00:41.335957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T22:00:15.259709Z digest=sha256:4e1fa00041ffc1769a273b7ce0381462f552d0f4fbd8a75ee216bcc8f7923eff

Observation 151f8100-3a5c-43fd-a11f-d3ff3ad44984 · inbound

Block-Sparse Global Attention for Efficient Multi-View Geometry Transformers cites this paper.

Block-Sparse Global Attention for Efficient Multi-View Geometry Transformers DINOv2: Learning Robust Visual Features without Supervision

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-21T22:24:23.574993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T22:22:23.320347Z digest=sha256:9c7f059261edb6b7729f044bb50b7611ac553f3b2e0ab751aa184d6b3641a749

Observation 48747c6e-fdcd-4b0e-80cd-33aed443f89f · inbound

One View, Many Worlds: Single-Image to 3D Object Meets Generative Domain Randomization for One-Shot 6D Pose Estimation cites this paper.

One View, Many Worlds: Single-Image to 3D Object Meets Generative Domain Randomization for One-Shot 6D Pose Estimation DINOv2: Learning Robust Visual Features without Supervision

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-04T21:31:08.235510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:31:08.235510Z digest=sha256:935486b6f177baca45847c44b57fda580ccde2a2a47d7cf230fc9621e29bc45a

Observation 18868781-093f-4806-b8ec-b558adfd1a18 · inbound

Domain Knowledge is Power: Leveraging Physiological Priors for Self Supervised Representation Learning in Electrocardiography cites this paper.

Domain Knowledge is Power: Leveraging Physiological Priors for Self Supervised Representation Learning in Electrocardiography DINOv2: Learning Robust Visual Features without Supervision

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T21:20:46.279827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:20:46.279827Z digest=sha256:3f15644b3690e663b83a496663064d8975eb7964d6295815b125d6506081adc0

Observation 3b690062-15d8-4053-9c37-476f70e7730b · inbound

Foundation Models for Autonomous Driving Perception: A Survey Through Core Capabilities cites this paper.

Foundation Models for Autonomous Driving Perception: A Survey Through Core Capabilities DINOv2: Learning Robust Visual Features without Supervision

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T20:52:06.161060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:52:06.161060Z digest=sha256:adcbfe8d33fcb00f4eda01a292643672e04a366dc4f0471a45844378f75a3b84

Observation e81d0d67-8d36-43ac-acb7-b7fd53065b58 · inbound

ViewSparsifier: Killing Redundancy in Multi-View Plant Phenotyping cites this paper.

ViewSparsifier: Killing Redundancy in Multi-View Plant Phenotyping DINOv2: Learning Robust Visual Features without Supervision

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T20:29:46.119905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:29:46.119905Z digest=sha256:c8f327f7a5cf503b3a31fe860235c23f14dbd4b207e2c918781a20c991c072e3

Observation c06163f8-3aff-44a6-b1ed-7e820b97f396 · inbound

RoboChemist: Long-Horizon and Safety-Compliant Robotic Chemical Experimentation cites this paper.

RoboChemist: Long-Horizon and Safety-Compliant Robotic Chemical Experimentation DINOv2: Learning Robust Visual Features without Supervision

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-04T20:10:13.655168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:10:13.655168Z digest=sha256:3e929a21bf0cf3fe00a144b2676acd34f4c8d3169932d6884c2bc8faf899118b

Observation 30cfddf2-47f1-4eb0-9d19-1375bfe987fb · inbound

Mind Meets Space: Rethinking Agentic Spatial Intelligence from a Neuroscience-inspired Perspective cites this paper.

Mind Meets Space: Rethinking Agentic Spatial Intelligence from a Neuroscience-inspired Perspective DINOv2: Learning Robust Visual Features without Supervision

Reference 150

Resolution
unresolved
no resolver link, observed 2026-08-04T19:39:13.303201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:39:13.303201Z digest=sha256:e4c209d686d00edf4861fd52ab48e9372ce414bd469ee535f8e65ea6ac404434

Observation bf5f5831-113e-4ebc-be0c-58225de6a9f9 · inbound

Learning Object-Centric Representations in SAR Images with Multi-Level Feature Fusion cites this paper.

Learning Object-Centric Representations in SAR Images with Multi-Level Feature Fusion DINOv2: Learning Robust Visual Features without Supervision

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-04T19:24:39.894167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:24:39.894167Z digest=sha256:1ce5626a29543ca8f3233dc19fcd3075e0993488778b8cb990e63931093e3c11

Observation f0fcf758-83ae-4563-850a-aaa7e5619bee · inbound

Image Recognition with Vision and Language Embeddings of VLMs cites this paper.

Image Recognition with Vision and Language Embeddings of VLMs DINOv2: Learning Robust Visual Features without Supervision

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T19:24:21.482265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:24:21.482265Z digest=sha256:406a2c7a874088e17058b51492ac79bb70d812a2574850578ffbc62d381b5c00

Observation 074847cd-7d7e-4632-90ab-04f6897547da · inbound

Unsupervised Integrated-Circuit Defect Segmentation via Image-Intrinsic Normality cites this paper.

Unsupervised Integrated-Circuit Defect Segmentation via Image-Intrinsic Normality DINOv2: Learning Robust Visual Features without Supervision

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T19:14:15.636242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:14:15.636242Z digest=sha256:236387b12affb166401d81032223adb1ff5db6e166c1bf434d4e58f7183c96ed

Observation b7a1aacb-ddf0-4850-9ced-6162638e0d3f · inbound

MimicDroid: In-Context Learning for Humanoid Robot Manipulation from Human Play Videos cites this paper.

MimicDroid: In-Context Learning for Humanoid Robot Manipulation from Human Play Videos DINOv2: Learning Robust Visual Features without Supervision

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-04T18:51:36.111093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:51:36.111093Z digest=sha256:899d062e665282da6f0403469aa36a39ec17f5a4b1244fd3e986fc198ef02d21

Observation f3708538-b30e-4440-a28d-55cdbdf4754c · inbound

Self-supervised Learning Of Visual Pose Estimation Without Pose Labels By Classifying LED States cites this paper.

Self-supervised Learning Of Visual Pose Estimation Without Pose Labels By Classifying LED States DINOv2: Learning Robust Visual Features without Supervision

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T17:59:14.758297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:59:14.758297Z digest=sha256:dd74f3ccc543dcc9b164175ba27df74eab9f82eaf98dd44fa8f2d809b900d9a4

Observation 1ba4bd1b-cc4e-4b7b-ab2e-342bd5ebda52 · inbound

ImMimic: Cross-Domain Imitation from Human Videos via Mapping and Interpolation cites this paper.

ImMimic: Cross-Domain Imitation from Human Videos via Mapping and Interpolation DINOv2: Learning Robust Visual Features without Supervision

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-04T17:26:19.669625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:26:19.669625Z digest=sha256:8b064a295d3a36985222947711cbf61f364d729240be0c4985fd8838a75ebb0d

Observation 7a283243-ebdd-43af-8725-57af9b3aa62b · inbound

RAPTOR: A Foundation Policy for Quadrotor Control cites this paper.

RAPTOR: A Foundation Policy for Quadrotor Control DINOv2: Learning Robust Visual Features without Supervision

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-05-18T17:31:41.749204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T17:28:32.778343Z digest=sha256:a979b0751e8e5a83a276a9a9999c14571a1a2662f8489d99744b9735bcc499a9