Pith. sign in

Paper Citation Record · LEDGER

Deep ViT Features as Dense Visual Descriptors

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 53 inbound Pith citation observations for arXiv:2112.05814.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2112.05814 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 53 of 53 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:33:34.922510Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-09T16:56:21.644793Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation eed0412e-546b-46d5-9640-b6a614a8de4a · inbound

ReKep: Spatio-Temporal Reasoning of Relational Keypoint Constraints for Robotic Manipulation cites this paper.

ReKep: Spatio-Temporal Reasoning of Relational Keypoint Constraints for Robotic Manipulation Deep ViT Features as Dense Visual Descriptors

Reference 119

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:25:18.019787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-16T08:25:17.847571Z digest=sha256:b9e8c55c49d93364445ca5ea7c87070127fe49e49378b56fe071854e72524824

Observation e06842bc-d1ce-49b8-a48d-d911c2f88299 · inbound

UrbanCAD: Towards Highly Controllable and Photorealistic 3D Vehicles for Urban Scene Simulation cites this paper.

UrbanCAD: Towards Highly Controllable and Photorealistic 3D Vehicles for Urban Scene Simulation Deep ViT Features as Dense Visual Descriptors

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T10:23:38.698760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:23:38.698760Z digest=sha256:7f8d4ad83e99622dc3463eadbaf630bc4c1bae84af4e89c2ec6a61b9bc0e291c

Observation 084c799b-533d-4215-9c35-03011c8fda8e · inbound

Multiview Equivariance Improves 3D Correspondence Understanding with Minimal Feature Finetuning cites this paper.

Multiview Equivariance Improves 3D Correspondence Understanding with Minimal Feature Finetuning Deep ViT Features as Dense Visual Descriptors

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T10:13:13.447417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:13:13.447417Z digest=sha256:bbd58aa2b5ee8092665232869d6a874955afe997d02cac7488fcb93059b4df5a

Observation 5562cb9f-ff28-4e93-a2e2-4860139fcf6e · inbound

Categorical Keypoint Positional Embedding for Robust Animal Re-Identification cites this paper.

Categorical Keypoint Positional Embedding for Robust Animal Re-Identification Deep ViT Features as Dense Visual Descriptors

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T05:02:50.981646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:02:50.981646Z digest=sha256:424ccbe7c8dabd244d6db6abd46dcab7c37b6a8c2a4f68b59bb37fa5ac4e8a53

Observation 7121af3f-0cea-4234-95e8-efbc59260239 · inbound

MAGMA: Manifold Regularization for MAEs cites this paper.

MAGMA: Manifold Regularization for MAEs Deep ViT Features as Dense Visual Descriptors

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:07.812734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:05:07.812734Z digest=sha256:ade8fbae605b09fe630c2fb5292a431ade095e5759db215c899919b15600c086

Observation 1195cb44-1cec-402d-bd32-1dea98310b8d · inbound

DIVE: Taming DINO for Subject-Driven Video Editing cites this paper.

DIVE: Taming DINO for Subject-Driven Video Editing Deep ViT Features as Dense Visual Descriptors

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T22:33:36.253762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:33:36.253762Z digest=sha256:cca620009c89feb82c5ebb5783e9dcd3c13f40758f6510c407997ee8eb4fc099

Observation be09a5fb-5527-4516-b6a1-5649e8772b4b · inbound

Distillation of Diffusion Features for Semantic Correspondence cites this paper.

Distillation of Diffusion Features for Semantic Correspondence Deep ViT Features as Dense Visual Descriptors

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T22:24:43.904338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:24:43.904338Z digest=sha256:fd9c0d1c40c3cb4ba0f22558a3d8d2a96add5c681d4707dd3eb45056b03da296

Observation 87e8edc2-f6c6-417a-9e77-265297dcc0ef · inbound

DenseMatcher: Learning 3D Semantic Correspondence for Category-Level Manipulation from a Single Demo cites this paper.

DenseMatcher: Learning 3D Semantic Correspondence for Category-Level Manipulation from a Single Demo Deep ViT Features as Dense Visual Descriptors

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T20:57:12.265574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:57:12.265574Z digest=sha256:964e75e7ee39fd1877fc9e51341108c2c42228b1e33bb60c9298cd993ee9662d

Observation b2df234b-f141-4692-a438-0e543ad55555 · inbound

Track4Gen: Teaching Video Diffusion Models to Track Points Improves Video Generation cites this paper.

Track4Gen: Teaching Video Diffusion Models to Track Points Improves Video Generation Deep ViT Features as Dense Visual Descriptors

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T20:10:47.357353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:10:47.357353Z digest=sha256:279893f04306756c9ee3e777254e1229bb87470cb9c1b209c8aee3d7fc331088

Observation 650dad2a-7c59-4daa-96a4-1cc9ad09630b · inbound

DrivingRecon: Large 4D Gaussian Reconstruction Model For Autonomous Driving cites this paper.

DrivingRecon: Large 4D Gaussian Reconstruction Model For Autonomous Driving Deep ViT Features as Dense Visual Descriptors

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-11T17:26:28.950433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:26:28.950433Z digest=sha256:a3a25bdd7895d039aa40bcc3e45aff5b891890d969956ad015f6849ff670d6ab

Observation a4ab7d16-9948-4a40-947d-d398aa7a4468 · inbound

Few-Shot Adaptation of Training-Free Foundation Model for 3D Medical Image Segmentation cites this paper.

Few-Shot Adaptation of Training-Free Foundation Model for 3D Medical Image Segmentation Deep ViT Features as Dense Visual Descriptors

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T20:15:26.470848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:15:26.470848Z digest=sha256:19b59d1684b2aaace6531d482f5b24af88e31fd0adeca3a0815cd1222c6eac9d

Observation ed917cb5-7e74-4cfa-93ef-19968c800fc9 · inbound

Surface-SOS: Self-Supervised Object Segmentation via Neural Surface Representation cites this paper.

Surface-SOS: Self-Supervised Object Segmentation via Neural Surface Representation Deep ViT Features as Dense Visual Descriptors

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T19:34:48.806833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:34:48.806833Z digest=sha256:a09a156cbf8d286771e4d78286bbc1973e50cca570fad633b24513239a33d3bd

Observation f1fb5de3-52e3-4e71-b754-9d14821a8089 · inbound

Exploring Temporally-Aware Features for Point Tracking cites this paper.

Exploring Temporally-Aware Features for Point Tracking Deep ViT Features as Dense Visual Descriptors

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T17:25:52.804922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:25:52.804922Z digest=sha256:3bd4e5f64de9f55ce32bde9822abd55682e411b272835817f0bf780de1804769

Observation cc12395c-daee-41de-a5f2-20262dae3c2d · inbound

Automated Measurement of Eczema Severity with Self-Supervised Learning cites this paper.

Automated Measurement of Eczema Severity with Self-Supervised Learning Deep ViT Features as Dense Visual Descriptors

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:34.922510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:33:34.922510Z digest=sha256:9a069d05a3d5a470cda583ca273a21000b87ad3867f96818d1617c4d263c6206

Observation 5874fc14-64dd-4796-abb1-018cd06aacc5 · inbound

Robotic Task Ambiguity Resolution via Natural Language Interaction cites this paper.

Robotic Task Ambiguity Resolution via Natural Language Interaction Deep ViT Features as Dense Visual Descriptors

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T10:36:59.065795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:36:59.065795Z digest=sha256:36e41c722062175d5c6f6096bad2275caa892e5d71e59ba9c141d728aa1a59b9

Observation 510a734a-66eb-45c7-85e8-bb482d4b0330 · inbound

Grounded Task Axes: Zero-Shot Semantic Skill Generalization via Task-Axis Controllers and Visual Foundation Models cites this paper.

Grounded Task Axes: Zero-Shot Semantic Skill Generalization via Task-Axis Controllers and Visual Foundation Models Deep ViT Features as Dense Visual Descriptors

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T20:53:25.664150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:53:25.664150Z digest=sha256:52b3245f0da91fff241f40f32129b0dbc8dede2657d7c89983559f348442a6d9

Observation 142abc24-a300-4ae2-83f6-371f4ead5fb8 · inbound

Guiding Diffusion with Deep Geometric Moments: Balancing Fidelity and Variation cites this paper.

Guiding Diffusion with Deep Geometric Moments: Balancing Fidelity and Variation Deep ViT Features as Dense Visual Descriptors

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T20:36:26.261952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:36:26.261952Z digest=sha256:82037faaecd6f039e26a14f1ce0b48ed582179a06ba51d2f361a5cb652ea0fb9

Observation 0c4c77bd-525f-4379-b8ac-ad7ba49c8b22 · inbound

Object-level Self-Distillation for Vision Pretraining cites this paper.

Object-level Self-Distillation for Vision Pretraining Deep ViT Features as Dense Visual Descriptors

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:54.657235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:52:54.657235Z digest=sha256:e7f995eafe1103a6dff38016e4fc19ffd3ba468cc0dd458b09e69703f12f9d15

Observation 60d41486-0e03-4083-8dd7-6de7fc5fc742 · inbound

ADAM: Autonomous Discovery and Annotation Model using LLMs for Context-Aware Annotations cites this paper.

ADAM: Autonomous Discovery and Annotation Model using LLMs for Context-Aware Annotations Deep ViT Features as Dense Visual Descriptors

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:11.372001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:11.372001Z digest=sha256:c394be2f2d32f83bf45189fb74516361e3d86476a867bb95e6a0c00f4924eebe

Observation 928be4f3-b957-4f07-9abc-f667ba638668 · inbound

Online Long-term Point Tracking in the Foundation Model Era cites this paper.

Online Long-term Point Tracking in the Foundation Model Era Deep ViT Features as Dense Visual Descriptors

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T18:06:17.431860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:06:17.431860Z digest=sha256:f79441ec17c7990519643d6b117f98dbbb0b8d56aa897c6aa3d17c4c0e3624fb

Observation e889d9bb-59c6-4de3-859a-6e4333d74a5e · inbound

PCR-GS: COLMAP-Free 3D Gaussian Splatting via Pose Co-Regularizations cites this paper.

PCR-GS: COLMAP-Free 3D Gaussian Splatting via Pose Co-Regularizations Deep ViT Features as Dense Visual Descriptors

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T16:18:24.841321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:18:24.841321Z digest=sha256:130fc431f22e79fa5bfe7d9494528b76816626c899a829712b39e4ebc80c36f3

Observation bbb79a4f-d31c-4cea-b127-90f5b0070a05 · inbound

MotionShot: Adaptive Motion Transfer across Arbitrary Objects for Text-to-Video Generation cites this paper.

MotionShot: Adaptive Motion Transfer across Arbitrary Objects for Text-to-Video Generation Deep ViT Features as Dense Visual Descriptors

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T15:20:00.044930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:20:00.044930Z digest=sha256:7dac59b01c86a567dbae0662ad607428617f659f8992bd8cdb55394611014693

Observation c75eca5b-3add-4908-a582-fc3522310408 · inbound

Unified Category-Level Object Detection and Pose Estimation from RGB Images using 3D Prototypes cites this paper.

Unified Category-Level Object Detection and Pose Estimation from RGB Images using 3D Prototypes Deep ViT Features as Dense Visual Descriptors

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T05:14:35.096273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:14:35.096273Z digest=sha256:bd94dc9cf68170ffcda1fdee6265e8e6246515efbefd5866d73c9f6913a0b8bc

Observation 0fd8203c-7202-4fef-bfcb-f36ae3b3f04a · inbound

No Masks Needed: Explainable AI for Deriving Segmentation from Classification cites this paper.

No Masks Needed: Explainable AI for Deriving Segmentation from Classification Deep ViT Features as Dense Visual Descriptors

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T00:01:56.587398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:01:56.587398Z digest=sha256:14139f9e2a06d3b7e225dc3157965dfd1784b9eee7e270dcbb133874a7018f78

Observation 71b76815-404a-4829-8cc6-d66127a07f2b · inbound

Structure-Preserving Medical Image Generation from a Latent Graph Representation cites this paper.

Structure-Preserving Medical Image Generation from a Latent Graph Representation Deep ViT Features as Dense Visual Descriptors

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-05T17:47:01.461255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:47:01.461255Z digest=sha256:62c537a3615bc75f7f68df8e62e29a44577005a70eaf78d244884262e0221879

Observation f49e918f-35f5-4eb9-996b-99abc183dc10 · inbound

Generalizable Object Re-Identification via Visual In-Context Prompting cites this paper.

Generalizable Object Re-Identification via Visual In-Context Prompting Deep ViT Features as Dense Visual Descriptors

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T14:33:20.462407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:33:20.462407Z digest=sha256:ca417ae6289ec8e327f05b9585c38975fb69609f395997c95965b3ea3b4a51dd

Observation f36cbf6a-aa6b-4343-89c2-ea0305796b62 · inbound

Maybe you don't need a U-Net: convolutional feature upsampling for materials micrograph segmentation cites this paper.

Maybe you don't need a U-Net: convolutional feature upsampling for materials micrograph segmentation Deep ViT Features as Dense Visual Descriptors

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T16:44:21.050130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:44:21.050130Z digest=sha256:a005d366f0d4455ed8a207574862beaab6227510c530b04b2866019bf4a56755

Observation 9dda9dd1-18d5-4c1b-b42f-e0fba1e583e1 · inbound

Weakly-Supervised Learning of Dense Functional Correspondences cites this paper.

Weakly-Supervised Learning of Dense Functional Correspondences Deep ViT Features as Dense Visual Descriptors

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T10:37:51.275765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:37:51.275765Z digest=sha256:c2e2137655d723a2328fbba602922f57fa2a5274b729402f715b61469c9c7a87

Observation 8d189666-5c52-47ae-af0f-783011e01d6d · inbound

Robotic Manipulation Framework Based on Semantic Keypoints for Packing Shoes of Different Sizes, Shapes, and Softness cites this paper.

Robotic Manipulation Framework Based on Semantic Keypoints for Packing Shoes of Different Sizes, Shapes, and Softness Deep ViT Features as Dense Visual Descriptors

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T04:36:20.760224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:36:20.760224Z digest=sha256:d4cf3227ebe58ba38884967c2721285d79eb42cef68d0d2b1aa1bd8454d2bc6f

Observation 41863813-330f-4116-8225-29db3a6c550b · inbound

From One-to-One to Many-to-Many: Dynamic Cross-Layer Injection for Deep Vision-Language Fusion cites this paper.

From One-to-One to Many-to-Many: Dynamic Cross-Layer Injection for Deep Vision-Language Fusion Deep ViT Features as Dense Visual Descriptors

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T10:18:31.361765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:18:31.361765Z digest=sha256:47d01230fb0ee29e489e5dadd0c3e81a82258596b6a39a47218f96f461681fac

Observation d9fc5d41-16d8-4161-9084-f14b8de71451 · inbound

UPLiFT: Efficient Pixel-Dense Feature Upsampling with Local Attenders cites this paper.

UPLiFT: Efficient Pixel-Dense Feature Upsampling with Local Attenders Deep ViT Features as Dense Visual Descriptors

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T08:13:39.247257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:13:39.247257Z digest=sha256:d9c026c8ebea6429c9d723ce0bdbe4786e075864ab359af10419b9c82e1a920c

Observation 504b71b5-f443-4e31-8795-55a7245c90a3 · inbound

Uncertainty-Aware Hierarchical Re-Localization in OpenStreetMap via Semantic Alignment cites this paper.

Uncertainty-Aware Hierarchical Re-Localization in OpenStreetMap via Semantic Alignment Deep ViT Features as Dense Visual Descriptors

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-02T19:36:19.815385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:36:19.815385Z digest=sha256:9bb5cf75cd2009fe44e863c72b1db7d62c2e19bf556a2bad9e497590e4dc5bdd

Observation e772a24d-ec4d-4ec9-ad39-298326d703c3 · inbound

Autonomous Search for Sparsely Distributed Visual Phenomena through Environmental Context Modeling cites this paper.

Autonomous Search for Sparsely Distributed Visual Phenomena through Environmental Context Modeling Deep ViT Features as Dense Visual Descriptors

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-14T23:50:28.153875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:50:28.153875Z digest=sha256:20a8f0438aa3c6d84ef1d916f2ec02395e13f5c1d04f887d8a60a692a7bdbd06

Observation 55c0e98a-11cd-4bf4-a8cb-76e93dbdadad · inbound

Human-like Object Grouping in Self-supervised Vision Transformers cites this paper.

Human-like Object Grouping in Self-supervised Vision Transformers Deep ViT Features as Dense Visual Descriptors

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-14T21:34:48.709465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T21:34:48.709465Z digest=sha256:2c050e8fb8e1004f9b5700446ca5b095e26bb5455b88492d714a7c6f0f879918

Observation 47c61a5c-b9c4-4ad8-aa62-0627a4ec43a3 · inbound

Zero-Shot DINOv3-Based Image Matching via Many-to-Many Association cites this paper.

Zero-Shot DINOv3-Based Image Matching via Many-to-Many Association Deep ViT Features as Dense Visual Descriptors

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:11:18.694655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-08T06:32:18.423143Z digest=sha256:37ec8265384eccf593ccdca4c350898d26f82d0759f05dfa7f01e2bc25e161fd

Observation 6f2a9136-6242-4406-bfe5-c6a25c9ca3aa · inbound

Zero-Shot DINOv3-Based Image Matching via Many-to-Many Association cites this paper.

Zero-Shot DINOv3-Based Image Matching via Many-to-Many Association Deep ViT Features as Dense Visual Descriptors

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T15:38:37.458871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T15:38:37.458871Z digest=sha256:a6cc75f40324ccebfa97f4c36302633ad847be2c9509f73117ea48c22369cb3d

Observation cacdf124-b7f6-4fba-9c21-2eb8edd60cc4 · inbound

VISION-SLS: Safe Perception-Based Control from Learned Visual Representations via System Level Synthesis cites this paper.

VISION-SLS: Safe Perception-Based Control from Learned Visual Representations via System Level Synthesis Deep ViT Features as Dense Visual Descriptors

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:26:15.116378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-08T02:44:01.250176Z digest=sha256:5381985917df2725b3acacdb2e8bc88a3ca18c9c375256aa62e7209154e45dc9

Observation ad55ea51-ba0c-4c6c-bd87-ec7b8156ce57 · inbound

Beyond the Last Layer: Multi-Layer Representation Fusion for Visual Tokenization cites this paper.

Beyond the Last Layer: Multi-Layer Representation Fusion for Visual Tokenization Deep ViT Features as Dense Visual Descriptors

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:31:23.422461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-12T05:31:06.963637Z digest=sha256:2d59167192b6e1354a9a5c5504d9556c561643ffbb763eb9b1815823574dee48

Observation b7de8118-5a6b-40b7-85bb-92db0592e9cb · inbound

Beyond the Last Layer: Multi-Layer Representation Fusion for Visual Tokenization cites this paper.

Beyond the Last Layer: Multi-Layer Representation Fusion for Visual Tokenization Deep ViT Features as Dense Visual Descriptors

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:42:30.587769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-13T07:40:16.927031Z digest=sha256:4dcc4be04fea678b800bde75ec6898a55bd807679efcbfadbdca5dbd0e596e1a

Observation d48c6675-3083-4b2d-b1f8-a3ccefc7e33e · inbound

Realiz3D: 3D Generation Made Photorealistic via Domain-Aware Learning cites this paper.

Realiz3D: 3D Generation Made Photorealistic via Domain-Aware Learning Deep ViT Features as Dense Visual Descriptors

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-15T07:39:50.413208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T07:36:16.124800Z digest=sha256:de0cce84e8f20a184b16798d773eb1717c298ce3659bac69e4c2ba7321849e92

Observation d17dd71a-628e-45d2-b230-6de38d4fbf5c · inbound

Registers Matter for Pixel-Space Diffusion Transformers cites this paper.

Registers Matter for Pixel-Space Diffusion Transformers Deep ViT Features as Dense Visual Descriptors

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:08:54.067427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T19:08:23.052621Z digest=sha256:1dbf320aaae2490e492472efa8e7d88aa2c79009c2bb1a92a1a48cfd923f019e

Observation bf6f00b2-ba86-4c10-9ad3-c6f2371e2a5a · inbound

PaintCopilot: Modeling Painting as Autonomous Artistic Continuation cites this paper.

PaintCopilot: Modeling Painting as Autonomous Artistic Continuation Deep ViT Features as Dense Visual Descriptors

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-21T05:29:39.311103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T05:29:27.751229Z digest=sha256:beb05ab280f94886f694da3dbb64f687ffd41790a33a659bd0ec1c9cc6557c47

Observation 096841c4-3df3-49ab-8bc2-a419b830d5d5 · inbound

Unsupervised Semantic Segmentation Facilitates Model Understanding cites this paper.

Unsupervised Semantic Segmentation Facilitates Model Understanding Deep ViT Features as Dense Visual Descriptors

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T08:43:15.483441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T08:37:02.350175Z digest=sha256:b826a5b63984cd0a328b09d64819e12fec92be5327672c70df61f815ee069e3a

Observation add15937-84a7-4852-9cc9-527f88040d52 · inbound

Unsupervised Semantic Segmentation Facilitates Model Understanding cites this paper.

Unsupervised Semantic Segmentation Facilitates Model Understanding Deep ViT Features as Dense Visual Descriptors

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T00:39:16.375764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-04T00:34:21.224797Z digest=sha256:d95ffd2191485d7f6e8272865aef056a875f032c9031fba6c6a447311de6ff36

Observation c113a4ef-3424-4e23-a27d-e7dcba2ae667 · inbound

Generate in Reconstruction Space, Match in Semantic Space: Transport Geometry for One-Step Generation cites this paper.

Generate in Reconstruction Space, Match in Semantic Space: Transport Geometry for One-Step Generation Deep ViT Features as Dense Visual Descriptors

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:02:34.265670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-28T18:58:13.956526Z digest=sha256:e9bf3819dd3f5d4cb887ba231db795ea81290e56df37f33fc449f2be9bb12abf

Observation 118fa68b-0cd4-4127-9d9f-24d150490879 · inbound

TUDSR: Twice Upsampling-Diffusion for Higher Super-Resolution cites this paper.

TUDSR: Twice Upsampling-Diffusion for Higher Super-Resolution Deep ViT Features as Dense Visual Descriptors

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:07:29.960791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-27T16:50:03.921332Z digest=sha256:f9b99cd45088df4ef78f27399a714d9c33a233fa528f14dec8d69614edcaa78e

Observation 4ceaa6ae-6e87-4ac8-9c6c-1f5d751731d2 · inbound

Lightweight 3D Feature Pretraining by Bayesian Inversion of 2D Foundation Models cites this paper.

Lightweight 3D Feature Pretraining by Bayesian Inversion of 2D Foundation Models Deep ViT Features as Dense Visual Descriptors

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-04T06:09:37.349741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-26T14:50:10.041000Z digest=sha256:45534ae59f6a7c3a959408d51c7a346bd5579faac253210be70a9d6e2c30c7fa

Observation e1b14f28-cd12-4463-90f3-9075f0277409 · inbound

FROST: Training-Free Few-Shot Segmentation with Frozen Features and Nonparametric Statistics cites this paper.

FROST: Training-Free Few-Shot Segmentation with Frozen Features and Nonparametric Statistics Deep ViT Features as Dense Visual Descriptors

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:55:35.969649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-01T06:48:55.341060Z digest=sha256:340ccc10daa503e683e31f31cefabe5f7a5e84c7f8afc0b7a97ac5dcd50c520b

Observation 9165e141-8c7a-492e-a41d-47bc12d63067 · inbound

Understanding Geometric Representations in Self-Supervised Vision Transformers via Subspace Intervention cites this paper.

Understanding Geometric Representations in Self-Supervised Vision Transformers via Subspace Intervention Deep ViT Features as Dense Visual Descriptors

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T15:38:33.115406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-03T15:34:54.954593Z digest=sha256:93c92bd5312619e9f656623cd5c22a0aeca6eda50d1df8471a3eca7f86af8dcd

Observation 723a8860-86c8-4b43-9c75-d173871c3f1c · inbound

Awakening Diffusion Transformers: Eliciting Stronger Generation and Understanding via Massive Activation Modulation cites this paper.

Awakening Diffusion Transformers: Eliciting Stronger Generation and Understanding via Massive Activation Modulation Deep ViT Features as Dense Visual Descriptors

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-12T05:46:45.293902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T05:46:45.293902Z digest=sha256:70300965e556e15d453580715c1678934c0081e2129e3a068ce8ba3204579dc5

Observation 90ea3029-786c-4f02-8b36-6e451092037e · inbound

`Attention-Guided Cross-Temporal Clustering for Self-Supervised Video Object Segmentation cites this paper.

`Attention-Guided Cross-Temporal Clustering for Self-Supervised Video Object Segmentation Deep ViT Features as Dense Visual Descriptors

Reference 45

Resolution
malformed identifier
local_arxiv, observed 2026-07-09T16:56:21.645918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T16:55:56.118110Z digest=sha256:1cccfe9eeada98c0e970f60c40a995d0ae26d220698722cbc06ab7b4751e53e9

Observation 97913c0c-7a46-4454-8f37-adad8f4aa698 · inbound

SeeSE3: Emergence of 3D Space in Vision Features cites this paper.

SeeSE3: Emergence of 3D Space in Vision Features Deep ViT Features as Dense Visual Descriptors

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T02:49:55.065331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:49:55.065331Z digest=sha256:cf65d9637bf68223b85c5bdeda5e21c2ed870be5dad230277bd101ed7585888e

Observation 08ffb6ba-dbcb-4766-86a0-56877d494a74 · inbound

Feature-Guided Diffusion for Non-Differentiable Inverse Rendering cites this paper.

Feature-Guided Diffusion for Non-Differentiable Inverse Rendering Deep ViT Features as Dense Visual Descriptors

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T18:06:36.322250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:06:36.322250Z digest=sha256:9bd00e81db9419c8717afef1d65765db2139530b1a62d3af6d4f52ced6c93a6b