Pith. sign in

Paper Citation Record · LEDGER

Deep ViT Features as Dense Visual Descriptors

As of 5 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 25 inbound Pith citation observations for arXiv:2112.05814.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2112.05814 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 25 of 25 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T10:18:31.361765Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-09T16:56:21.644793Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation eed0412e-546b-46d5-9640-b6a614a8de4a · inbound

ReKep: Spatio-Temporal Reasoning of Relational Keypoint Constraints for Robotic Manipulation cites this paper.

ReKep: Spatio-Temporal Reasoning of Relational Keypoint Constraints for Robotic Manipulation Deep ViT Features as Dense Visual Descriptors

Reference 119

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:25:18.019787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:25:17.847571Z digest=sha256:ec945205cede662c74739cde428d6078a5bfbc08c504f76455210215c2a5585a

Observation 41863813-330f-4116-8225-29db3a6c550b · inbound

From One-to-One to Many-to-Many: Dynamic Cross-Layer Injection for Deep Vision-Language Fusion cites this paper.

From One-to-One to Many-to-Many: Dynamic Cross-Layer Injection for Deep Vision-Language Fusion Deep ViT Features as Dense Visual Descriptors

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T10:18:31.361765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:18:31.361765Z digest=sha256:6b063b954ed364ff3b0bfc4b4075f3713895b9ac26f2647875bfbdf640984de7

Observation d9fc5d41-16d8-4161-9084-f14b8de71451 · inbound

UPLiFT: Efficient Pixel-Dense Feature Upsampling with Local Attenders cites this paper.

UPLiFT: Efficient Pixel-Dense Feature Upsampling with Local Attenders Deep ViT Features as Dense Visual Descriptors

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T08:13:39.247257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:13:39.247257Z digest=sha256:cd90acb6d1ec8071a5d343da3aae586a422c8f9902e96188633033c79c994058

Observation 504b71b5-f443-4e31-8795-55a7245c90a3 · inbound

Uncertainty-Aware Hierarchical Re-Localization in OpenStreetMap via Semantic Alignment cites this paper.

Uncertainty-Aware Hierarchical Re-Localization in OpenStreetMap via Semantic Alignment Deep ViT Features as Dense Visual Descriptors

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-02T19:36:19.815385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:36:19.815385Z digest=sha256:dd57575fff780158fa3cf6f1b29cecc08f5d15deb41218390b5548430838e42e

Observation e772a24d-ec4d-4ec9-ad39-298326d703c3 · inbound

Autonomous Search for Sparsely Distributed Visual Phenomena through Environmental Context Modeling cites this paper.

Autonomous Search for Sparsely Distributed Visual Phenomena through Environmental Context Modeling Deep ViT Features as Dense Visual Descriptors

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-14T23:50:28.153875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:50:28.153875Z digest=sha256:39e8089d67677a0658e69837eb4a04a9f8a451a92a44c9fb7382f49a2c863c92

Observation 55c0e98a-11cd-4bf4-a8cb-76e93dbdadad · inbound

Human-like Object Grouping in Self-supervised Vision Transformers cites this paper.

Human-like Object Grouping in Self-supervised Vision Transformers Deep ViT Features as Dense Visual Descriptors

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-14T21:34:48.709465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T21:34:48.709465Z digest=sha256:ba9f2d5aaaec1a4d2ada5961633c97b04d901e2be91b782f634ef1cf6db3c72e

Observation 47c61a5c-b9c4-4ad8-aa62-0627a4ec43a3 · inbound

Zero-Shot DINOv3-Based Image Matching via Many-to-Many Association cites this paper.

Zero-Shot DINOv3-Based Image Matching via Many-to-Many Association Deep ViT Features as Dense Visual Descriptors

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:11:18.694655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T06:32:18.423143Z digest=sha256:99d2fdc3856e86da10034709a55788fb4afc85a36dcb4aa4d67519b83d21928b

Observation 6f2a9136-6242-4406-bfe5-c6a25c9ca3aa · inbound

Zero-Shot DINOv3-Based Image Matching via Many-to-Many Association cites this paper.

Zero-Shot DINOv3-Based Image Matching via Many-to-Many Association Deep ViT Features as Dense Visual Descriptors

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T15:38:37.458871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T15:38:37.458871Z digest=sha256:0d69abf4fc22f0ad5a1e1516afdec51094e4b3d79f82e58b52146142687c6213

Observation cacdf124-b7f6-4fba-9c21-2eb8edd60cc4 · inbound

VISION-SLS: Safe Perception-Based Control from Learned Visual Representations via System Level Synthesis cites this paper.

VISION-SLS: Safe Perception-Based Control from Learned Visual Representations via System Level Synthesis Deep ViT Features as Dense Visual Descriptors

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:26:15.116378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T02:44:01.250176Z digest=sha256:cb4f39227405c4109e9afa4d687a3c648f57dd649de93a376f7afdb7d39f753d

Observation ad55ea51-ba0c-4c6c-bd87-ec7b8156ce57 · inbound

Beyond the Last Layer: Multi-Layer Representation Fusion for Visual Tokenization cites this paper.

Beyond the Last Layer: Multi-Layer Representation Fusion for Visual Tokenization Deep ViT Features as Dense Visual Descriptors

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:31:23.422461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T05:31:06.963637Z digest=sha256:f870269a7aa41d836762a649c4d86686bec1ebe01a6342d3552782d5984add10

Observation b7de8118-5a6b-40b7-85bb-92db0592e9cb · inbound

Beyond the Last Layer: Multi-Layer Representation Fusion for Visual Tokenization cites this paper.

Beyond the Last Layer: Multi-Layer Representation Fusion for Visual Tokenization Deep ViT Features as Dense Visual Descriptors

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:42:30.587769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T07:40:16.927031Z digest=sha256:50027e0f1e9fdbdea935f32e26e2ea4366f9fe5b2f69d3c9237f646bc1d7b708

Observation d48c6675-3083-4b2d-b1f8-a3ccefc7e33e · inbound

Realiz3D: 3D Generation Made Photorealistic via Domain-Aware Learning cites this paper.

Realiz3D: 3D Generation Made Photorealistic via Domain-Aware Learning Deep ViT Features as Dense Visual Descriptors

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-15T07:39:50.413208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T07:36:16.124800Z digest=sha256:0cbf81288adbecd056f2ff11696c218ae7d21dd5ec8a23421cd1f174579c1d2b

Observation d17dd71a-628e-45d2-b230-6de38d4fbf5c · inbound

Registers Matter for Pixel-Space Diffusion Transformers cites this paper.

Registers Matter for Pixel-Space Diffusion Transformers Deep ViT Features as Dense Visual Descriptors

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:08:54.067427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T19:08:23.052621Z digest=sha256:274ab4506ebed3d545284f9f41c8fb12b9df4d9ee54f184a13c90e3a40325486

Observation bf6f00b2-ba86-4c10-9ad3-c6f2371e2a5a · inbound

PaintCopilot: Modeling Painting as Autonomous Artistic Continuation cites this paper.

PaintCopilot: Modeling Painting as Autonomous Artistic Continuation Deep ViT Features as Dense Visual Descriptors

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-21T05:29:39.311103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T05:29:27.751229Z digest=sha256:367f24261b1e0cbe189958d28d8ad140fef602521120e79a811f68c7d16758f2

Observation 096841c4-3df3-49ab-8bc2-a419b830d5d5 · inbound

Unsupervised Semantic Segmentation Facilitates Model Understanding cites this paper.

Unsupervised Semantic Segmentation Facilitates Model Understanding Deep ViT Features as Dense Visual Descriptors

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T08:43:15.483441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T08:37:02.350175Z digest=sha256:7e8760f84637be08e356b2a138ec9b89b5e26b139f42a28e09685bd059d0bfb7

Observation add15937-84a7-4852-9cc9-527f88040d52 · inbound

Unsupervised Semantic Segmentation Facilitates Model Understanding cites this paper.

Unsupervised Semantic Segmentation Facilitates Model Understanding Deep ViT Features as Dense Visual Descriptors

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T00:39:16.375764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-04T00:34:21.224797Z digest=sha256:c5d81c9c35d8495b540044089944fae8b2a7ab33e87a9b938fc0b85cbda4cf71

Observation c113a4ef-3424-4e23-a27d-e7dcba2ae667 · inbound

Generate in Reconstruction Space, Match in Semantic Space: Transport Geometry for One-Step Generation cites this paper.

Generate in Reconstruction Space, Match in Semantic Space: Transport Geometry for One-Step Generation Deep ViT Features as Dense Visual Descriptors

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:02:34.265670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T18:58:13.956526Z digest=sha256:0df0a70861ff7b906413ad76b4fa4a6b5de84b955dc9fd89b72c1b3321f39a50

Observation 118fa68b-0cd4-4127-9d9f-24d150490879 · inbound

TUDSR: Twice Upsampling-Diffusion for Higher Super-Resolution cites this paper.

TUDSR: Twice Upsampling-Diffusion for Higher Super-Resolution Deep ViT Features as Dense Visual Descriptors

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:07:29.960791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-27T16:50:03.921332Z digest=sha256:de478b966e1de06f3dc423dd79f2bc939bf192af919dd8d1c920c67c1e5f1ed5

Observation 4ceaa6ae-6e87-4ac8-9c6c-1f5d751731d2 · inbound

Lightweight 3D Feature Pretraining by Bayesian Inversion of 2D Foundation Models cites this paper.

Lightweight 3D Feature Pretraining by Bayesian Inversion of 2D Foundation Models Deep ViT Features as Dense Visual Descriptors

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-04T06:09:37.349741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-26T14:50:10.041000Z digest=sha256:0e3b2e9afe9b6737c85b219553e9e155cda9aea1f441c2693d3b608cd9f0a40a

Observation e1b14f28-cd12-4463-90f3-9075f0277409 · inbound

FROST: Training-Free Few-Shot Segmentation with Frozen Features and Nonparametric Statistics cites this paper.

FROST: Training-Free Few-Shot Segmentation with Frozen Features and Nonparametric Statistics Deep ViT Features as Dense Visual Descriptors

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:55:35.969649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-01T06:48:55.341060Z digest=sha256:18335fc174254bb0b82b4fd7e09d2d0372db3ba44d691872341ac130cfb94d1d

Observation 9165e141-8c7a-492e-a41d-47bc12d63067 · inbound

Understanding Geometric Representations in Self-Supervised Vision Transformers via Subspace Intervention cites this paper.

Understanding Geometric Representations in Self-Supervised Vision Transformers via Subspace Intervention Deep ViT Features as Dense Visual Descriptors

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T15:38:33.115406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-03T15:34:54.954593Z digest=sha256:9e8178b877111e98ed85de0671a15a3ffea2d333da189427928b219bff92b7bf

Observation 723a8860-86c8-4b43-9c75-d173871c3f1c · inbound

Awakening Diffusion Transformers: Eliciting Stronger Generation and Understanding via Massive Activation Modulation cites this paper.

Awakening Diffusion Transformers: Eliciting Stronger Generation and Understanding via Massive Activation Modulation Deep ViT Features as Dense Visual Descriptors

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-12T05:46:45.293902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T05:46:45.293902Z digest=sha256:3cf74961fabd52626eab64c7f00986448099fdcfe8237f5f82fe701e30da40e4

Observation 90ea3029-786c-4f02-8b36-6e451092037e · inbound

`Attention-Guided Cross-Temporal Clustering for Self-Supervised Video Object Segmentation cites this paper.

`Attention-Guided Cross-Temporal Clustering for Self-Supervised Video Object Segmentation Deep ViT Features as Dense Visual Descriptors

Reference 45

Resolution
malformed identifier
local_arxiv, observed 2026-07-09T16:56:21.645918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-09T16:55:56.118110Z digest=sha256:efcdb3602796344d2788b6cf384cddee64228c1d8994adc6ad0dff1d3bdce66d

Observation 97913c0c-7a46-4454-8f37-adad8f4aa698 · inbound

SeeSE3: Emergence of 3D Space in Vision Features cites this paper.

SeeSE3: Emergence of 3D Space in Vision Features Deep ViT Features as Dense Visual Descriptors

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T02:49:55.065331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:49:55.065331Z digest=sha256:c6ccc4dec651672cf0f9e55f3b5afd48366233201da07166744d8a851218103c

Observation 08ffb6ba-dbcb-4766-86a0-56877d494a74 · inbound

Feature-Guided Diffusion for Non-Differentiable Inverse Rendering cites this paper.

Feature-Guided Diffusion for Non-Differentiable Inverse Rendering Deep ViT Features as Dense Visual Descriptors

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T18:06:36.322250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:06:36.322250Z digest=sha256:3bb07b3b78772395e5318e3125133178e9fe03821ac142df2d81ade953b19972