Pith. sign in

Paper Citation Record · LEDGER

CLIP-DINOiser: Teaching CLIP a few DINO tricks for open-vocabulary semantic segmentation

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2312.12359.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2312.12359 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:34:51.113165Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-15T19:10:15.655884Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 906a4119-549c-4f0c-bd06-546106e4190b · inbound

Harnessing Vision Foundation Models for High-Performance, Training-Free Open Vocabulary Segmentation cites this paper.

Harnessing Vision Foundation Models for High-Performance, Training-Free Open Vocabulary Segmentation CLIP-DINOiser: Teaching CLIP a few DINO tricks for open-vocabulary semantic segmentation

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-12T21:02:01.039742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:02:01.039742Z digest=sha256:4d682e941b3f5ab61576a9ff18efd2b1b369d74ad3f2e353721b1c51f48d182c

Observation ead9691a-9e08-4023-85c4-3cbcf69610f7 · inbound

Distilling Spectral Graph for Object-Context Aware Open-Vocabulary Semantic Segmentation cites this paper.

Distilling Spectral Graph for Object-Context Aware Open-Vocabulary Semantic Segmentation CLIP-DINOiser: Teaching CLIP a few DINO tricks for open-vocabulary semantic segmentation

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T12:33:52.302895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:33:52.302895Z digest=sha256:711bd28e7bbf23901b7dd9a14468596fdb1cfdd31ca38a00864c990c7afb01fc

Observation 69530690-b4b4-4d70-b092-d26aa33fe939 · inbound

Causal Graphical Models for Vision-Language Compositional Understanding cites this paper.

Causal Graphical Models for Vision-Language Compositional Understanding CLIP-DINOiser: Teaching CLIP a few DINO tricks for open-vocabulary semantic segmentation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T17:10:44.364395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:10:44.364395Z digest=sha256:b0d1abbe2846e859c0eacd3541b76ff6a7ac1526ab01811c30a42a9f61c75c21

Observation bcacc359-537f-40ee-ba76-661dedc8a5a1 · inbound

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation cites this paper.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation CLIP-DINOiser: Teaching CLIP a few DINO tricks for open-vocabulary semantic segmentation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T12:28:54.577649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:28:54.577649Z digest=sha256:e9bff8f982cca51e11ee2cbd86274e1656f1dde65d522fa8bec3366d19809389

Observation 83bd5aa6-70ee-4da8-8514-d9fa7599938e · inbound

DeCLIP: Decoupled Learning for Open-Vocabulary Dense Perception cites this paper.

DeCLIP: Decoupled Learning for Open-Vocabulary Dense Perception CLIP-DINOiser: Teaching CLIP a few DINO tricks for open-vocabulary semantic segmentation

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-15T23:34:51.113165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:34:51.113165Z digest=sha256:5672f959169cfed84818de5d0494f23fe9c27cdefd908a9a2dce576cb6ad6211

Observation 32ddfacc-5600-41c8-9c72-85f9cbef18d3 · inbound

Plug-in Feedback Self-adaptive Attention in CLIP for Training-free Open-Vocabulary Segmentation cites this paper.

Plug-in Feedback Self-adaptive Attention in CLIP for Training-free Open-Vocabulary Segmentation CLIP-DINOiser: Teaching CLIP a few DINO tricks for open-vocabulary semantic segmentation

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-05T15:18:46.555130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:18:46.555130Z digest=sha256:1d62561cdfe28c03623e196c96fb4b568398b218b139c04c83364b5fe3720b1e

Observation b65c4bf7-0c81-4bc4-ac79-79689399890a · inbound

Vision Transformers Need More Than Registers cites this paper.

Vision Transformers Need More Than Registers CLIP-DINOiser: Teaching CLIP a few DINO tricks for open-vocabulary semantic segmentation

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-15T19:10:15.659677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T19:08:50.579190Z digest=sha256:9b38d63675ca7755a30aade93b87e8f06f21f300aea55594abed8a764a8b496b

Observation 0d1b3e93-86f6-4244-944b-671e9c1c640d · inbound

FUS3DMaps: Scalable and Accurate Open-Vocabulary Semantic Mapping by 3D Fusion of Voxel- and Instance-Level Layers cites this paper.

FUS3DMaps: Scalable and Accurate Open-Vocabulary Semantic Mapping by 3D Fusion of Voxel- and Instance-Level Layers CLIP-DINOiser: Teaching CLIP a few DINO tricks for open-vocabulary semantic segmentation

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:56:30.052572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-07T15:47:11.780364Z digest=sha256:dab8ebbd44781c09e786d840a21222cb0955fdb05370b45ead0ceab26a6e919c