Pith. sign in

Paper Citation Record · LEDGER

Beyond-Labels: Advancing Open-Vocabulary Segmentation With Vision-Language Models

As of 12 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 0 inbound Pith citation observations for arXiv:2501.16769.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.16769 v5

Coverage vector

measured 32 of 32 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T10:54:06.184306Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

32 of 32 outbound references displayed

  • verified exact0
  • verified fuzzy25
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6f91db66-a518-4fdf-860f-52a5e47d5eb1 · outbound

This paper cites Flamingo: a Visual Language Model for Few-Shot Learning.

Beyond-Labels: Advancing Open-Vocabulary Segmentation With Vision-Language Models Flamingo: a Visual Language Model for Few-Shot Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T10:54:06.048139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:54:06.048139Z digest=sha256:7ea21a6c782a08c476561d05ec5e33709bd99baa882c2cdb68d265fcbd61f403

Observation 75a1b2de-421c-4634-b4f1-63f5649ac2c3 · outbound

This paper cites Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, and Amanda Askell et al.

Beyond-Labels: Advancing Open-Vocabulary Segmentation With Vision-Language Models Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, and Amanda Askell et al

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:54:06.690980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T10:54:06.053021Z digest=sha256:17d59949a0bd01cb39530df9c44e38a8417f9587e6c4ef9aebfa7298a29ebfde

Observation 14b24a7e-c4b8-45a2-8fe7-8882bff9b5dd · outbound

This paper cites Zero-shot semantic segmentation.

Beyond-Labels: Advancing Open-Vocabulary Segmentation With Vision-Language Models Zero-shot semantic segmentation

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:54:06.674025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T10:54:06.057178Z digest=sha256:fa456864731a64e33c4ad0f3f128c3a4a4f04e96432bbffac3754dc962ee43d2

Observation 6592b3b3-f1d7-4c29-b710-2bf5831a6497 · outbound

This paper cites Emerging properties in self-supervised vision transformers.

Beyond-Labels: Advancing Open-Vocabulary Segmentation With Vision-Language Models Emerging properties in self-supervised vision transformers

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:54:06.655694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T10:54:06.061264Z digest=sha256:31d7c2fb8ad311d276aa4c9128bed5137b1be0867f1d439c52c22cc5ea3357f8

Observation 494d0fdf-2882-40a7-9ad2-78a5bd9a130f · outbound

This paper cites An empirical study of training self-supervised vision transformers.

Beyond-Labels: Advancing Open-Vocabulary Segmentation With Vision-Language Models An empirical study of training self-supervised vision transformers

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:54:06.640062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T10:54:06.065554Z digest=sha256:8d9cbe93c59949107fe692df8eff795b2b5f6c51466e580462409c1fc249066f

Observation e0d3caa3-c80b-4217-bb43-ab3005c413db · outbound

This paper cites Bert: Pre-training of deep bidirectional transformers for language un- derstanding.

Beyond-Labels: Advancing Open-Vocabulary Segmentation With Vision-Language Models Bert: Pre-training of deep bidirectional transformers for language un- derstanding

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:54:06.624670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T10:54:06.069805Z digest=sha256:267674f5b759c5f60627f6a83354966026e9a4567ac8f2cb8af5f79cb92a0b46

Observation e064962a-193e-4628-a3a3-5882205b3903 · outbound

This paper cites PaLM: Scaling Language Modeling with Pathways.

Beyond-Labels: Advancing Open-Vocabulary Segmentation With Vision-Language Models PaLM: Scaling Language Modeling with Pathways

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T10:54:06.074250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:54:06.074250Z digest=sha256:2a9ad86f75e8e65ef91c220f4af868571525e5a2b081fd71351469a1e8564bfa

Observation aeab8e8e-f33f-4dba-a973-9bb8a0982768 · outbound

This paper cites On the Opportunities and Risks of Foundation Models.

Beyond-Labels: Advancing Open-Vocabulary Segmentation With Vision-Language Models On the Opportunities and Risks of Foundation Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T10:54:06.078400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:54:06.078400Z digest=sha256:f1a5d987e18c86de318cd5aeed6cdbeba1aca405c4c9382f72a6eed55d372435

Observation aea3d2e5-b556-4e1b-893e-29fce153770f · outbound

This paper cites The pascal visual object classes (voc) challenge.

Beyond-Labels: Advancing Open-Vocabulary Segmentation With Vision-Language Models The pascal visual object classes (voc) challenge

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:54:06.609581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T10:54:06.082565Z digest=sha256:92280c249963b56d26cd264499900b92619917a5855d5bd7160698fab1f11200

Observation 1ff063a3-5c47-47a7-87c1-5bb2d5cce502 · outbound

This paper cites Bootstrap your own latent-a new approach to self-supervised learning.

Beyond-Labels: Advancing Open-Vocabulary Segmentation With Vision-Language Models Bootstrap your own latent-a new approach to self-supervised learning

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:54:06.595119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T10:54:06.086268Z digest=sha256:d15bb4f2d35ef34a65fb6d4b76545566f2d28be4632371d62b4bda86922b401e

Observation 5c880cc9-ef6c-45ab-9088-a5222b9a5247 · outbound

This paper cites Context-aware feature generation for zero-shot semantic segmentation.

Beyond-Labels: Advancing Open-Vocabulary Segmentation With Vision-Language Models Context-aware feature generation for zero-shot semantic segmentation

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:54:06.579730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T10:54:06.090301Z digest=sha256:4d1240860696c60c109161fe4cd9c1e6305ca645e2ff185c067db07c9e4e56a8

Observation 4ee4c634-c2d1-45ab-9c05-fd954edc82c1 · outbound

This paper cites Simultaneous detection and segmentation.

Beyond-Labels: Advancing Open-Vocabulary Segmentation With Vision-Language Models Simultaneous detection and segmentation

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:54:06.563893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T10:54:06.094096Z digest=sha256:e3e26cdc3e2a54fa80ab817fb31b16a2fcec05ac75b8adaf5dfb399788cb769d

Observation a61a390e-fbd0-4240-ab9c-1f6b8eecca94 · outbound

This paper cites Perceiver IO: A General Architecture for Structured Inputs & Outputs.

Beyond-Labels: Advancing Open-Vocabulary Segmentation With Vision-Language Models Perceiver IO: A General Architecture for Structured Inputs & Outputs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T10:54:06.098569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:54:06.098569Z digest=sha256:4bf55b2246fec848621f7c47c85be939429399a1e2b40f4b49a15b8707687f16

Observation 3bda8739-f5dc-4d22-bccd-2de83f9df1ad · outbound

This paper cites Scaling up visual and vision-language representation learning with noisy text supervision.

Beyond-Labels: Advancing Open-Vocabulary Segmentation With Vision-Language Models Scaling up visual and vision-language representation learning with noisy text supervision

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:54:06.549229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T10:54:06.103351Z digest=sha256:088284b2f5da8b885b6c6ad0248ef358483f88c5ede19c5558baefd74d70f4e2

Observation 52a6fcf3-2f09-4983-acdc-ff0452739398 · outbound

This paper cites Deep spectral methods: A surprisingly strong baseline for unsupervised semantic segmentation and localization.

Beyond-Labels: Advancing Open-Vocabulary Segmentation With Vision-Language Models Deep spectral methods: A surprisingly strong baseline for unsupervised semantic segmentation and localization

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:54:06.536186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T10:54:06.107703Z digest=sha256:0e167a822f449bf8d07e2d0e6c1b57c05317a3b6e7ba5a832fd6e5583d59a82e

Observation aeca5109-b1ce-4b1d-be25-f52da5c75bf6 · outbound

This paper cites NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis.

Beyond-Labels: Advancing Open-Vocabulary Segmentation With Vision-Language Models NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T10:54:06.112024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:54:06.112024Z digest=sha256:f0418a704284be2815b88a85cb44759219cc1751ccbfc445d25ded6726347337

Observation 7bf68296-1297-4492-9a37-1627f9515d2e · outbound

This paper cites Learning transferable visual models from natural language supervision.

Beyond-Labels: Advancing Open-Vocabulary Segmentation With Vision-Language Models Learning transferable visual models from natural language supervision

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:54:06.522822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T10:54:06.116677Z digest=sha256:d912e637c7917fd18524665581b952ab88ab7d3c94c05150e676403d969e3efd

Observation 62cbd38c-5f81-4833-ada5-b73883eda648 · outbound

This paper cites Conditional networks for few-shot semantic segmenta- tion.

Beyond-Labels: Advancing Open-Vocabulary Segmentation With Vision-Language Models Conditional networks for few-shot semantic segmenta- tion

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:54:06.508628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T10:54:06.120860Z digest=sha256:9fc48fafe7a0aead875fba40441b9133903e3d7143942f5340a0c539411fcfca

Observation 89ddb77d-8a7a-46de-a332-b0c8c9969170 · outbound

This paper cites Vision trans- formers for dense prediction.

Beyond-Labels: Advancing Open-Vocabulary Segmentation With Vision-Language Models Vision trans- formers for dense prediction

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:54:06.494736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T10:54:06.125173Z digest=sha256:c332d36a9cd9aa220f2cd0c84801d8dca375b3203457857be9d4a6da03655466

Observation 9f59eb25-6a67-435b-8c08-e17b1739f039 · outbound

This paper cites U-net: Convolu- tional networks for biomedical image segmentation.

Beyond-Labels: Advancing Open-Vocabulary Segmentation With Vision-Language Models U-net: Convolu- tional networks for biomedical image segmentation

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:54:06.480252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T10:54:06.129392Z digest=sha256:83627be906b3c9f20da16348718e8c119616ed685caf6c54ac99e2dd30b9d665

Observation 215dbaaf-f008-412d-9e1f-e8cc61a7ec10 · outbound

This paper cites One-shot learning for semantic segmentation.

Beyond-Labels: Advancing Open-Vocabulary Segmentation With Vision-Language Models One-shot learning for semantic segmentation

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:54:06.464797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T10:54:06.133810Z digest=sha256:c337c17332707e223511343a643a00f4f36e87897e9a6c08802e2bdde50eaa25

Observation f59dcb86-ce78-4e8f-9a82-93f5a7a8aa4b · outbound

This paper cites Fully convolutional networks for semantic segmentation.

Beyond-Labels: Advancing Open-Vocabulary Segmentation With Vision-Language Models Fully convolutional networks for semantic segmentation

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:54:06.449246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T10:54:06.138184Z digest=sha256:622cfa29c3dec80e50ac4d0cf0f1a8d697373a495fe623f0e0d543bb14b5ef42

Observation 1b89e5fb-d2f1-4467-8b86-95314d4991b9 · outbound

This paper cites Unsupervised salient object detection with spectral cluster voting.

Beyond-Labels: Advancing Open-Vocabulary Segmentation With Vision-Language Models Unsupervised salient object detection with spectral cluster voting

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:54:06.433134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T10:54:06.142625Z digest=sha256:e0ac007c370b866e01f3f178efddb57e7d95d22d020e48aeda9d4c3a2138754b

Observation 713261a4-52b4-4cbb-bfc8-26da241b6462 · outbound

This paper cites Oreshkin, and Martin Jagersand.

Beyond-Labels: Advancing Open-Vocabulary Segmentation With Vision-Language Models Oreshkin, and Martin Jagersand

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:54:06.418859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T10:54:06.147009Z digest=sha256:c900b872d35b20de694888d5ce8fcb4a4018e476384b853e0b433823e4126f95

Observation 30a4347c-7849-4d82-a040-d8d4fcc0dfab · outbound

This paper cites Implicit neural representations with periodic activation functions.

Beyond-Labels: Advancing Open-Vocabulary Segmentation With Vision-Language Models Implicit neural representations with periodic activation functions

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:54:06.404206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T10:54:06.153490Z digest=sha256:4e2210d10f8b964d3145fb3d9580522fa106e159d8a014665257ccb23afb7a9f

Observation 92d4fcef-9cba-492e-a833-59694bc3adc8 · outbound

This paper cites Fourier features let networks learn high frequency functions in low dimensional domains.

Beyond-Labels: Advancing Open-Vocabulary Segmentation With Vision-Language Models Fourier features let networks learn high frequency functions in low dimensional domains

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:54:06.387498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T10:54:06.157801Z digest=sha256:add76c2ec8455b4f555c04fe0378d97e995eb6e4ae4732d4cd54813b2f1910ae

Observation dd9e2a18-daa9-4bda-af36-143e4919a601 · outbound

This paper cites Menick, Serkan Cabi, S.M.

Beyond-Labels: Advancing Open-Vocabulary Segmentation With Vision-Language Models Menick, Serkan Cabi, S.M

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:54:06.370664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T10:54:06.162156Z digest=sha256:eb7300b1e2be4b9a6df6838a7316b68be7f6c47d1e605ee2cb466d12b801f9e3

Observation 2dff4929-f9d5-43f0-87b2-dfefc5905b92 · outbound

This paper cites Attention is all you need.

Beyond-Labels: Advancing Open-Vocabulary Segmentation With Vision-Language Models Attention is all you need

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:54:06.355776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T10:54:06.166384Z digest=sha256:5a460c11b649c978e1fa7f55eedd11555e364e6df4d6cf1181925fc3d728586e

Observation 9e57a93a-c9bb-4a9f-baff-d7e282489966 · outbound

This paper cites Panet: Few-shot image semantic segmentation with prototype alignment.

Beyond-Labels: Advancing Open-Vocabulary Segmentation With Vision-Language Models Panet: Few-shot image semantic segmentation with prototype alignment

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:54:06.341091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T10:54:06.171032Z digest=sha256:08c7314f49b20e1d4374941d1b6052bef7daaac9b720a7d9cb66cad70bfa5cd0

Observation 47b08435-b0c7-4d22-9af2-e8c8c5767cb5 · outbound

This paper cites Semantic projection network for zero-and few-label semantic segmentation.

Beyond-Labels: Advancing Open-Vocabulary Segmentation With Vision-Language Models Semantic projection network for zero-and few-label semantic segmentation

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:54:06.326957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T10:54:06.175446Z digest=sha256:196da734a171546a527bb7ced530a60abb254c801ef327d0de5255d31d3995a2

Observation 2c1aaa83-0174-4e1f-885d-b82037429461 · outbound

This paper cites Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language.

Beyond-Labels: Advancing Open-Vocabulary Segmentation With Vision-Language Models Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T10:54:06.179983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:54:06.179983Z digest=sha256:1e6e7c9e50d9778d327f0f507ddc68a449839dc200be2d7ab5630f4058c49c41

Observation 078f95ff-dde8-4c1c-bb09-0c4d4628d449 · outbound

This paper cites an unresolved cited work.

Beyond-Labels: Advancing Open-Vocabulary Segmentation With Vision-Language Models Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-10T10:54:06.311218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T10:54:06.184306Z digest=sha256:31cbe7a109c86f6230700172b907662da90368f834c1833ba92effa4fff3af4c

Pith citing papers

No inbound Pith citation observations are available.