Pith. sign in

Paper Citation Record · LEDGER

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval

As of 15 August 2026, this Paper Citation Record lists 85 of 85 outbound references and 0 inbound Pith citation observations for arXiv:2412.18806.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.18806 v1

Coverage vector

measured 85 of 85 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T04:32:44.459932Z

measured 85 of 85 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

85 of 85 outbound references displayed

  • verified exact0
  • verified fuzzy61
  • unresolved22
  • parse uncertain1
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ec488dff-669b-41a2-bbff-64aa37573a2a · outbound

This paper cites Label-embedding for image classification.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Label-embedding for image classification

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T04:32:44.172563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:32:44.172563Z digest=sha256:f6633dd52fdf1825438751e3ebfed29e313e8f44f744a37df4d4355d4a4dd5ac

Observation c125ee8e-1fac-4ae5-93ba-7684d01cdf5d · outbound

This paper cites Bottom-up and top-down attention for image captioning and visual question answering.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Bottom-up and top-down attention for image captioning and visual question answering

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T04:32:44.176758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:32:44.176758Z digest=sha256:2bb12addeeb160c3b486a6570921c18d6dcde73b37f99916af85d3c1e0c515b2

Observation 9e82b47a-cec5-48bb-8e04-6d544622e4e8 · outbound

This paper cites Pseudo-labeling and confirmation bias in deep semi-supervised learning.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Pseudo-labeling and confirmation bias in deep semi-supervised learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T04:32:44.184879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:32:44.184879Z digest=sha256:64de5df4682050fd271e5ae10c8546260b0d274a015a7d64c05384608273d6f8

Observation d26f7b55-1ea4-47bd-b0f5-230404ba5725 · outbound

This paper cites Bridg- ing the gap between object and image-level representations for open-vocabulary detection.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Bridg- ing the gap between object and image-level representations for open-vocabulary detection

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T04:32:44.188636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:32:44.188636Z digest=sha256:a498a08204489d74c6e3ac6522aeb5ccd6228c6457a3cc4f0f271089f5636da6

Observation 41970f72-8636-4b34-8d16-633b727c8d03 · outbound

This paper cites Zero-shot object detection.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Zero-shot object detection

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:45.411789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T04:32:44.192362Z digest=sha256:75630b77d2c4be04ca661befc03de3d56d57f4ed5d2769afb1929100ad3e8a1f

Observation 05dc9e84-02a1-4631-a38e-71b6cfa5fae3 · outbound

This paper cites Mixmatch: A holistic approach to semi-supervised learning.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Mixmatch: A holistic approach to semi-supervised learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T04:32:44.196412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:32:44.196412Z digest=sha256:b55b355f190cc26d2689e6841fa8dd944e5379ae4f799ebb7129df63659de871

Observation 17a5de76-97ca-4339-98f7-13b0d9685fbc · outbound

This paper cites Lang, Sourabh V ora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Gi- ancarlo Baldan, and Oscar Beijbom.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Lang, Sourabh V ora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Gi- ancarlo Baldan, and Oscar Beijbom

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:45.394749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T04:32:44.200071Z digest=sha256:dcfad14c010eeaa2a01cc2375675f174669f2b234d5490461871372315f8fcb3

Observation 4e90383b-b9f4-4eec-a88a-bb0a7e0f2dbe · outbound

This paper cites X-detr: A versatile architecture for instance-wise vision- language tasks.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval X-detr: A versatile architecture for instance-wise vision- language tasks

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:45.384161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T04:32:44.203649Z digest=sha256:2d36d2642d3ab2cad0f91d19f485991628dcbafd35d4e91efe99f9d16a253915

Observation f87f0292-183b-4672-acc3-4919ee74a76a · outbound

This paper cites End-to- end object detection with transformers.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval End-to- end object detection with transformers

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:45.373408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T04:32:44.207011Z digest=sha256:898cf39d949c93f182c811279afcf27bb1d707a588db3f01110dacf51fe22b51

Observation 56821bf1-e47b-455d-95fc-595fd94361ff · outbound

This paper cites Big self-supervised mod- els are strong semi-supervised learners.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Big self-supervised mod- els are strong semi-supervised learners

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:45.363471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T04:32:44.210464Z digest=sha256:f9746198cf4e9e53bf19f17bfeff5cac6d3372448753a26f5c411a664c92dd57

Observation 0a833bd5-93f0-4192-95a1-42d607bf4af8 · outbound

This paper cites UNITER: universal image-text representation learning.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval UNITER: universal image-text representation learning

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:45.352804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T04:32:44.214032Z digest=sha256:997a3908fbbc293d2be4d84b11c927938a06660860ee581c34de87be6e188d9e

Observation 61919452-1d70-4fe3-9c12-8542943725a4 · outbound

This paper cites Proba- bilistic embeddings for cross-modal retrieval.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Proba- bilistic embeddings for cross-modal retrieval

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:45.341194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T04:32:44.217445Z digest=sha256:6abac7056d08e5561eba9dcf67336fe756145eb56d8f92209e51b61aa7d63ca8

Observation 903289dd-14d6-4b02-9e5f-95e5d290c92b · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Imagenet: A large-scale hierarchical image database

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:45.330048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T04:32:44.220762Z digest=sha256:b2584c00d9e6995033b4f84b9f4d37138f1113367d6f35c879f55bf3a5c4383c

Observation 6ce47396-f825-43a8-a6cf-66eae25fd29f · outbound

This paper cites Finding beans in burgers: Deep semantic- visual embedding with localization.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Finding beans in burgers: Deep semantic- visual embedding with localization

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:45.318051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T04:32:44.224063Z digest=sha256:1f3aec47b6735df99465be9cd0eaa0ed5a1c77932986455981bdce6b37cc117b

Observation 953e852d-7825-4d4a-95a3-827dbc5c7084 · outbound

This paper cites Fleet, Jamie Ryan Kiros, and Sanja Fidler.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Fleet, Jamie Ryan Kiros, and Sanja Fidler

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:45.306459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T04:32:44.227298Z digest=sha256:73789e0cbece8b9cdccd0a1f75c088d4663fc03538586e839813a61a190bed1a

Observation 06144a13-524b-496b-b7e6-1f4c216a364c · outbound

This paper cites De- vise: A deep visual-semantic embedding model.Advances in neural information processing systems (NeurIPS), 26, 2013.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval De- vise: A deep visual-semantic embedding model.Advances in neural information processing systems (NeurIPS), 26, 2013

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:45.294893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T04:32:44.230559Z digest=sha256:94029d3612abafb7f88f5ee6d599c74e7efddc2892877c6df9b9fb2204cffcb4

Observation 3313c946-2ae3-4911-999a-8236d4c16ea0 · outbound

This paper cites Understanding the diffi- culty of training deep feedforward neural networks.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Understanding the diffi- culty of training deep feedforward neural networks

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:45.283263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T04:32:44.233871Z digest=sha256:a6f55d148db278b5615861b35fd35673ea223e681dbb6967dfabbfc373e82f61

Observation 6863ac69-5e23-4c13-8d6e-3a8a42c5e877 · outbound

This paper cites Improving image-sentence embeddings using large weakly annotated photo collections.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Improving image-sentence embeddings using large weakly annotated photo collections

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:45.272919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T04:32:44.237175Z digest=sha256:6421b3af15f7e539aa911606a778e8c3be63327c9b3faa0e2f2e2ce42c2fa755

Observation 719f7dc7-c08f-42a5-9630-e5623dc54cf3 · outbound

This paper cites Open-vocabulary object detection via vision and language knowledge distillation.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Open-vocabulary object detection via vision and language knowledge distillation

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:45.262866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T04:32:44.240663Z digest=sha256:6b4c452f350f51a8a671606019ef16782ae8f2a0dac2f6b9e540263f93d00546

Observation 143294bd-f34c-424d-8c02-172b4b6160b1 · outbound

This paper cites Girshick.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Girshick

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:45.251143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T04:32:44.243834Z digest=sha256:38440d2d7be13abb8e1f0145dea53da667a65835f3bf6a80804654824f69fc11

Observation 875578ae-d670-4d24-a181-b266a7f3887d · outbound

This paper cites Generative multi-label zero-shot learning.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Generative multi-label zero-shot learning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:45.240567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T04:32:44.247322Z digest=sha256:45496fea40dca7f098d99ffbae039f3938ba95dd16bff5f029a9228acd51999e

Observation e3af98b8-c68a-4ff9-a994-01a437976edf · outbound

This paper cites Mean average precision map@k metric explained code.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Mean average precision map@k metric explained code

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:45.229970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T04:32:44.250616Z digest=sha256:bcdd7e695191d4c7c561219227d3eb9f277b6ea0d5214584d2e357b92cffe38f

Observation eadb2934-3082-4975-8320-8af75c355b0a · outbound

This paper cites Instance-aware im- age and sentence matching with selective multimodal LSTM.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Instance-aware im- age and sentence matching with selective multimodal LSTM

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:45.219450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T04:32:44.253776Z digest=sha256:7eae976718ffb5ae8b10a7a82fc20054a9bd075d5b753e6436498d2eb4b56d09

Observation 997aa0cb-9f17-408b-b1a0-f8a892d9fe35 · outbound

This paper cites Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T04:32:44.256855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:32:44.256855Z digest=sha256:bb6d59d405aea62dd5ea6d09c48a8c5edfb2c05209ab551e4d83024a363bd870

Observation 9ec5a05a-8bdd-4ae1-9026-fae029bf1a22 · outbound

This paper cites A shared multi-attention framework for multi-label zero-shot learning.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval A shared multi-attention framework for multi-label zero-shot learning

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:45.208593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T04:32:44.260734Z digest=sha256:df62ae795a7e9cf15b655db71c12728bb9c9908b99a87d43b7caa15c89e19f36

Observation 3fa3cc9b-d828-48df-83e3-961b39cba4ff · outbound

This paper cites Saliency-guided attention network for image-sentence matching.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Saliency-guided attention network for image-sentence matching

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:45.197759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T04:32:44.264069Z digest=sha256:9fd7edecfd813f21f94ec4b0d2d236d311c96e379435341e02161bbaa8afb95c

Observation 686a888e-1eba-4f25-8004-13e084b54a3f · outbound

This paper cites Scaling up visual and vision-language representa- tion learning with noisy text supervision.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Scaling up visual and vision-language representa- tion learning with noisy text supervision

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T04:32:44.268216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:32:44.268216Z digest=sha256:79ecca4635c02cac1fa09c4561834f4c8020f2fbf9da13aa12bb2612f18312a0

Observation eafb4ef9-c241-417d-8886-47d9e74930e1 · outbound

This paper cites Billion- scale similarity search with GPUs.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Billion- scale similarity search with GPUs

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:45.180726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T04:32:44.271625Z digest=sha256:f930edcfcf6ffe423c1fd21eabcf0180fc8bccaebd80c84429a76e52238f6899

Observation d8e8ddab-b78e-4966-b790-3786abc2f5d7 · outbound

This paper cites Deep fragment embeddings for bidirectional image sentence map- ping.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Deep fragment embeddings for bidirectional image sentence map- ping

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:45.169751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T04:32:44.274901Z digest=sha256:d176f453af0cd7619e84fe56c2ecb8c1050ce12c16a0f34ceffb1f4a24757d63

Observation 21fc5f81-024b-4a2c-b6a7-7b36f78de260 · outbound

This paper cites Kingma and Jimmy Ba.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Kingma and Jimmy Ba

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T04:32:44.278151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:32:44.278151Z digest=sha256:146bb844938acffd1db5540735af4db9095e851ae4fe6d98ab1cc3dc9d49a6b9

Observation 8a12dc4f-a81c-44cc-a008-49f6370ae8d7 · outbound

This paper cites Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T04:32:44.281134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:32:44.281134Z digest=sha256:85675cadb8c464d2c719356d175e7a2d79c2e0f12c077e511dc2938721316cb0

Observation 95c20d47-1580-4854-a49b-3dc08b8d5f11 · outbound

This paper cites Shamma, Michael S.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Shamma, Michael S

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:45.152272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T04:32:44.284372Z digest=sha256:7e35b4873e4a076544e226652da15f333ba424fcc4ea9b7e228cc18de2ed1ce5

Observation 6df81d8e-eff7-40a4-a593-cdfc267d3d75 · outbound

This paper cites an unresolved cited work.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-11T04:32:45.141789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T04:32:44.287351Z digest=sha256:741b52f08d58cac683a3222cb0e71b1b852cdc85b8e2138f65b6b8f59cd583eb

Observation b3b67e33-dcd7-43e6-9c36-f586355d5870 · outbound

This paper cites an unresolved cited work.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-11T04:32:45.131234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T04:32:44.290556Z digest=sha256:da938dd525155e1ee56fcf9e16a1bb0544c643537b1635d1eacbce52fbd99306

Observation 60c05cb3-e6c2-405b-af93-126efa21482b · outbound

This paper cites Temporal ensembling for semi-supervised learning.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Temporal ensembling for semi-supervised learning

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:45.119941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T04:32:44.293820Z digest=sha256:9cdcd1f56b8aa8e5500508f5cafcebe17141cd8c96ca9a5886c88e146d0bae12

Observation de4c5e1e-d551-435e-a233-d19e650eb4ce · outbound

This paper cites Pseudo-label: The simple and effi- cient semi-supervised learning method for deep neural net- works.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Pseudo-label: The simple and effi- cient semi-supervised learning method for deep neural net- works

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:45.108387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T04:32:44.296862Z digest=sha256:0f40a2f00b1291d66dc02f5b61a6bb5f48f2e98fda63140c594817edb36fa5d8

Observation e3ead0e3-686a-4c35-89a1-81cc86ebb2b4 · outbound

This paper cites Stacked cross attention for image-text match- ing.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Stacked cross attention for image-text match- ing

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:45.097553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T04:32:44.299959Z digest=sha256:f1e62812db48cbab84a688f3481f8fb3a02dd79cadfb57f1fe1af76878177fe8

Observation 9dc3596a-371b-4a4f-9ae1-751f4c6e4a23 · outbound

This paper cites Object- centric open-vocabulary image retrieval with aggregated fea- tures.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Object- centric open-vocabulary image retrieval with aggregated fea- tures

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:45.086738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T04:32:44.303233Z digest=sha256:51f34d5102089a3947d50d085bf97099999aecee123340ba797902d5d56b923a

Observation 146a311d-96ca-4ecb-873c-1483b1c4eb96 · outbound

This paper cites an unresolved cited work.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-11T04:32:45.066222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T04:32:44.310751Z digest=sha256:c12aaa0df83d6b9c5a3148682e63f698ed79b8f076bbb42acfe732db41aa8970

Observation e34b8b82-f596-4a95-a4ee-346ed7268fc4 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for uni- fied vision-language understanding and generation.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Blip: Bootstrapping language-image pre-training for uni- fied vision-language understanding and generation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T04:32:44.314386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:32:44.314386Z digest=sha256:eca87d99ec0fe0b95dc2c55dafc2498ca996f7650c5a9312ee8e08c2d415f389

Observation 59b5766e-5fd1-4b2e-b8f9-6b2f750695a0 · outbound

This paper cites Selvaraju, Akhilesh Gotmare, Shafiq R.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Selvaraju, Akhilesh Gotmare, Shafiq R

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:45.050149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T04:32:44.317593Z digest=sha256:073dcd211791dda7266c681dece2ed9045f8a86e049cfed254bbbaf68c61ae65

Observation fd915810-0a36-4df9-a6ea-3e382c155e70 · outbound

This paper cites Adapting CLIP For Phrase Localization Without Further Training.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Adapting CLIP For Phrase Localization Without Further Training

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T04:32:44.321352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:32:44.321352Z digest=sha256:2dbbd7488cc6b5b5626f89b3347b1cc80211eed513e114d23b38e73d4d1b62e4

Observation d046fad0-f1d1-480c-8022-6a12fe3da4a4 · outbound

This paper cites Visual semantic reasoning for image-text matching.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Visual semantic reasoning for image-text matching

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:45.039224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T04:32:44.325446Z digest=sha256:86f5f9a94062347d918703a81bb53e76abc6d97ced6d5b64fbd43d381f57d7f1

Observation 73bd46c6-5fc3-4ab1-a26a-26cf35164df4 · outbound

This paper cites Oscar: Object-semantics aligned pre-training for vision-language tasks.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Oscar: Object-semantics aligned pre-training for vision-language tasks

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:45.028368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T04:32:44.328831Z digest=sha256:eec77f84acc176a2df964cad822d670becdb2fdab64d449c35d1b332ae9c2bd6

Observation a9d6e88f-414b-42f1-b0b8-507a899234a1 · outbound

This paper cites Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll ´ar, and C.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll ´ar, and C

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:44.897271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T04:32:44.336016Z digest=sha256:aa057e3188807dade52a4e89a91446f9d6dfc63d4f7d045cca44015ddb39578f

Observation 2b349096-06b7-42cb-a4c6-9cb48cf6f698 · outbound

This paper cites OVIS: open-vocabulary visual instance search via visual-semantic aligned representation learning.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval OVIS: open-vocabulary visual instance search via visual-semantic aligned representation learning

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:44.887428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T04:32:44.340173Z digest=sha256:55939d3a3a878e0a467b9d12f00c31bb7ddce88bc7bb5430cb5a93ea28e85c30

Observation 16a27aa1-d413-4d77-8973-cf788d4a5f0c · outbound

This paper cites Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:44.877132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T04:32:44.343611Z digest=sha256:37afff4617dc63a7827ac4cc1c207cd350cb63771fd5b10bdc33e4ee1b16bc31

Observation d94d9d91-28c5-4225-994f-2c6853b905e0 · outbound

This paper cites Simple open-vocabulary object detection with vi- sion transformers.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Simple open-vocabulary object detection with vi- sion transformers

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:44.867278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T04:32:44.347063Z digest=sha256:c1dcbc3da6cccb103b5d4b3a920158701a79b8f17bd8809f2dc8bdcd682970a8

Observation 7588f2e9-508f-4bb2-a304-3e9c275e8002 · outbound

This paper cites Gritsenko, and Neil Houlsby.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Gritsenko, and Neil Houlsby

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:44.856527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T04:32:44.350738Z digest=sha256:bceefa5fa3d044a425143a66078b5b34701f7e2d4ec7204ec43d6e30d7c2ad89

Observation 5bf0f9be-7dff-4b7f-8e14-5aa812923dba · outbound

This paper cites an unresolved cited work.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-11T04:32:44.845186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T04:32:44.354036Z digest=sha256:ce5766c1784026d0ce703bc860032d5e144b23df749f374b6f9aa869c5bcdf82

Observation 09142528-ce7a-4f0b-ac31-6b5639a54405 · outbound

This paper cites Zero-Shot Learning by Convex Combination of Semantic Embeddings.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Zero-Shot Learning by Convex Combination of Semantic Embeddings

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T04:32:44.357178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:32:44.357178Z digest=sha256:a908ef55b5137ba794a7baab29548eff500d4b4e6c0a42c0f8c5b485414cd90e

Observation 1fc4f01e-9860-440f-ae21-fb0089197ab5 · outbound

This paper cites Meta pseudo labels.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Meta pseudo labels

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:44.834732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T04:32:44.361301Z digest=sha256:7bb4dc95dbc79abca6c6bb4ddccf67c8f9365cb688af635b9315dd19a514353b

Observation 3fa9b5d8-4ad4-4319-8180-845013a1315b · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Learning transferable visual models from natural language supervi- sion

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:44.823825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T04:32:44.364047Z digest=sha256:bc8fb3f8ca3deb0288e532ef67b192250acb441615afa927a2f4046b1ac87801

Observation 8279678d-be3a-4a0f-907b-8b02f10f7380 · outbound

This paper cites Learning transferable visual models from natural language supervision.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Learning transferable visual models from natural language supervision

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:44.812922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T04:32:44.366883Z digest=sha256:ba3b32a6dcffadfd830015ebcb545db618e0a691931af32d6dd6b45b403a94a9

Observation 32507526-2ecc-4074-846f-32b1062fe148 · outbound

This paper cites Denseclip: Language-guided dense prediction with context- aware prompting.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Denseclip: Language-guided dense prediction with context- aware prompting

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:44.802015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T04:32:44.369964Z digest=sha256:a4d7b9302514d4653638c121097eceaf16c19de5fb790d81a1dbaf502628c572

Observation 344bb416-94ed-4c4b-afe8-07253d2c8211 · outbound

This paper cites Regularization with stochastic transformations and perturba- tions for deep semi-supervised learning.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Regularization with stochastic transformations and perturba- tions for deep semi-supervised learning

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T04:32:44.372916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:32:44.372916Z digest=sha256:97b3a1e0634443241c5591d69c61e37a42c5da13dd759034df5cde5338940940

Observation f9390e8f-d0e3-4351-a377-a724965289cc · outbound

This paper cites Objects365: A large-scale, high-quality dataset for object detection.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Objects365: A large-scale, high-quality dataset for object detection

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:44.783806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T04:32:44.376152Z digest=sha256:7a48971d362b616a0dfc316fd30a6b7b8d91f18d29fa6a0a1a23507416f82896

Observation 709102e8-1cd6-4c97-8075-a6903002902b · outbound

This paper cites Fixmatch: Simplifying semi-supervised learning with consistency and confidence.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Fixmatch: Simplifying semi-supervised learning with consistency and confidence

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T04:32:44.379094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:32:44.379094Z digest=sha256:960effbcf5257e568babbf3c8f3d1a56404e0cb87c3be4b4468c5d359afea96d

Observation 0a9d0bf0-888d-426d-89a1-2fc076229288 · outbound

This paper cites A Simple Semi-Supervised Learning Framework for Object Detection.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval A Simple Semi-Supervised Learning Framework for Object Detection

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T04:32:44.381983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:32:44.381983Z digest=sha256:8ff3c337d1948dbc0edf434998b8c28f5b615c473f179784708a68e5446fac2a

Observation ea46336a-b1d4-4448-9650-8cdf05a0cf7b · outbound

This paper cites Dualcoop: Fast adaptation to multi-label recognition with limited annotations.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Dualcoop: Fast adaptation to multi-label recognition with limited annotations

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:44.765945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T04:32:44.385620Z digest=sha256:a6bbe13e2785115d0af4515307e7d515b298e64e903d323e5016ded3a2d8fe60

Observation 4b43bf7a-9756-4073-aead-773778dcfe73 · outbound

This paper cites LXMERT: learning cross- modality encoder representations from transformers.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval LXMERT: learning cross- modality encoder representations from transformers

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:44.755377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T04:32:44.388942Z digest=sha256:058781240319a61f4f4d97b0cbe733cd5515c8986a3d5cd3d02afafe254676e6

Observation c7ad6b8d-16d5-47a5-a1e2-e70438683180 · outbound

This paper cites GIT: A generative image-to-text transformer for vision and language.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval GIT: A generative image-to-text transformer for vision and language

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:44.744179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T04:32:44.392434Z digest=sha256:64f047dfd8bf8a3c3a47207d08448f0f7e0852ff963d3fd75829f83df662008b

Observation ef0f87d3-b8a2-405e-a4c5-f16b7fdc7551 · outbound

This paper cites Object-aware dis- tillation pyramid for open-vocabulary object detection.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Object-aware dis- tillation pyramid for open-vocabulary object detection

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:44.733306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T04:32:44.395836Z digest=sha256:b6e1d24bfc313beb02e35bc596bf89689e8c07b93746db59d91eb221f14159cd

Observation a81d8b9c-e971-4e63-a19f-8ba16a6494e7 · outbound

This paper cites Simvlm: Simple visual language model pretraining with weak supervision.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Simvlm: Simple visual language model pretraining with weak supervision

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-11T04:32:44.399222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:32:44.399222Z digest=sha256:19ce236233a6c415b90b0bd8dd21802d460bd9fec8e6b11d401021a13c745a56

Observation 49996732-0701-4f2d-9d60-70578c39dc16 · outbound

This paper cites Aligning bag of regions for open- vocabulary object detection.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Aligning bag of regions for open- vocabulary object detection

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:44.716100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T04:32:44.402977Z digest=sha256:11929583f6b80e57b46ac6c190fbefad12ab11d9224b618b893340342bf11799

Observation 498f067e-e35f-4230-8853-ecb759bcd7db · outbound

This paper cites CLIPSelf: Vision transformer distills itself for open-vocabulary dense predic- tion.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval CLIPSelf: Vision transformer distills itself for open-vocabulary dense predic- tion

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:44.706543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T04:32:44.406238Z digest=sha256:e50015675c7f54ea2da104dd348fa70d99e3733750c7a77a8119dc57063fbb85

Observation 81b8bdef-5eb0-46d9-9bba-cabc7ba1b374 · outbound

This paper cites CLIM: contrastive language-image mosaic for region representation.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval CLIM: contrastive language-image mosaic for region representation

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:44.696334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T04:32:44.409651Z digest=sha256:8bb740caeacf54bbcedf57f9d1fe20fa43bba6613563af9959370a7d59361f1f

Observation 580337d6-a7cf-4fc9-b58b-182f5997df2c · outbound

This paper cites CORA: adapting CLIP for open-vocabulary detection with region prompting and anchor pre-matching.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval CORA: adapting CLIP for open-vocabulary detection with region prompting and anchor pre-matching

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:44.686378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T04:32:44.412748Z digest=sha256:491b321c8023cf84f272cd5db059ba4a2d59b656f235b0e833a6f0fc4ca0b03e

Observation 700ef036-57bb-4c77-ab99-8e0878c5917f · outbound

This paper cites Unsupervised data augmentation for consistency training.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Unsupervised data augmentation for consistency training

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:44.676599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T04:32:44.415985Z digest=sha256:3323d60b5695ef60971cdbab2e960abf308cd0e14a1ea7fa1c7920084759384a

Observation b744b86f-6194-4436-b81d-dd6704429373 · outbound

This paper cites Self-training with noisy student improves imagenet classification.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Self-training with noisy student improves imagenet classification

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:44.666904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T04:32:44.419323Z digest=sha256:37a52803f13d6023f4d01f4b89c279ad176470433bca1ad25077a7feba9bc4ba

Observation 048912da-d9df-44e4-b709-f355f2b41bd4 · outbound

This paper cites From image descriptions to visual denotations: New similarity metrics for semantic inference over event descrip- tions.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval From image descriptions to visual denotations: New similarity metrics for semantic inference over event descrip- tions

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:44.657502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T04:32:44.422649Z digest=sha256:cd81af764d55dc07575d05481c9366929a09a95bd336f0fa7e5fe6ae90de7b29

Observation bd62beab-7379-4888-8130-214531072f82 · outbound

This paper cites Coca: Contrastive captioners are image-text foundation models.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Coca: Contrastive captioners are image-text foundation models

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:44.647759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T04:32:44.426262Z digest=sha256:eae217f623d50459130e41ab0d400dab652b1cef8bf33787ddf11f53797d24b0

Observation b4d89a9c-6988-460f-b56c-92e848b601fd · outbound

This paper cites Florence: A New Foundation Model for Computer Vision.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Florence: A New Foundation Model for Computer Vision

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-11T04:32:44.429626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:32:44.429626Z digest=sha256:59c93c10be61fad3e146ac8dc6c1a37b272a87cfcb1f9f264e2c9f5be94822a4

Observation d208642c-8131-4507-af4f-51998798543d · outbound

This paper cites Lit: Zero-shot transfer with locked-image text tuning.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Lit: Zero-shot transfer with locked-image text tuning

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:44.637917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T04:32:44.433395Z digest=sha256:db08f7b6e5868f6c2b507803106b13ff1ef19445a7e65c76c8dbb53b8ecc30a1

Observation ef404a21-badb-4213-b54b-cac7d18684fa · outbound

This paper cites Fast zero- shot image tagging.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Fast zero- shot image tagging

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:44.627313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T04:32:44.436720Z digest=sha256:fe54717fe09c569aa8d22f89db5695ebfe15cafa26bc37d44bb3294fae427ec0

Observation ee36940d-78f8-4565-a92f-1b331c27776f · outbound

This paper cites Exploiting unlabeled data with vision and language models for object detection.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Exploiting unlabeled data with vision and language models for object detection

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:44.616505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T04:32:44.439922Z digest=sha256:538acfc7108a58faba12144ecc5d64b8f77d75e9cb6b88034e5c25122a841335

Observation e2e330fd-2ab1-4fbe-94ce-3da4a3e6a41e · outbound

This paper cites Regionclip: Region-based language-image pretraining.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Regionclip: Region-based language-image pretraining

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:44.605565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T04:32:44.443115Z digest=sha256:a7f13931f6a80f0b2ab1db6dd744ebff8e89418eecbb3d62be59d9fb6c01267e

Observation b62ee730-0a93-42a7-87f4-f0bb529cb6f6 · outbound

This paper cites Extract free dense labels from CLIP.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Extract free dense labels from CLIP

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:44.594224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T04:32:44.446293Z digest=sha256:e6563d75c44ce0251867aeaf39fbd29da76ff4a939e5a857ff862b81f3ce371e

Observation eb1e4021-4501-4b31-ad38-015d563c7db8 · outbound

This paper cites Detecting twenty-thousand classes using image-level supervision.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Detecting twenty-thousand classes using image-level supervision

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:44.583536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T04:32:44.449513Z digest=sha256:99b9b049cc3d134a8634bff64072ce8e1654b8b8160d29609261c07e00af9cb2

Observation 95111b72-017c-45d8-bace-6ee30db8b557 · outbound

This paper cites Semi-supervised learning literature survey.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Semi-supervised learning literature survey

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:44.573034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T04:32:44.452756Z digest=sha256:81d9f75e3ef69a95e9dcaf069e06319338c50c8f5ee3cdc5557223bbf93da498

Observation 351aa591-742d-4d2d-89d0-32ec6b46c002 · outbound

This paper cites Rethinking pre- training and self-training.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Rethinking pre- training and self-training

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:44.562509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T04:32:44.456090Z digest=sha256:301e394d26371cd8d8249f4e48d5fb66162a78b0d22fa2731b6ef7a56048800a

Observation e59c07a9-5278-4433-89ee-a27e483a05aa · outbound

This paper cites an unresolved cited work.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Unresolved cited work

Reference 85

Resolution
malformed identifier
raw_fallback, observed 2026-08-11T04:32:44.550453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T04:32:44.459932Z digest=sha256:ff6d681232a9174bec4859e7140192a59b1c94940e2d16697c6fe571448c8af9

Observation bf8ebb87-d2c6-425d-a6f2-0de489193ef3 · outbound

This paper cites an unresolved cited work.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Unresolved cited work

Reference 137

Resolution
parse uncertain
raw_fallback, observed 2026-08-11T04:32:44.906987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T04:32:44.332198Z digest=sha256:093757aead2f0cdf2fe74ea27293547a4212c819dc0d2bc8220deb8aef132860

Observation c7121743-9217-492c-aba3-3e4c4fe05e1a · outbound

This paper cites 1, 2, 3, 4, 6, 8, 13.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval 1, 2, 3, 4, 6, 8, 13

Reference 608

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:32:45.076156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T04:32:44.307277Z digest=sha256:9690452ecff739db2ba1b627322caf7d89e6c44f3d96ab74ac9eb24109ff28ae

Observation 17f61c72-0b65-4ef8-9c9d-831b1e11f4cd · outbound

This paper cites an unresolved cited work.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Unresolved cited work

Reference 6086

Resolution
unresolved
no resolver link, observed 2026-08-11T04:32:44.180725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:32:44.180725Z digest=sha256:a0c2fb11d5cad2a3703906a47f0c48f6bac478065b3665b1a2518bb03c9a6c70

Pith citing papers

No inbound Pith citation observations are available.