Pith. sign in

Paper Citation Record · LEDGER

Fine-Grained Open-Vocabulary Object Recognition via User-Guided Segmentation

As of 13 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2411.15620.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.15620 v1

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T14:08:51.045860Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact0
  • verified fuzzy17
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 50138ab6-374a-454f-bca6-0960389ac1e5 · outbound

This paper cites Yolo-world: Real-time open- vocabulary object detection.

Fine-Grained Open-Vocabulary Object Recognition via User-Guided Segmentation Yolo-world: Real-time open- vocabulary object detection

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:52.782620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:08:50.799557Z digest=sha256:2fdec1db0e16e25c91f5419bafee7593812ef6d6af4a1ae149a849eb29ecca6b

Observation 3ad24b57-c0b5-4f24-a988-9c873b24aef4 · outbound

This paper cites Everingham, S.

Fine-Grained Open-Vocabulary Object Recognition via User-Guided Segmentation Everingham, S

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T14:08:50.808537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:08:50.808537Z digest=sha256:5c64d42349b48dec936da04b837c1122987b6dd9246db91038a5ff07a96fe382

Observation fa83f4eb-2b2a-4bde-9434-f4c78f9a6aed · outbound

This paper cites Salience detr: Enhancing detection trans- former with hierarchical salience filtering refinement.

Fine-Grained Open-Vocabulary Object Recognition via User-Guided Segmentation Salience detr: Enhancing detection trans- former with hierarchical salience filtering refinement

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:52.619603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:08:50.816298Z digest=sha256:9f4a9c5f70e716e531c4c4dc965bead9f2e48c596c64042cd530b5d04ac56721

Observation 210ce89d-a7c3-4dc2-9d01-d0851234644f · outbound

This paper cites Relation detr: Exploring explicit position relation prior for object detection.

Fine-Grained Open-Vocabulary Object Recognition via User-Guided Segmentation Relation detr: Exploring explicit position relation prior for object detection

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:52.547675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:08:50.826991Z digest=sha256:80ac2844ac02b9f0425abcf5712ed490586b7db58123e7b1ddb6fe3502a35f6a

Observation 79c4197b-d459-4276-963d-830d183012e8 · outbound

This paper cites Scaling up visual and vision-language representation learning with noisy text supervision.

Fine-Grained Open-Vocabulary Object Recognition via User-Guided Segmentation Scaling up visual and vision-language representation learning with noisy text supervision

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:52.495515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:08:50.832716Z digest=sha256:2f318a7bc7b7c760503e91f847e4ec2b10cbcea9d015f00dde2fb88406fb8289

Observation 337d7243-e6e6-4a86-a567-435a4b096caa · outbound

This paper cites Rea- soning grasping via multimodal large language model.

Fine-Grained Open-Vocabulary Object Recognition via User-Guided Segmentation Rea- soning grasping via multimodal large language model

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:52.387944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:08:50.840529Z digest=sha256:1a5befaf23fe018df0371031ba2e84ffb86a32cc645a5888baf8101013541780

Observation 35747dbc-4335-47ce-81f3-21932321792e · outbound

This paper cites Segment Anything.

Fine-Grained Open-Vocabulary Object Recognition via User-Guided Segmentation Segment Anything

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T14:08:50.850774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:08:50.850774Z digest=sha256:5621e6b90407ac255b6d3a3d6ed887e3834dfe666e3d9ac7a9a16afad33fa8f8

Observation 6219f2b7-5734-4e37-b39b-e90fc803bdd4 · outbound

This paper cites Lisa: Reasoning segmentation via large language model.

Fine-Grained Open-Vocabulary Object Recognition via User-Guided Segmentation Lisa: Reasoning segmentation via large language model

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:52.231673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:08:50.861734Z digest=sha256:087af10afa5c1bd340f7d5f5d41b82dac65e8f2845b0bc4418582eb613d6217d

Observation fb9f3cba-f79d-4b78-807a-58ce65695933 · outbound

This paper cites Grounded language-image pre-training.

Fine-Grained Open-Vocabulary Object Recognition via User-Guided Segmentation Grounded language-image pre-training

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:52.193813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:08:50.873299Z digest=sha256:52cbba1af0338389f4514eed227e04f7bc3699b905d2ca8d4aa46e717503d6f2

Observation c30fb905-77d4-4b98-b6a0-18c71d00e3c2 · outbound

This paper cites Microsoft COCO: Common Objects in Context.

Fine-Grained Open-Vocabulary Object Recognition via User-Guided Segmentation Microsoft COCO: Common Objects in Context

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T14:08:50.882876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:08:50.882876Z digest=sha256:c2645c8d07aaff58ca78c9e9d90d70f3d7dee0d546f01e9d07424f90623b7c1a

Observation ec1aa3e9-5be7-43c4-88dd-05116502b2bd · outbound

This paper cites Grounding dino: Marrying dino with grounded pre-training for open-set object detection.

Fine-Grained Open-Vocabulary Object Recognition via User-Guided Segmentation Grounding dino: Marrying dino with grounded pre-training for open-set object detection

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:52.153977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:08:50.892988Z digest=sha256:1b4d85ee07991cd4b2759bb6c5ac08861573e190db4592a079ddddfb34ac4b24

Observation be5e2057-faa5-49aa-a18c-ca5c110d3ac6 · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows.

Fine-Grained Open-Vocabulary Object Recognition via User-Guided Segmentation Swin transformer: Hierarchical vision transformer using shifted windows

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:52.021915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:08:50.903311Z digest=sha256:dad102714f407612d17038f26b50fc9f3db31e2d50389721d08ddd64aedce9d2

Observation 7f1415ec-09b3-4030-b373-465973c58d77 · outbound

This paper cites Scaling open-vocabulary object detection.

Fine-Grained Open-Vocabulary Object Recognition via User-Guided Segmentation Scaling open-vocabulary object detection

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:51.904431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:08:50.909609Z digest=sha256:dde724684603ca29af16981f1692ded77baf513500d4f3937dd1314ed9a470c8

Observation 92d8ffe7-0b9f-4867-8d12-9ae52b9c1c4b · outbound

This paper cites GPT-4 Technical Report.

Fine-Grained Open-Vocabulary Object Recognition via User-Guided Segmentation GPT-4 Technical Report

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T14:08:50.917523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:08:50.917523Z digest=sha256:20398c371f614a536a83d6f4ff9fdfddf3d2e3b896853cd50ce2ae86510680c4

Observation ad72e87f-d283-4eb8-b513-eb863308019b · outbound

This paper cites Learning transferable visual models from natural language supervision.

Fine-Grained Open-Vocabulary Object Recognition via User-Guided Segmentation Learning transferable visual models from natural language supervision

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T14:08:50.925751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:08:50.925751Z digest=sha256:d50bd7c6a67beb0969e62d62f9f947d601a33bc639991d3becb0b698dcb77fc9

Observation 8f65144c-6981-4181-b610-1fb1c1de4d63 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

Fine-Grained Open-Vocabulary Object Recognition via User-Guided Segmentation Learning Transferable Visual Models From Natural Language Supervision

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T14:08:50.933022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:08:50.933022Z digest=sha256:7e7adf378b38da97251e6fae7db799b3262afcc854e071578b528899462e5040

Observation 3b04b5d5-8898-4ba3-99ab-4ee6f4312242 · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

Fine-Grained Open-Vocabulary Object Recognition via User-Guided Segmentation SAM 2: Segment Anything in Images and Videos

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T14:08:50.940066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:08:50.940066Z digest=sha256:33637c8bc1a20aae879c2cf0e0751a959bf3e79d6e7822e9d71652758bb90e61

Observation b0b50802-1b23-45ee-bac7-b2f08ca6c85a · outbound

This paper cites Yolov3: An incremental improvement.

Fine-Grained Open-Vocabulary Object Recognition via User-Guided Segmentation Yolov3: An incremental improvement

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:51.855646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:08:50.947093Z digest=sha256:e7257b9f06db6257dfdcd0cb2e9ab2f085591be0068277638d1ab6c28b8abedf

Observation 1bacba83-b9aa-4812-ad7f-277872890c04 · outbound

This paper cites Faster r-cnn: Towards real-time object detection with region proposal networks.

Fine-Grained Open-Vocabulary Object Recognition via User-Guided Segmentation Faster r-cnn: Towards real-time object detection with region proposal networks

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:51.817071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:08:50.955086Z digest=sha256:da44df4ea3ffc1dc52b0ab60f03cc68962f9059eba81e437e4170c8bc91d515d

Observation 41ba8b2b-e49b-48e4-a1d4-31daaef17d8e · outbound

This paper cites Grounded sam: Assembling open-world models for diverse visual tasks,.

Fine-Grained Open-Vocabulary Object Recognition via User-Guided Segmentation Grounded sam: Assembling open-world models for diverse visual tasks,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T14:08:50.961143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:08:50.961143Z digest=sha256:744341a3d25fa972a42ed83396eff01e1f7ae3881f12e4ab4cd5679ca30fae0c

Observation d51c7704-469e-420d-81d3-8d856f51345a · outbound

This paper cites Grounding language models for visual entity recognition.

Fine-Grained Open-Vocabulary Object Recognition via User-Guided Segmentation Grounding language models for visual entity recognition

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:51.660960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:08:50.972021Z digest=sha256:fed58332fbfb76cd7415c6fbcc994c7a2212e8384b4e7fa1a9f413edb16d4129

Observation 1bd54eb3-1b2f-4d2e-ac0a-c57aea135967 · outbound

This paper cites Visa: Reasoning video object segmentation via large language models.

Fine-Grained Open-Vocabulary Object Recognition via User-Guided Segmentation Visa: Reasoning video object segmentation via large language models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:51.547226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:08:50.987492Z digest=sha256:8be1905f00076930b7592c06b7adb257a15d8505e23de82c56cce57d5013f61d

Observation acc33548-8d83-4de0-98c3-27d49d270014 · outbound

This paper cites Detclipv3: Towards versatile generative open-vocabulary object detec- tion.

Fine-Grained Open-Vocabulary Object Recognition via User-Guided Segmentation Detclipv3: Towards versatile generative open-vocabulary object detec- tion

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:51.499368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:08:51.001752Z digest=sha256:6c73a82121527f7d9aee5790c2459f94dccc79cc34b4e496f29aef827ec46096

Observation ea0853c3-9ac0-4c9b-a826-d060fdf95993 · outbound

This paper cites Ferret: Refer and ground anything anywhere at any granularity.

Fine-Grained Open-Vocabulary Object Recognition via User-Guided Segmentation Ferret: Refer and ground anything anywhere at any granularity

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T14:08:51.016112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:08:51.016112Z digest=sha256:885cf803b87ae71a9d61234978bcfa21ccb6f78519140428858adcde3c7a9c55

Observation c7a90017-2aa7-4879-b148-a75a3c757406 · outbound

This paper cites A simple framework for open-vocabulary segmentation and detection.

Fine-Grained Open-Vocabulary Object Recognition via User-Guided Segmentation A simple framework for open-vocabulary segmentation and detection

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:51.422003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:08:51.024807Z digest=sha256:eaf84212e490b6b8ff382f26b0191973c7ec1275912a780916d33303cd09b9e4

Observation 7e5658a2-92f1-4125-939a-109bbe38a5c6 · outbound

This paper cites Recognize Anything: A Strong Image Tagging Model.

Fine-Grained Open-Vocabulary Object Recognition via User-Guided Segmentation Recognize Anything: A Strong Image Tagging Model

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T14:08:51.038158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:08:51.038158Z digest=sha256:b15a8bf76705528b9755b5f6aa7c7669a66065036cffd91ae004b9d687016f94

Observation 16020361-37af-4ae2-bde3-c36d1b42f02e · outbound

This paper cites Detrs with col- laborative hybrid assignments training.

Fine-Grained Open-Vocabulary Object Recognition via User-Guided Segmentation Detrs with col- laborative hybrid assignments training

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:51.373602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:08:51.045860Z digest=sha256:507b92dbc3e46bf8db93e53a77a6b6afb570f0001812ca9360155ca9068c3b36

Pith citing papers

No inbound Pith citation observations are available.