Pith. sign in

Paper Citation Record · LEDGER

Teaching VLMs to Localize Specific Objects from In-context Examples

As of 13 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 0 inbound Pith citation observations for arXiv:2411.13317.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.13317 v2

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T16:39:40.687439Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

61 of 61 outbound references displayed

  • verified exact1
  • verified fuzzy34
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 37ddf28d-cbe8-43cd-9fdf-5bf0c0da1bee · outbound

This paper cites Pixtral 12B.

Teaching VLMs to Localize Specific Objects from In-context Examples Pixtral 12B

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T16:39:40.456544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:39:40.456544Z digest=sha256:c135c63270cd2cc35bea2951fdcc4c6a0e32cfef31293fedbc7a64d4adfe53f6

Observation 6a83b50c-0057-4a65-98c3-638a881f30a4 · outbound

This paper cites Flamingo: a Visual Language Model for Few-Shot Learning.

Teaching VLMs to Localize Specific Objects from In-context Examples Flamingo: a Visual Language Model for Few-Shot Learning

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:39:41.549331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:39:40.461608Z digest=sha256:1b3ee38fce1e89986efb95e5e39532cc2d2cbb74731f8159bbf7f2484da8cfcc

Observation acb2a615-02de-4be6-8190-6425fbe40cbd · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Teaching VLMs to Localize Specific Objects from In-context Examples Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T16:39:40.465499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:39:40.465499Z digest=sha256:4de71f8973ba49f250c74f3b0ba8e647cf5810d775402995d4d998e6fc1dc526

Observation 7a938626-bc81-4701-9e58-4cf7c375e58d · outbound

This paper cites DeciMamba: Exploring the Length Extrapolation Potential of Mamba.

Teaching VLMs to Localize Specific Objects from In-context Examples DeciMamba: Exploring the Length Extrapolation Potential of Mamba

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T16:39:40.469752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:39:40.469752Z digest=sha256:2a96cfe84fb3e449432f34d4f1f9e411d06345720f069e6bcf7ed9f8a64cbdaf

Observation 3c5f6a9b-5fae-489e-a060-19bbac3f903d · outbound

This paper cites Lan- guage models are few-shot learners.

Teaching VLMs to Localize Specific Objects from In-context Examples Lan- guage models are few-shot learners

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T16:39:40.474895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:39:40.474895Z digest=sha256:7e33258908b92b2c67bc5c3fb63a33b5a1e704f3d16276a0981027ed6f7407fd

Observation 1ce28c8d-a6d9-4b19-b1b8-e17f57bd4801 · outbound

This paper cites an unresolved cited work.

Teaching VLMs to Localize Specific Objects from In-context Examples Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-12T16:39:41.529173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:39:40.478781Z digest=sha256:edd539539de5652a5ac14f2da321c982d14df59018ddc5d4e9a4acf841e185cf

Observation a547f13d-6ce2-4733-aece-8b0ed6f60aee · outbound

This paper cites MiniGPT-v2: Large Language Model as a Unified Interface for Vision-Language Multi-task Learning.

Teaching VLMs to Localize Specific Objects from In-context Examples MiniGPT-v2: Large Language Model as a Unified Interface for Vision-Language Multi-task Learning

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:39:41.517511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:39:40.482901Z digest=sha256:96e02bbc20e4e5fc41b42984b1d91af834fc377311de5b95c6b527827a0a5da1

Observation 4bb4b678-9d7f-4dc3-965d-839f697798bd · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

Teaching VLMs to Localize Specific Objects from In-context Examples How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T16:39:40.487264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:39:40.487264Z digest=sha256:f8fea053a41221d6ad3705a61291088cbcd6501661b827feec688bc29c9559a5

Observation 660e31e1-6833-46ff-aaaf-41de340fa4aa · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

Teaching VLMs to Localize Specific Objects from In-context Examples Gonzalez, Ion Stoica, and Eric P

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:39:41.506480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:39:40.491205Z digest=sha256:a0a3c0c08e01ac127f6d90cfa74cb65f64aed0e51e6d1816fe77f1d8ddab9819

Observation 56991c9d-8caa-4af0-b942-95d8e8e38f2f · outbound

This paper cites InstructBLIP: Towards General-purpose Vision- Language Models with Instruction Tuning.

Teaching VLMs to Localize Specific Objects from In-context Examples InstructBLIP: Towards General-purpose Vision- Language Models with Instruction Tuning

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:39:41.494396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:39:40.494839Z digest=sha256:0a223c595d9f9d429253cdb93bf1f25c6f4822ef3efdc65484e5162facae9a26

Observation fefc5034-5bc1-4149-a8c1-52a3968fae0d · outbound

This paper cites Tao: A large-scale bench- mark for tracking any object.

Teaching VLMs to Localize Specific Objects from In-context Examples Tao: A large-scale bench- mark for tracking any object

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:39:41.482721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:39:40.498165Z digest=sha256:5a333f343e27894e282e1df92e518c506e23ea58624d467a5667fae5f0aa23e5

Observation 72c2c290-9dbf-4842-897e-d2185b5ee624 · outbound

This paper cites Dense and Aligned Captions (DAC) Promote Compositional Reasoning in VL Models.

Teaching VLMs to Localize Specific Objects from In-context Examples Dense and Aligned Captions (DAC) Promote Compositional Reasoning in VL Models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:39:41.469865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:39:40.501315Z digest=sha256:37d178e06c270d7b0107718be2a19cbe085cdc821cc1406d5c5b2c2616c63383

Observation aa7b3741-3448-489a-819e-b80b7ec9db98 · outbound

This paper cites Teaching structured vision & language concepts to vision & language models.

Teaching VLMs to Localize Specific Objects from In-context Examples Teaching structured vision & language concepts to vision & language models

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:39:41.456962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:39:40.505137Z digest=sha256:7d880519f0c6e03ea377026c01b960ba7516bb071db1defa27a1f09cb4c8108b

Observation f9cbc167-0a17-4e7e-8298-5527f44f95f3 · outbound

This paper cites Towards Multimodal In-Context Learning for Vision & Language Models.

Teaching VLMs to Localize Specific Objects from In-context Examples Towards Multimodal In-Context Learning for Vision & Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T16:39:40.508301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:39:40.508301Z digest=sha256:68196d80666444eabff3164512e466dc7f934ffad0811cfe6ae15fa0647ad179

Observation 8b7704a9-01ed-4ebe-adf9-a5a4c901f8a0 · outbound

This paper cites The Llama 3 Herd of Models.

Teaching VLMs to Localize Specific Objects from In-context Examples The Llama 3 Herd of Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T16:39:40.512501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:39:40.512501Z digest=sha256:dc219006398fc15a57def602c35f2f1df17da2f44129eae27fb4290bc47ce83e

Observation 8c326ce3-f619-4baa-80a9-b62541567166 · outbound

This paper cites Lasot: A high-quality benchmark for large-scale single ob- ject tracking.

Teaching VLMs to Localize Specific Objects from In-context Examples Lasot: A high-quality benchmark for large-scale single ob- ject tracking

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T16:39:40.516794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:39:40.516794Z digest=sha256:97b032dadb2ef2de81571abaa62aa40d0d94bfc8e51b7fa905a607419dd78534

Observation 03581bc5-4c3b-42cb-819c-1eeece0d12ca · outbound

This paper cites SEED: Self-supervised Dis- 9 tillation for Visual Representation.

Teaching VLMs to Localize Specific Objects from In-context Examples SEED: Self-supervised Dis- 9 tillation for Visual Representation

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:39:41.438365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:39:40.520549Z digest=sha256:f04eccb01d5a1b5c2a371cf0b635da52057f431b5927ad398c50f8e6ac4617c3

Observation 3124bf49-f3ba-4c40-aa7a-62ec7484cc3f · outbound

This paper cites Cross-domain few-shot object detection via enhanced open-set object detector.

Teaching VLMs to Localize Specific Objects from In-context Examples Cross-domain few-shot object detection via enhanced open-set object detector

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:39:41.424973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:39:40.524498Z digest=sha256:c3e669c3a99ae4b232db5aee65d7212da9aff6e1c1752e3a66abf821f1dbccf2

Observation d571833c-e27a-4c74-b083-dcc1f19b7e47 · outbound

This paper cites Vision-Language Models Create Cross-Modal Task Representations.

Teaching VLMs to Localize Specific Objects from In-context Examples Vision-Language Models Create Cross-Modal Task Representations

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-12T16:39:41.022727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:39:40.528370Z digest=sha256:3d7dbb4fd7a993ed91fdb1b76f07a0fb05323c53804b549f418385db2a0ead9f

Observation f1cbd9f0-fa7b-4745-bc9c-bbb8d444ac47 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Teaching VLMs to Localize Specific Objects from In-context Examples LoRA: Low-Rank Adaptation of Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T16:39:40.532371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:39:40.532371Z digest=sha256:3d6f34f9490ab44a534f17420beaa303abe2a3f19c2302b2029fd39fb4b04270

Observation da2197ba-17e3-46ba-8e97-32736727f1de · outbound

This paper cites Multimodal Task Vectors Enable Many-Shot Multimodal In-Context Learning.

Teaching VLMs to Localize Specific Objects from In-context Examples Multimodal Task Vectors Enable Many-Shot Multimodal In-Context Learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T16:39:40.536776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:39:40.536776Z digest=sha256:4cc4fd79efb072d7d0f80270b9301ef356dd8e993c0575c9a1035fede4399102

Observation 8a132cbb-8e01-4271-975d-a27d5e071e5f · outbound

This paper cites ConMe: Rethinking Evaluation of Compositional Reasoning for Modern VLMs.

Teaching VLMs to Localize Specific Objects from In-context Examples ConMe: Rethinking Evaluation of Compositional Reasoning for Modern VLMs

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T16:39:40.541079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:39:40.541079Z digest=sha256:03119f6c81ad23e12b1b622842914f840e5653b4032d57482c12cf34c6b21103

Observation c545965a-a60e-44d3-bf91-d8ae4a21144c · outbound

This paper cites Got-10k: A large high-diversity benchmark for generic object tracking in the wild.

Teaching VLMs to Localize Specific Objects from In-context Examples Got-10k: A large high-diversity benchmark for generic object tracking in the wild

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:39:41.412751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:39:40.545070Z digest=sha256:c1f0e3fac5ae36748cbb189cf735ddf2b2cc5e6ae962bdef003f71915b20c906

Observation 33af747b-5247-4e84-8b00-3d5771dd7a2a · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

Teaching VLMs to Localize Specific Objects from In-context Examples Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:39:41.400012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:39:40.549059Z digest=sha256:d267c99626f4799f43c116e4ba6e983b191e7636ecce5ee4caf55970be02cf54

Observation 94fe1d9f-b150-4776-acd0-b0a29c10cbf6 · outbound

This paper cites Le, Yunhsuan Sung, Zhen Li, and Tom Duerig.

Teaching VLMs to Localize Specific Objects from In-context Examples Le, Yunhsuan Sung, Zhen Li, and Tom Duerig

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:39:41.387969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:39:40.552750Z digest=sha256:644991cbaa61a069c641e49e5b44463301d37036f0fa9fc502a2c5c75b82c157

Observation 35fc6464-8f6e-4c8d-bec8-7506fc282719 · outbound

This paper cites Improving Zero-Shot Models with Label Distribution Priors.

Teaching VLMs to Localize Specific Objects from In-context Examples Improving Zero-Shot Models with Label Distribution Priors

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T16:39:40.557149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:39:40.557149Z digest=sha256:612bdf3248bf8837b23f492896a5dacaef3532854639354e95a00c2f57c3a31c

Observation 6fdff393-4014-45ea-b112-dbb80debdda9 · outbound

This paper cites Building and better understanding vision- language models: insights and future directions., 2024.

Teaching VLMs to Localize Specific Objects from In-context Examples Building and better understanding vision- language models: insights and future directions., 2024

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:39:41.376507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:39:40.560467Z digest=sha256:fc1508e9cd8e08f0f59681aa2f5159c96009843966d1309852d5f4423b4fa164

Observation 01145f1d-b456-4daa-b31b-da539b857769 · outbound

This paper cites Rush, Douwe Kiela, Matthieu Cord, and Victor Sanh.

Teaching VLMs to Localize Specific Objects from In-context Examples Rush, Douwe Kiela, Matthieu Cord, and Victor Sanh

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T16:39:40.563667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:39:40.563667Z digest=sha256:367b9ce8c4b310afc2c980bd99e71f3964a721e16d2607171789e49bdf2a8d7f

Observation aa5e7dce-cfde-4538-917b-f031368508ec · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Teaching VLMs to Localize Specific Objects from In-context Examples LLaVA-OneVision: Easy Visual Task Transfer

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T16:39:40.567068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:39:40.567068Z digest=sha256:87998a1761e7d1e1e2e84b56999aacdbad7ea24c06f1db8c0317c9ada43f6649

Observation c84b9ba8-0073-4f37-be49-5db383960f33 · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

Teaching VLMs to Localize Specific Objects from In-context Examples BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:39:41.357258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:39:40.570471Z digest=sha256:83e2159b70d32e0c35bc06deb6234da963bc39148b498ec16842d200fcf42f10

Observation b0813bb4-f2e8-4f14-8a16-a738e03b66a6 · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

Teaching VLMs to Localize Specific Objects from In-context Examples Evaluating Object Hallucination in Large Vision-Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T16:39:40.573595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:39:40.573595Z digest=sha256:f96e5cc217e1835268ceb6531d3489b750ca52c7b059a9e2df619270f3752358

Observation c91f8eed-6ed1-4d0c-a355-a4dbce10b525 · outbound

This paper cites Video-LLaV A: Learning united visual repre- sentation by alignment before projection.

Teaching VLMs to Localize Specific Objects from In-context Examples Video-LLaV A: Learning united visual repre- sentation by alignment before projection

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:39:41.345079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:39:40.577439Z digest=sha256:c7d5453072a3465fe2e3b1b19b25884913469daf5c506d0c1f118e78dec6b5fb

Observation 519c5fd0-9609-44ef-a658-9fd720182095 · outbound

This paper cites Microsoft coco: Common objects in context.

Teaching VLMs to Localize Specific Objects from In-context Examples Microsoft coco: Common objects in context

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:39:41.333635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:39:40.581006Z digest=sha256:96e8f724b48b8b77038fdb141185b736dab83d5eaf2da38d4c71a8f3b709814c

Observation 4c5fb39b-bd6b-4146-86d0-5ce1a19cf6a5 · outbound

This paper cites MAtch, eXpand and Improve: Unsupervised Finetuning for Zero-Shot Action Recognition with Language Knowledge.

Teaching VLMs to Localize Specific Objects from In-context Examples MAtch, eXpand and Improve: Unsupervised Finetuning for Zero-Shot Action Recognition with Language Knowledge

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:39:41.322994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:39:40.584841Z digest=sha256:e642d18e40d36fc105d1dc700c7e06cac772f7cfa3e5f2235a016266e3726ddc

Observation fe242f4f-7192-4569-b606-d8581bcd90fe · outbound

This paper cites LLaV A-NeXT: Improved reasoning, OCR, and world knowl- edge, 2023.

Teaching VLMs to Localize Specific Objects from In-context Examples LLaV A-NeXT: Improved reasoning, OCR, and world knowl- edge, 2023

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:39:41.311111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:39:40.588422Z digest=sha256:11c3f2f1b16f4035598b82588e53b9aaa2d62313be15d5c373f2b36603c6707a

Observation 6dd04bcd-bac5-47bd-a1e7-44bb346c8192 · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

Teaching VLMs to Localize Specific Objects from In-context Examples Improved Baselines with Visual Instruction Tuning

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:39:41.299613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:39:40.592000Z digest=sha256:cc14a6f8021e8dadcdf108730deb47e33424fe4708dc64c7016c43446de20c9e

Observation baff4608-4631-4f08-af17-81cd5f94941c · outbound

This paper cites Visual Instruction Tuning.

Teaching VLMs to Localize Specific Objects from In-context Examples Visual Instruction Tuning

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:39:41.288260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:39:40.595416Z digest=sha256:ea3471a5d3111ce4c97c2f9c1f2cc314afea80dc89128a287d70fb8e605315ee

Observation 0c31c0e0-c363-424b-95bd-4fe8a5fa9407 · outbound

This paper cites MetaICL: Learning to Learn In Context.

Teaching VLMs to Localize Specific Objects from In-context Examples MetaICL: Learning to Learn In Context

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:39:41.276171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:39:40.598963Z digest=sha256:01ca166f6b03f744d7668af068f59eb7aa189c7550b50b79aa8548ac0304e77c

Observation 9fa3ac82-a1f5-4457-8a1c-5a3dbc7a2aa3 · outbound

This paper cites Simple open-vocabulary object detection.

Teaching VLMs to Localize Specific Objects from In-context Examples Simple open-vocabulary object detection

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:39:41.264856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:39:40.602374Z digest=sha256:bd9bcb0e2b9fc6e4f009459cf3bca6ed7803ca068c666a2a005bcfc74ef12d20

Observation c8fb27d6-9170-4d02-886a-c7274d494a5b · outbound

This paper cites Jehanzeb Mirza, Leonid Karlinsky, Wei Lin, Sivan Doveh, , Jakub Micorek, Mateusz Kozinski, Hilde Kuhene, and Horst Possegger.

Teaching VLMs to Localize Specific Objects from In-context Examples Jehanzeb Mirza, Leonid Karlinsky, Wei Lin, Sivan Doveh, , Jakub Micorek, Mateusz Kozinski, Hilde Kuhene, and Horst Possegger

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:39:41.251810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:39:40.605839Z digest=sha256:b3d257f22c6161f3008a7db4cc51fc9742e3fd8618a0f5807f42ded865ca8b57

Observation 5164a560-0e1b-4fc1-95e4-5b9e5ae55f2b · outbound

This paper cites TAP: Targeted Prompting for Task Adaptive Generation of Textual Training Instances for Visual Classification.

Teaching VLMs to Localize Specific Objects from In-context Examples TAP: Targeted Prompting for Task Adaptive Generation of Textual Training Instances for Visual Classification

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T16:39:40.609623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:39:40.609623Z digest=sha256:6c775cea2deb2e97bbc633673995c8a49c3c4c43f76dfe3bf020b696ec52a794

Observation 175230ee-ede7-4953-8990-398a4235376f · outbound

This paper cites LaFTer: Label-Free Tuning of Zero-shot Classifier using Language and Unlabeled Image Collections.

Teaching VLMs to Localize Specific Objects from In-context Examples LaFTer: Label-Free Tuning of Zero-shot Classifier using Language and Unlabeled Image Collections

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:39:41.238798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:39:40.614016Z digest=sha256:6736358f1853bda01d1f06dacf7a87dbab58b7a507f4a1373e617a436c8d5299

Observation a11f86cc-1d84-4f14-86dc-26392ea575c6 · outbound

This paper cites GLOV: Guided Large Language Models as Implicit Optimizers for Vision Language Models.

Teaching VLMs to Localize Specific Objects from In-context Examples GLOV: Guided Large Language Models as Implicit Optimizers for Vision Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T16:39:40.617774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:39:40.617774Z digest=sha256:27727feff5e0cdf9411ad05f4b3875e5b07b4b9794cebd7c0827a5697db02b31

Observation e8e46f47-cb9b-4ddd-84da-400aa6b54e99 · outbound

This paper cites GPT-4 Technical Report.

Teaching VLMs to Localize Specific Objects from In-context Examples GPT-4 Technical Report

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T16:39:40.621780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:39:40.621780Z digest=sha256:a75b65e63028152e185086a8c5f392d1cf5d959874b3db9e5c0d0f586298c6af

Observation f3e373b9-d24c-4297-aa9b-3283af0c9793 · outbound

This paper cites Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation.

Teaching VLMs to Localize Specific Objects from In-context Examples Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T16:39:40.625552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:39:40.625552Z digest=sha256:b8651e9dd91c0be0f821facdcdf6eae390917fffd512ed8c7167ad4cad3132b1

Observation 160657ba-b93e-49f5-a829-2bcf62741f84 · outbound

This paper cites Learning Transferable Visual Models from Natural Language Supervision.

Teaching VLMs to Localize Specific Objects from In-context Examples Learning Transferable Visual Models from Natural Language Supervision

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:39:41.227513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:39:40.629796Z digest=sha256:4e846800325327ae5a6852604d16b6104fd297cd5e4fa7fd0333ccbac639e0de

Observation 3b28e437-91ac-4fe9-a4a3-f28f1477b76d · outbound

This paper cites Where’s waldo: Diffusion features for person- alized segmentation and retrieval.

Teaching VLMs to Localize Specific Objects from In-context Examples Where’s waldo: Diffusion features for person- alized segmentation and retrieval

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:39:41.214913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:39:40.633609Z digest=sha256:4b3b995ae8db9fb712f4ce15700dda0ef58c5cde2b3a9200a1a1c93d018a3f3e

Observation 0b6132a6-ce8a-4202-aa88-0afec2a19689 · outbound

This paper cites LAION-5b: An open large-scale dataset for training next generation image-text models.

Teaching VLMs to Localize Specific Objects from In-context Examples LAION-5b: An open large-scale dataset for training next generation image-text models

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:39:41.202871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:39:40.637531Z digest=sha256:0cdf50d7d96dd4ae0f5cbe5d52b1dcbed8572e4c874917af6fc478fc1b736222

Observation ba8e2d6a-1097-4033-9fbb-5e5a2cc5a152 · outbound

This paper cites Generative Multimodal Models are In-Context Learners.

Teaching VLMs to Localize Specific Objects from In-context Examples Generative Multimodal Models are In-Context Learners

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T16:39:40.641230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:39:40.641230Z digest=sha256:f6b6af9e5a7a810026d1b5ff09f7fbe8ac556044f4b8bfb46e922f99825cdad0

Observation 3924a722-b553-4b89-8913-59da0fdae312 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Teaching VLMs to Localize Specific Objects from In-context Examples Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T16:39:40.645655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:39:40.645655Z digest=sha256:feaef3cb44733e2900dd54e7ae1f98dea14bc2e97692b0a671b88032cfb7fe9a

Observation ef08ee48-1c8a-4169-b68c-9fe9d13dac31 · outbound

This paper cites Frustratingly Simple Few-Shot Object Detection.

Teaching VLMs to Localize Specific Objects from In-context Examples Frustratingly Simple Few-Shot Object Detection

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T16:39:40.649624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:39:40.649624Z digest=sha256:cdab3cc4af2093bc08a9e00918e770b8f699c7e26a0d405cd9fcab4324b36edf

Observation 9fa68b21-9f5e-464f-99dc-765642284e80 · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

Teaching VLMs to Localize Specific Objects from In-context Examples Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:39:41.190180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:39:40.653529Z digest=sha256:a3884c27b74d6f1afaf3c782b48c879f19d83eb715e42d161aea06d305c99307

Observation 205c90c8-e051-41f4-845c-0666c1fb871c · outbound

This paper cites Larger language models do in-context learning differently.

Teaching VLMs to Localize Specific Objects from In-context Examples Larger language models do in-context learning differently

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T16:39:40.656379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:39:40.656379Z digest=sha256:417cd4be3e6ff84127a285dbd9243c79720a554a481b7da427cd5ce5142873df

Observation 77021848-34fa-4134-9011-7cbe8b15e93a · outbound

This paper cites Demystify- ing CLIP Data.

Teaching VLMs to Localize Specific Objects from In-context Examples Demystify- ing CLIP Data

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:39:41.175239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:39:40.659764Z digest=sha256:6e70946307a42d8a2f3afc684b5e6deca3a790524b32519fbc238a08fa27bcd6

Observation d4c46cb3-2366-49d2-8401-fc9c3fc6659f · outbound

This paper cites Tree of Thoughts: Deliberate Problem Solving with Large Language Models.

Teaching VLMs to Localize Specific Objects from In-context Examples Tree of Thoughts: Deliberate Problem Solving with Large Language Models

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:39:41.160513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:39:40.662991Z digest=sha256:d6ff57603452165ad13b0d0e872688f91137f45f7509628a3d0dcbcbbbb76f66

Observation 9f94954e-ca22-44fb-8a66-c9b02056ee0e · outbound

This paper cites Sigmoid Loss for Language Image Pre- training.

Teaching VLMs to Localize Specific Objects from In-context Examples Sigmoid Loss for Language Image Pre- training

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:39:41.148368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:39:40.666254Z digest=sha256:58478b35348616df92867b100589a3f49ccacf9c5feb4e9aa7823bad8cd6ec9b

Observation d274c548-9ea8-40aa-9335-2e17f8c9b0ab · outbound

This paper cites Personalize Segment Anything Model with One Shot.

Teaching VLMs to Localize Specific Objects from In-context Examples Personalize Segment Anything Model with One Shot

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-12T16:39:40.670404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:39:40.670404Z digest=sha256:3188fd71b385cc1d1afc0d83c1809512e1b3762b6548f1fdc0a887b6cc405c3b

Observation 938945a8-1bf8-4757-9a8a-2697a5c556ff · outbound

This paper cites MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning.

Teaching VLMs to Localize Specific Objects from In-context Examples MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-12T16:39:40.674841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:39:40.674841Z digest=sha256:bb6a85d62387e766e57ac1dec72ce4d5226b8dbf81297a0e6255d149d5f1392f

Observation c2ab052d-2b49-4d7b-9950-fa3c2055ab67 · outbound

This paper cites Llamafac- tory: Unified efficient fine-tuning of 100+ language mod- els.

Teaching VLMs to Localize Specific Objects from In-context Examples Llamafac- tory: Unified efficient fine-tuning of 100+ language mod- els

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:39:41.134100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:39:40.678763Z digest=sha256:f432d292963423518797e60a6084788a1c4e859164ac9846c6ef4a9fc613a231

Observation 94ae1924-e7c2-43f3-830a-1e3029e61081 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Teaching VLMs to Localize Specific Objects from In-context Examples MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:39:41.122145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:39:40.682746Z digest=sha256:ecac088f18aef5451adfff7fa8b4785a9e19c8ac89ee17400913c411d8cf03de

Observation 2778b1e1-baad-4af4-83b2-95b3c377f8ec · outbound

This paper cites <ref>category</ref>.

Teaching VLMs to Localize Specific Objects from In-context Examples <ref>category</ref>

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:39:41.109720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:39:40.687439Z digest=sha256:758b13f6d19b102371b5fcd803916d1cfafe84679e0c86aeb62a5be9b495379b

Pith citing papers

No inbound Pith citation observations are available.