Pith. sign in

Paper Citation Record · LEDGER

Teaching VLMs to Localize Specific Objects from In-context Examples

As of 14 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 0 inbound Pith citation observations for arXiv:2411.13317.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.13317 v2

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T16:39:40.687439Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

61 of 61 outbound references displayed

  • verified exact1
  • verified fuzzy34
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 37ddf28d-cbe8-43cd-9fdf-5bf0c0da1bee · outbound

This paper cites Pixtral 12B.

Teaching VLMs to Localize Specific Objects from In-context Examples Pixtral 12B

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T16:39:40.456544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:39:40.456544Z digest=sha256:75250a7b5903fbfc0e1c34f90d3efc0a2960f9ec19d9abb64aaf9268b8be78ca

Observation 6a83b50c-0057-4a65-98c3-638a881f30a4 · outbound

This paper cites Flamingo: a Visual Language Model for Few-Shot Learning.

Teaching VLMs to Localize Specific Objects from In-context Examples Flamingo: a Visual Language Model for Few-Shot Learning

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:39:41.549331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T16:39:40.461608Z digest=sha256:71c032955c4551ec057c1d167439e096d73fe937faeb4bd231c72a5d1e4cbb62

Observation acb2a615-02de-4be6-8190-6425fbe40cbd · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Teaching VLMs to Localize Specific Objects from In-context Examples Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T16:39:40.465499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:39:40.465499Z digest=sha256:338bf90c0f1dcab6a1cae7176016a9cde8c24c4df0c68f20155e56caa4e82b4a

Observation 7a938626-bc81-4701-9e58-4cf7c375e58d · outbound

This paper cites DeciMamba: Exploring the Length Extrapolation Potential of Mamba.

Teaching VLMs to Localize Specific Objects from In-context Examples DeciMamba: Exploring the Length Extrapolation Potential of Mamba

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T16:39:40.469752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:39:40.469752Z digest=sha256:7ed24759d98d17201b56b4ff2618d8b438ef48dc9b4265d55659578192d0b358

Observation 3c5f6a9b-5fae-489e-a060-19bbac3f903d · outbound

This paper cites Lan- guage models are few-shot learners.

Teaching VLMs to Localize Specific Objects from In-context Examples Lan- guage models are few-shot learners

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T16:39:40.474895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:39:40.474895Z digest=sha256:5dc925f48dab7939cce569e601af3829fa719658b5a6a8611c82a29e0834cb4a

Observation 1ce28c8d-a6d9-4b19-b1b8-e17f57bd4801 · outbound

This paper cites an unresolved cited work.

Teaching VLMs to Localize Specific Objects from In-context Examples Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-12T16:39:41.529173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T16:39:40.478781Z digest=sha256:31ba74c6eddb7352a15a1698a69a9cc3854b9261e5796f67d2849fd3e68d2b59

Observation a547f13d-6ce2-4733-aece-8b0ed6f60aee · outbound

This paper cites MiniGPT-v2: Large Language Model as a Unified Interface for Vision-Language Multi-task Learning.

Teaching VLMs to Localize Specific Objects from In-context Examples MiniGPT-v2: Large Language Model as a Unified Interface for Vision-Language Multi-task Learning

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:39:41.517511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T16:39:40.482901Z digest=sha256:f3a7db539692e6688f62f460614d7bf37a901367c4796dc4e281ece0b75915be

Observation 4bb4b678-9d7f-4dc3-965d-839f697798bd · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

Teaching VLMs to Localize Specific Objects from In-context Examples How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T16:39:40.487264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:39:40.487264Z digest=sha256:c4c50c17530b8671b385f9543c731a01402100c71ad99c763ea7dbdb2827c0c0

Observation 660e31e1-6833-46ff-aaaf-41de340fa4aa · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

Teaching VLMs to Localize Specific Objects from In-context Examples Gonzalez, Ion Stoica, and Eric P

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:39:41.506480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T16:39:40.491205Z digest=sha256:e50c69b8cb13da624d8c8cbe168f7514e2fe59aead52593b00a72dbbcf9d87d0

Observation 56991c9d-8caa-4af0-b942-95d8e8e38f2f · outbound

This paper cites InstructBLIP: Towards General-purpose Vision- Language Models with Instruction Tuning.

Teaching VLMs to Localize Specific Objects from In-context Examples InstructBLIP: Towards General-purpose Vision- Language Models with Instruction Tuning

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:39:41.494396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T16:39:40.494839Z digest=sha256:513adc5b9fb7680c8a57a153343ff76965e34a6f417d41da5f794b597fb2b3dc

Observation fefc5034-5bc1-4149-a8c1-52a3968fae0d · outbound

This paper cites Tao: A large-scale bench- mark for tracking any object.

Teaching VLMs to Localize Specific Objects from In-context Examples Tao: A large-scale bench- mark for tracking any object

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:39:41.482721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T16:39:40.498165Z digest=sha256:32ba31c6adf43c4285e1671f353cf88c359be6ba6ec42cbc475cdb8fc0aa41d8

Observation 72c2c290-9dbf-4842-897e-d2185b5ee624 · outbound

This paper cites Dense and Aligned Captions (DAC) Promote Compositional Reasoning in VL Models.

Teaching VLMs to Localize Specific Objects from In-context Examples Dense and Aligned Captions (DAC) Promote Compositional Reasoning in VL Models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:39:41.469865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T16:39:40.501315Z digest=sha256:58b3f381310e2f8c382acee50936bc8613f037913e95eb3d9771c2b25c0eff7a

Observation aa7b3741-3448-489a-819e-b80b7ec9db98 · outbound

This paper cites Teaching structured vision & language concepts to vision & language models.

Teaching VLMs to Localize Specific Objects from In-context Examples Teaching structured vision & language concepts to vision & language models

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:39:41.456962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T16:39:40.505137Z digest=sha256:5f4a36b5486513cb2211ded287993ce30e58e60e3bbb990b5f1d6562934c9180

Observation f9cbc167-0a17-4e7e-8298-5527f44f95f3 · outbound

This paper cites Towards Multimodal In-Context Learning for Vision & Language Models.

Teaching VLMs to Localize Specific Objects from In-context Examples Towards Multimodal In-Context Learning for Vision & Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T16:39:40.508301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:39:40.508301Z digest=sha256:cf2c516c184bb18aa24007c96319ec1141cd27cf011d0485cc1fa842b4afc9ce

Observation 8b7704a9-01ed-4ebe-adf9-a5a4c901f8a0 · outbound

This paper cites The Llama 3 Herd of Models.

Teaching VLMs to Localize Specific Objects from In-context Examples The Llama 3 Herd of Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T16:39:40.512501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:39:40.512501Z digest=sha256:9cd3daeb439d3dedbdff6374340e0238f8965ad9524d2df9f2cc4f96ccbe964a

Observation 8c326ce3-f619-4baa-80a9-b62541567166 · outbound

This paper cites Lasot: A high-quality benchmark for large-scale single ob- ject tracking.

Teaching VLMs to Localize Specific Objects from In-context Examples Lasot: A high-quality benchmark for large-scale single ob- ject tracking

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T16:39:40.516794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:39:40.516794Z digest=sha256:4ecb2730016e00006ba3a9f6e435f0c7a4d2d0878647e6dc4a4216212f65d254

Observation 03581bc5-4c3b-42cb-819c-1eeece0d12ca · outbound

This paper cites SEED: Self-supervised Dis- 9 tillation for Visual Representation.

Teaching VLMs to Localize Specific Objects from In-context Examples SEED: Self-supervised Dis- 9 tillation for Visual Representation

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:39:41.438365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T16:39:40.520549Z digest=sha256:bc3ab7b7d3112eab0281fbcaa72706de6daeadd7e924a4a8687d57d64f037a5f

Observation 3124bf49-f3ba-4c40-aa7a-62ec7484cc3f · outbound

This paper cites Cross-domain few-shot object detection via enhanced open-set object detector.

Teaching VLMs to Localize Specific Objects from In-context Examples Cross-domain few-shot object detection via enhanced open-set object detector

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:39:41.424973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T16:39:40.524498Z digest=sha256:26166cf59d084b85c22b2c7422e871339db7d4f0aaad36e22824fba36a1c5383

Observation d571833c-e27a-4c74-b083-dcc1f19b7e47 · outbound

This paper cites Vision-Language Models Create Cross-Modal Task Representations.

Teaching VLMs to Localize Specific Objects from In-context Examples Vision-Language Models Create Cross-Modal Task Representations

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-12T16:39:41.022727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T16:39:40.528370Z digest=sha256:146f8e8d7d5774ff247f430791355cf737ee983eba31515351c141cfeee6c451

Observation f1cbd9f0-fa7b-4745-bc9c-bbb8d444ac47 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Teaching VLMs to Localize Specific Objects from In-context Examples LoRA: Low-Rank Adaptation of Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T16:39:40.532371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:39:40.532371Z digest=sha256:285cf2de63f3926a8c9835ecca51b44c4922be0d3e4c63af782224ca86b05f16

Observation da2197ba-17e3-46ba-8e97-32736727f1de · outbound

This paper cites Multimodal Task Vectors Enable Many-Shot Multimodal In-Context Learning.

Teaching VLMs to Localize Specific Objects from In-context Examples Multimodal Task Vectors Enable Many-Shot Multimodal In-Context Learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T16:39:40.536776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:39:40.536776Z digest=sha256:df335f4e68b15649f331a2548ae1e7f2981ad43945dc3d97ee1cb34f15afe0e1

Observation 8a132cbb-8e01-4271-975d-a27d5e071e5f · outbound

This paper cites ConMe: Rethinking Evaluation of Compositional Reasoning for Modern VLMs.

Teaching VLMs to Localize Specific Objects from In-context Examples ConMe: Rethinking Evaluation of Compositional Reasoning for Modern VLMs

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T16:39:40.541079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:39:40.541079Z digest=sha256:b3baf9ff0f0129e78fdba7251fdbeface7198327a7c51c10e4527e0e472e72cd

Observation c545965a-a60e-44d3-bf91-d8ae4a21144c · outbound

This paper cites Got-10k: A large high-diversity benchmark for generic object tracking in the wild.

Teaching VLMs to Localize Specific Objects from In-context Examples Got-10k: A large high-diversity benchmark for generic object tracking in the wild

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:39:41.412751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T16:39:40.545070Z digest=sha256:96a49231c5e36bbd35ca7e2688f7326decef5ee970bdb5fcbf6a06cde61a336c

Observation 33af747b-5247-4e84-8b00-3d5771dd7a2a · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

Teaching VLMs to Localize Specific Objects from In-context Examples Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:39:41.400012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T16:39:40.549059Z digest=sha256:1898562ec836993c4c14b11004752ded135a94edc30990174011850187c003cd

Observation 94fe1d9f-b150-4776-acd0-b0a29c10cbf6 · outbound

This paper cites Le, Yunhsuan Sung, Zhen Li, and Tom Duerig.

Teaching VLMs to Localize Specific Objects from In-context Examples Le, Yunhsuan Sung, Zhen Li, and Tom Duerig

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:39:41.387969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T16:39:40.552750Z digest=sha256:c4ecc5953aea850042060b6492caae3f04ba3715e21c0ac15f9ee18ac04144ad

Observation 35fc6464-8f6e-4c8d-bec8-7506fc282719 · outbound

This paper cites Improving Zero-Shot Models with Label Distribution Priors.

Teaching VLMs to Localize Specific Objects from In-context Examples Improving Zero-Shot Models with Label Distribution Priors

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T16:39:40.557149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:39:40.557149Z digest=sha256:6acd296d0b991c6f09f5aa17ccd7c45e7c06fa6cd14edf333e4a49f2dcc4949e

Observation 6fdff393-4014-45ea-b112-dbb80debdda9 · outbound

This paper cites Building and better understanding vision- language models: insights and future directions., 2024.

Teaching VLMs to Localize Specific Objects from In-context Examples Building and better understanding vision- language models: insights and future directions., 2024

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:39:41.376507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T16:39:40.560467Z digest=sha256:fc002b9d0e7d8d7dea46c7d1fd482001c11dfb9188924b4d2b5b741481023726

Observation 01145f1d-b456-4daa-b31b-da539b857769 · outbound

This paper cites Rush, Douwe Kiela, Matthieu Cord, and Victor Sanh.

Teaching VLMs to Localize Specific Objects from In-context Examples Rush, Douwe Kiela, Matthieu Cord, and Victor Sanh

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T16:39:40.563667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:39:40.563667Z digest=sha256:1806f5ac0fc759f82cd40958931396a6730a98809425cfffbce291be88434ce1

Observation aa5e7dce-cfde-4538-917b-f031368508ec · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Teaching VLMs to Localize Specific Objects from In-context Examples LLaVA-OneVision: Easy Visual Task Transfer

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T16:39:40.567068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:39:40.567068Z digest=sha256:0578295bba0bc2adb4d042731eafc737fbc9cd65fa736d7105a50d03d4d9ba57

Observation c84b9ba8-0073-4f37-be49-5db383960f33 · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

Teaching VLMs to Localize Specific Objects from In-context Examples BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:39:41.357258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T16:39:40.570471Z digest=sha256:348156987e0b0e056cdd8f244daa0c8afb1366fdf70aa77fd75674b09cda9880

Observation b0813bb4-f2e8-4f14-8a16-a738e03b66a6 · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

Teaching VLMs to Localize Specific Objects from In-context Examples Evaluating Object Hallucination in Large Vision-Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T16:39:40.573595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:39:40.573595Z digest=sha256:9262d4e2fdc0522f7b875ada39921b58cff9e33077232cd2915e9dac207a7878

Observation c91f8eed-6ed1-4d0c-a355-a4dbce10b525 · outbound

This paper cites Video-LLaV A: Learning united visual repre- sentation by alignment before projection.

Teaching VLMs to Localize Specific Objects from In-context Examples Video-LLaV A: Learning united visual repre- sentation by alignment before projection

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:39:41.345079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T16:39:40.577439Z digest=sha256:d8d8b5569fdd6bc36cb70c1674999984371b3ccf932474d5c4f27df258154bd8

Observation 519c5fd0-9609-44ef-a658-9fd720182095 · outbound

This paper cites Microsoft coco: Common objects in context.

Teaching VLMs to Localize Specific Objects from In-context Examples Microsoft coco: Common objects in context

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:39:41.333635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T16:39:40.581006Z digest=sha256:f99bfd0f8ccf1685cda86efed4d22f06c635dd5c86443ddeb6a93feaad6d87b7

Observation 4c5fb39b-bd6b-4146-86d0-5ce1a19cf6a5 · outbound

This paper cites MAtch, eXpand and Improve: Unsupervised Finetuning for Zero-Shot Action Recognition with Language Knowledge.

Teaching VLMs to Localize Specific Objects from In-context Examples MAtch, eXpand and Improve: Unsupervised Finetuning for Zero-Shot Action Recognition with Language Knowledge

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:39:41.322994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T16:39:40.584841Z digest=sha256:b4fa3d92d53faef5b0c31b031d1e80e3426b5898c69d6689c48b2e27fd2e7465

Observation fe242f4f-7192-4569-b606-d8581bcd90fe · outbound

This paper cites LLaV A-NeXT: Improved reasoning, OCR, and world knowl- edge, 2023.

Teaching VLMs to Localize Specific Objects from In-context Examples LLaV A-NeXT: Improved reasoning, OCR, and world knowl- edge, 2023

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:39:41.311111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T16:39:40.588422Z digest=sha256:0e9f7554010508a125a8a8d031957c3de9d1b176b9f011098366f8d385ef499e

Observation 6dd04bcd-bac5-47bd-a1e7-44bb346c8192 · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

Teaching VLMs to Localize Specific Objects from In-context Examples Improved Baselines with Visual Instruction Tuning

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:39:41.299613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T16:39:40.592000Z digest=sha256:070a6e5bf1d59976ba93da38679bd5af4686aac7a96e0e5f6252512db79f776f

Observation baff4608-4631-4f08-af17-81cd5f94941c · outbound

This paper cites Visual Instruction Tuning.

Teaching VLMs to Localize Specific Objects from In-context Examples Visual Instruction Tuning

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:39:41.288260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T16:39:40.595416Z digest=sha256:afbce9679b3c51e239479f792a4b58e6393b510fba88289615a374a6d9108ad6

Observation 0c31c0e0-c363-424b-95bd-4fe8a5fa9407 · outbound

This paper cites MetaICL: Learning to Learn In Context.

Teaching VLMs to Localize Specific Objects from In-context Examples MetaICL: Learning to Learn In Context

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:39:41.276171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T16:39:40.598963Z digest=sha256:5f4fa71512b3f8259595ec77543542becdb077bbbb7f7cae5b25119f49d46928

Observation 9fa3ac82-a1f5-4457-8a1c-5a3dbc7a2aa3 · outbound

This paper cites Simple open-vocabulary object detection.

Teaching VLMs to Localize Specific Objects from In-context Examples Simple open-vocabulary object detection

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:39:41.264856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T16:39:40.602374Z digest=sha256:13c730f88177b7c92620e1df667a92a21b9d1ee417a85a36611ebbeae701d75e

Observation c8fb27d6-9170-4d02-886a-c7274d494a5b · outbound

This paper cites Jehanzeb Mirza, Leonid Karlinsky, Wei Lin, Sivan Doveh, , Jakub Micorek, Mateusz Kozinski, Hilde Kuhene, and Horst Possegger.

Teaching VLMs to Localize Specific Objects from In-context Examples Jehanzeb Mirza, Leonid Karlinsky, Wei Lin, Sivan Doveh, , Jakub Micorek, Mateusz Kozinski, Hilde Kuhene, and Horst Possegger

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:39:41.251810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T16:39:40.605839Z digest=sha256:474accc249830cb0545e7596a7f56a98bc7757b28c11d8b147655e124c77f53e

Observation 5164a560-0e1b-4fc1-95e4-5b9e5ae55f2b · outbound

This paper cites TAP: Targeted Prompting for Task Adaptive Generation of Textual Training Instances for Visual Classification.

Teaching VLMs to Localize Specific Objects from In-context Examples TAP: Targeted Prompting for Task Adaptive Generation of Textual Training Instances for Visual Classification

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T16:39:40.609623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:39:40.609623Z digest=sha256:b11d51096014d4429330af359212ffc71c5be81420db2b46a39772c206d5b941

Observation 175230ee-ede7-4953-8990-398a4235376f · outbound

This paper cites LaFTer: Label-Free Tuning of Zero-shot Classifier using Language and Unlabeled Image Collections.

Teaching VLMs to Localize Specific Objects from In-context Examples LaFTer: Label-Free Tuning of Zero-shot Classifier using Language and Unlabeled Image Collections

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:39:41.238798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T16:39:40.614016Z digest=sha256:d2cf22e5373e51ed3226ca95978284037a59a495255f1f9772386aeb0219eb6a

Observation a11f86cc-1d84-4f14-86dc-26392ea575c6 · outbound

This paper cites GLOV: Guided Large Language Models as Implicit Optimizers for Vision Language Models.

Teaching VLMs to Localize Specific Objects from In-context Examples GLOV: Guided Large Language Models as Implicit Optimizers for Vision Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T16:39:40.617774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:39:40.617774Z digest=sha256:cb51fb5625bb47b29107410f0e0fbc9320c5936d9f3aab7d2838a3318d88b6de

Observation e8e46f47-cb9b-4ddd-84da-400aa6b54e99 · outbound

This paper cites GPT-4 Technical Report.

Teaching VLMs to Localize Specific Objects from In-context Examples GPT-4 Technical Report

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T16:39:40.621780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:39:40.621780Z digest=sha256:d7a084f5779528010353eb44e0c536cbf99e1826761d4b584083bc94a7c199a2

Observation f3e373b9-d24c-4297-aa9b-3283af0c9793 · outbound

This paper cites Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation.

Teaching VLMs to Localize Specific Objects from In-context Examples Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T16:39:40.625552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:39:40.625552Z digest=sha256:8a96a537eeacf04d4691dc7d22fae0144edcea34a227e8f7e2c87797963182a4

Observation 160657ba-b93e-49f5-a829-2bcf62741f84 · outbound

This paper cites Learning Transferable Visual Models from Natural Language Supervision.

Teaching VLMs to Localize Specific Objects from In-context Examples Learning Transferable Visual Models from Natural Language Supervision

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:39:41.227513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T16:39:40.629796Z digest=sha256:543d86c7cd6570d0122a97e927e566252ec7f602f5509992bbb05918baa5a227

Observation 3b28e437-91ac-4fe9-a4a3-f28f1477b76d · outbound

This paper cites Where’s waldo: Diffusion features for person- alized segmentation and retrieval.

Teaching VLMs to Localize Specific Objects from In-context Examples Where’s waldo: Diffusion features for person- alized segmentation and retrieval

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:39:41.214913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T16:39:40.633609Z digest=sha256:161ec3e6ce642210694a370c1b8f7e6413fe4825db9025b42a7405da9c03441e

Observation 0b6132a6-ce8a-4202-aa88-0afec2a19689 · outbound

This paper cites LAION-5b: An open large-scale dataset for training next generation image-text models.

Teaching VLMs to Localize Specific Objects from In-context Examples LAION-5b: An open large-scale dataset for training next generation image-text models

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:39:41.202871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T16:39:40.637531Z digest=sha256:33f95c46a62a325306b70b3a5e0c8db1ac5c64ad782a2148c167ef41abdff999

Observation ba8e2d6a-1097-4033-9fbb-5e5a2cc5a152 · outbound

This paper cites Generative Multimodal Models are In-Context Learners.

Teaching VLMs to Localize Specific Objects from In-context Examples Generative Multimodal Models are In-Context Learners

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T16:39:40.641230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:39:40.641230Z digest=sha256:84d952d4b153eb070ee10ef4281c2e57b54edd999376d84f76c425d8990dad56

Observation 3924a722-b553-4b89-8913-59da0fdae312 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Teaching VLMs to Localize Specific Objects from In-context Examples Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T16:39:40.645655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:39:40.645655Z digest=sha256:c0ebf58fe60961b6a0c9664ca7069e7a8b5e1a548391e977e791ff12c48eaa0d

Observation ef08ee48-1c8a-4169-b68c-9fe9d13dac31 · outbound

This paper cites Frustratingly Simple Few-Shot Object Detection.

Teaching VLMs to Localize Specific Objects from In-context Examples Frustratingly Simple Few-Shot Object Detection

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T16:39:40.649624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:39:40.649624Z digest=sha256:31e418101b437b50526a214e1d57f13bbd25984bd9413dc09b8ceeb9ac2f699d

Observation 9fa68b21-9f5e-464f-99dc-765642284e80 · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

Teaching VLMs to Localize Specific Objects from In-context Examples Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:39:41.190180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T16:39:40.653529Z digest=sha256:ee03355714c71d3437f575a75613416961ab7597ab2de46624b9d9b13070c274

Observation 205c90c8-e051-41f4-845c-0666c1fb871c · outbound

This paper cites Larger language models do in-context learning differently.

Teaching VLMs to Localize Specific Objects from In-context Examples Larger language models do in-context learning differently

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T16:39:40.656379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:39:40.656379Z digest=sha256:0bcc9bf7d0a518e41625f516de6ee64f1b977919ead475e138e225d6c11dfb13

Observation 77021848-34fa-4134-9011-7cbe8b15e93a · outbound

This paper cites Demystify- ing CLIP Data.

Teaching VLMs to Localize Specific Objects from In-context Examples Demystify- ing CLIP Data

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:39:41.175239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T16:39:40.659764Z digest=sha256:163538a41b9313d9c55385e198ae77966f0c6e6ee789f215277866304dbb1fde

Observation d4c46cb3-2366-49d2-8401-fc9c3fc6659f · outbound

This paper cites Tree of Thoughts: Deliberate Problem Solving with Large Language Models.

Teaching VLMs to Localize Specific Objects from In-context Examples Tree of Thoughts: Deliberate Problem Solving with Large Language Models

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:39:41.160513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T16:39:40.662991Z digest=sha256:4cf43f6a2ccfd9f8653dde46f9bd2f6631a427376e9ba3bc6bc15ebd98f74b96

Observation 9f94954e-ca22-44fb-8a66-c9b02056ee0e · outbound

This paper cites Sigmoid Loss for Language Image Pre- training.

Teaching VLMs to Localize Specific Objects from In-context Examples Sigmoid Loss for Language Image Pre- training

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:39:41.148368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T16:39:40.666254Z digest=sha256:dda28594d1bc90fc7b4112688b1092598b104edf06735e52113cf4240e188ad1

Observation d274c548-9ea8-40aa-9335-2e17f8c9b0ab · outbound

This paper cites Personalize Segment Anything Model with One Shot.

Teaching VLMs to Localize Specific Objects from In-context Examples Personalize Segment Anything Model with One Shot

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-12T16:39:40.670404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:39:40.670404Z digest=sha256:2ff3a3b30f6fd833769a598a72d1d3f697ba0bb2b1eb1946e20953b67cbf60f2

Observation 938945a8-1bf8-4757-9a8a-2697a5c556ff · outbound

This paper cites MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning.

Teaching VLMs to Localize Specific Objects from In-context Examples MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-12T16:39:40.674841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:39:40.674841Z digest=sha256:fc6351cf46e8dd393e02609d999a85babc02d88e69bf4d49a01fab82f5eed01e

Observation c2ab052d-2b49-4d7b-9950-fa3c2055ab67 · outbound

This paper cites Llamafac- tory: Unified efficient fine-tuning of 100+ language mod- els.

Teaching VLMs to Localize Specific Objects from In-context Examples Llamafac- tory: Unified efficient fine-tuning of 100+ language mod- els

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:39:41.134100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T16:39:40.678763Z digest=sha256:efb54bdd33b1d8fef2e63ed086dde3ef3b96db39190009f6e0e0d73c09ed2129

Observation 94ae1924-e7c2-43f3-830a-1e3029e61081 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Teaching VLMs to Localize Specific Objects from In-context Examples MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:39:41.122145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T16:39:40.682746Z digest=sha256:13e668a8d49aa3d7e0809f02403945bf5b748585583348d8e6c46db8011041ac

Observation 2778b1e1-baad-4af4-83b2-95b3c377f8ec · outbound

This paper cites <ref>category</ref>.

Teaching VLMs to Localize Specific Objects from In-context Examples <ref>category</ref>

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:39:41.109720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T16:39:40.687439Z digest=sha256:adcfc0affb39c9467a7220e23569a9f365048c7275ba67fb169c3c7e30f5d7cb

Pith citing papers

No inbound Pith citation observations are available.