Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T16:39:40.687439Z
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 0 inbound Pith citation observations for arXiv:2411.13317.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T16:39:40.687439Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
61 of 61 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 37ddf28d-cbe8-43cd-9fdf-5bf0c0da1bee · outbound
Teaching VLMs to Localize Specific Objects from In-context Examples Pixtral 12B
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a83b50c-0057-4a65-98c3-638a881f30a4 · outbound
Teaching VLMs to Localize Specific Objects from In-context Examples Flamingo: a Visual Language Model for Few-Shot Learning
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation acb2a615-02de-4be6-8190-6425fbe40cbd · outbound
Teaching VLMs to Localize Specific Objects from In-context Examples Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a938626-bc81-4701-9e58-4cf7c375e58d · outbound
Teaching VLMs to Localize Specific Objects from In-context Examples DeciMamba: Exploring the Length Extrapolation Potential of Mamba
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c5f6a9b-5fae-489e-a060-19bbac3f903d · outbound
Teaching VLMs to Localize Specific Objects from In-context Examples Lan- guage models are few-shot learners
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ce28c8d-a6d9-4b19-b1b8-e17f57bd4801 · outbound
Teaching VLMs to Localize Specific Objects from In-context Examples Unresolved cited work
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation a547f13d-6ce2-4733-aece-8b0ed6f60aee · outbound
Teaching VLMs to Localize Specific Objects from In-context Examples MiniGPT-v2: Large Language Model as a Unified Interface for Vision-Language Multi-task Learning
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 4bb4b678-9d7f-4dc3-965d-839f697798bd · outbound
Teaching VLMs to Localize Specific Objects from In-context Examples How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 660e31e1-6833-46ff-aaaf-41de340fa4aa · outbound
Teaching VLMs to Localize Specific Objects from In-context Examples Gonzalez, Ion Stoica, and Eric P
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 56991c9d-8caa-4af0-b942-95d8e8e38f2f · outbound
Teaching VLMs to Localize Specific Objects from In-context Examples InstructBLIP: Towards General-purpose Vision- Language Models with Instruction Tuning
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation fefc5034-5bc1-4149-a8c1-52a3968fae0d · outbound
Teaching VLMs to Localize Specific Objects from In-context Examples Tao: A large-scale bench- mark for tracking any object
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 72c2c290-9dbf-4842-897e-d2185b5ee624 · outbound
Teaching VLMs to Localize Specific Objects from In-context Examples Dense and Aligned Captions (DAC) Promote Compositional Reasoning in VL Models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation aa7b3741-3448-489a-819e-b80b7ec9db98 · outbound
Teaching VLMs to Localize Specific Objects from In-context Examples Teaching structured vision & language concepts to vision & language models
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation f9cbc167-0a17-4e7e-8298-5527f44f95f3 · outbound
Teaching VLMs to Localize Specific Objects from In-context Examples Towards Multimodal In-Context Learning for Vision & Language Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b7704a9-01ed-4ebe-adf9-a5a4c901f8a0 · outbound
Teaching VLMs to Localize Specific Objects from In-context Examples The Llama 3 Herd of Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c326ce3-f619-4baa-80a9-b62541567166 · outbound
Teaching VLMs to Localize Specific Objects from In-context Examples Lasot: A high-quality benchmark for large-scale single ob- ject tracking
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03581bc5-4c3b-42cb-819c-1eeece0d12ca · outbound
Teaching VLMs to Localize Specific Objects from In-context Examples SEED: Self-supervised Dis- 9 tillation for Visual Representation
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 3124bf49-f3ba-4c40-aa7a-62ec7484cc3f · outbound
Teaching VLMs to Localize Specific Objects from In-context Examples Cross-domain few-shot object detection via enhanced open-set object detector
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation d571833c-e27a-4c74-b083-dcc1f19b7e47 · outbound
Teaching VLMs to Localize Specific Objects from In-context Examples Vision-Language Models Create Cross-Modal Task Representations
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation f1cbd9f0-fa7b-4745-bc9c-bbb8d444ac47 · outbound
Teaching VLMs to Localize Specific Objects from In-context Examples LoRA: Low-Rank Adaptation of Large Language Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da2197ba-17e3-46ba-8e97-32736727f1de · outbound
Teaching VLMs to Localize Specific Objects from In-context Examples Multimodal Task Vectors Enable Many-Shot Multimodal In-Context Learning
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a132cbb-8e01-4271-975d-a27d5e071e5f · outbound
Teaching VLMs to Localize Specific Objects from In-context Examples ConMe: Rethinking Evaluation of Compositional Reasoning for Modern VLMs
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c545965a-a60e-44d3-bf91-d8ae4a21144c · outbound
Teaching VLMs to Localize Specific Objects from In-context Examples Got-10k: A large high-diversity benchmark for generic object tracking in the wild
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 33af747b-5247-4e84-8b00-3d5771dd7a2a · outbound
Teaching VLMs to Localize Specific Objects from In-context Examples Gqa: A new dataset for real-world visual reasoning and compositional question answering
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 94fe1d9f-b150-4776-acd0-b0a29c10cbf6 · outbound
Teaching VLMs to Localize Specific Objects from In-context Examples Le, Yunhsuan Sung, Zhen Li, and Tom Duerig
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 35fc6464-8f6e-4c8d-bec8-7506fc282719 · outbound
Teaching VLMs to Localize Specific Objects from In-context Examples Improving Zero-Shot Models with Label Distribution Priors
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6fdff393-4014-45ea-b112-dbb80debdda9 · outbound
Teaching VLMs to Localize Specific Objects from In-context Examples Building and better understanding vision- language models: insights and future directions., 2024
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 01145f1d-b456-4daa-b31b-da539b857769 · outbound
Teaching VLMs to Localize Specific Objects from In-context Examples Rush, Douwe Kiela, Matthieu Cord, and Victor Sanh
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa5e7dce-cfde-4538-917b-f031368508ec · outbound
Teaching VLMs to Localize Specific Objects from In-context Examples LLaVA-OneVision: Easy Visual Task Transfer
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c84b9ba8-0073-4f37-be49-5db383960f33 · outbound
Teaching VLMs to Localize Specific Objects from In-context Examples BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation b0813bb4-f2e8-4f14-8a16-a738e03b66a6 · outbound
Teaching VLMs to Localize Specific Objects from In-context Examples Evaluating Object Hallucination in Large Vision-Language Models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c91f8eed-6ed1-4d0c-a355-a4dbce10b525 · outbound
Teaching VLMs to Localize Specific Objects from In-context Examples Video-LLaV A: Learning united visual repre- sentation by alignment before projection
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 519c5fd0-9609-44ef-a658-9fd720182095 · outbound
Teaching VLMs to Localize Specific Objects from In-context Examples Microsoft coco: Common objects in context
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 4c5fb39b-bd6b-4146-86d0-5ce1a19cf6a5 · outbound
Teaching VLMs to Localize Specific Objects from In-context Examples MAtch, eXpand and Improve: Unsupervised Finetuning for Zero-Shot Action Recognition with Language Knowledge
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation fe242f4f-7192-4569-b606-d8581bcd90fe · outbound
Teaching VLMs to Localize Specific Objects from In-context Examples LLaV A-NeXT: Improved reasoning, OCR, and world knowl- edge, 2023
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 6dd04bcd-bac5-47bd-a1e7-44bb346c8192 · outbound
Teaching VLMs to Localize Specific Objects from In-context Examples Improved Baselines with Visual Instruction Tuning
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation baff4608-4631-4f08-af17-81cd5f94941c · outbound
Teaching VLMs to Localize Specific Objects from In-context Examples Visual Instruction Tuning
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 0c31c0e0-c363-424b-95bd-4fe8a5fa9407 · outbound
Teaching VLMs to Localize Specific Objects from In-context Examples MetaICL: Learning to Learn In Context
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 9fa3ac82-a1f5-4457-8a1c-5a3dbc7a2aa3 · outbound
Teaching VLMs to Localize Specific Objects from In-context Examples Simple open-vocabulary object detection
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation c8fb27d6-9170-4d02-886a-c7274d494a5b · outbound
Teaching VLMs to Localize Specific Objects from In-context Examples Jehanzeb Mirza, Leonid Karlinsky, Wei Lin, Sivan Doveh, , Jakub Micorek, Mateusz Kozinski, Hilde Kuhene, and Horst Possegger
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 5164a560-0e1b-4fc1-95e4-5b9e5ae55f2b · outbound
Teaching VLMs to Localize Specific Objects from In-context Examples TAP: Targeted Prompting for Task Adaptive Generation of Textual Training Instances for Visual Classification
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 175230ee-ede7-4953-8990-398a4235376f · outbound
Teaching VLMs to Localize Specific Objects from In-context Examples LaFTer: Label-Free Tuning of Zero-shot Classifier using Language and Unlabeled Image Collections
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation a11f86cc-1d84-4f14-86dc-26392ea575c6 · outbound
Teaching VLMs to Localize Specific Objects from In-context Examples GLOV: Guided Large Language Models as Implicit Optimizers for Vision Language Models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8e46f47-cb9b-4ddd-84da-400aa6b54e99 · outbound
Teaching VLMs to Localize Specific Objects from In-context Examples GPT-4 Technical Report
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3e373b9-d24c-4297-aa9b-3283af0c9793 · outbound
Teaching VLMs to Localize Specific Objects from In-context Examples Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 160657ba-b93e-49f5-a829-2bcf62741f84 · outbound
Teaching VLMs to Localize Specific Objects from In-context Examples Learning Transferable Visual Models from Natural Language Supervision
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 3b28e437-91ac-4fe9-a4a3-f28f1477b76d · outbound
Teaching VLMs to Localize Specific Objects from In-context Examples Where’s waldo: Diffusion features for person- alized segmentation and retrieval
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 0b6132a6-ce8a-4202-aa88-0afec2a19689 · outbound
Teaching VLMs to Localize Specific Objects from In-context Examples LAION-5b: An open large-scale dataset for training next generation image-text models
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation ba8e2d6a-1097-4033-9fbb-5e5a2cc5a152 · outbound
Teaching VLMs to Localize Specific Objects from In-context Examples Generative Multimodal Models are In-Context Learners
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3924a722-b553-4b89-8913-59da0fdae312 · outbound
Teaching VLMs to Localize Specific Objects from In-context Examples Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef08ee48-1c8a-4169-b68c-9fe9d13dac31 · outbound
Teaching VLMs to Localize Specific Objects from In-context Examples Frustratingly Simple Few-Shot Object Detection
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9fa68b21-9f5e-464f-99dc-765642284e80 · outbound
Teaching VLMs to Localize Specific Objects from In-context Examples Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 205c90c8-e051-41f4-845c-0666c1fb871c · outbound
Teaching VLMs to Localize Specific Objects from In-context Examples Larger language models do in-context learning differently
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77021848-34fa-4134-9011-7cbe8b15e93a · outbound
Teaching VLMs to Localize Specific Objects from In-context Examples Demystify- ing CLIP Data
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation d4c46cb3-2366-49d2-8401-fc9c3fc6659f · outbound
Teaching VLMs to Localize Specific Objects from In-context Examples Tree of Thoughts: Deliberate Problem Solving with Large Language Models
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 9f94954e-ca22-44fb-8a66-c9b02056ee0e · outbound
Teaching VLMs to Localize Specific Objects from In-context Examples Sigmoid Loss for Language Image Pre- training
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation d274c548-9ea8-40aa-9335-2e17f8c9b0ab · outbound
Teaching VLMs to Localize Specific Objects from In-context Examples Personalize Segment Anything Model with One Shot
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 938945a8-1bf8-4757-9a8a-2697a5c556ff · outbound
Teaching VLMs to Localize Specific Objects from In-context Examples MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2ab052d-2b49-4d7b-9950-fa3c2055ab67 · outbound
Teaching VLMs to Localize Specific Objects from In-context Examples Llamafac- tory: Unified efficient fine-tuning of 100+ language mod- els
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 94ae1924-e7c2-43f3-830a-1e3029e61081 · outbound
Teaching VLMs to Localize Specific Objects from In-context Examples MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 2778b1e1-baad-4af4-83b2-95b3c377f8ec · outbound
Teaching VLMs to Localize Specific Objects from In-context Examples <ref>category</ref>
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
No inbound Pith citation observations are available.