Pith. sign in

Paper Citation Record · LEDGER

Flamingo: a Visual Language Model for Few-Shot Learning

As of 12 August 2026, this Paper Citation Record lists 100 of 165 outbound references and 100 inbound Pith citation observations for arXiv:2204.14198.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2204.14198 v2

Coverage vector

measured 100 of 165 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-12T04:22:30.008355Z

measured 200 of 200 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 100 of 164 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T21:10:16.529113Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

100 of 165 outbound references displayed

  • verified exact37
  • verified fuzzy51
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch6

External citation measurements

1244
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 70778e2e-984c-40f1-9fb0-0c0b8d3eead3 · outbound

This paper cites CM3: A Causal Masked Multimodal Model of the Internet.

Flamingo: a Visual Language Model for Few-Shot Learning CM3: A Causal Masked Multimodal Model of the Internet

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T04:22:30.133071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:9bea2f36c50529d3bcb58c720b25ef723d08f0dac6a002ab1aef620e6e6959cc

Observation 9ac766f7-5b89-4aac-98e4-0329f5641652 · outbound

This paper cites Self- supervised multimodal versatile networks.

Flamingo: a Visual Language Model for Few-Shot Learning Self- supervised multimodal versatile networks

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:22:30.747983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:36e22e097ae15d6e442df185bc3a1f2419c7df9b93ce219212a9b49ed1d3a45f

Observation a1062d83-a7ab-42f2-99b0-704822c4e37a · outbound

This paper cites Lawrence Zitnick, and Devi Parikh.

Flamingo: a Visual Language Model for Few-Shot Learning Lawrence Zitnick, and Devi Parikh

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:22:30.751062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:4eca1b341fb7b320c0e2605ada2323b5e042e60960588c5b67f0cc81214c85dc

Observation 27d0936e-919a-43a9-99ee-0847ab0ec610 · outbound

This paper cites ReZero is all you need: Fast convergence at large depth.

Flamingo: a Visual Language Model for Few-Shot Learning ReZero is all you need: Fast convergence at large depth

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:22:30.756766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:52fb05e005730253c5fce8e70ddc54976bc6eb229e6b34a6261e8fe4b9302bea

Observation a9be39e4-523d-40d9-bb69-bbedfaaa0cc5 · outbound

This paper cites Frozen in time: A joint video and image encoder for end-to-end retrieval.

Flamingo: a Visual Language Model for Few-Shot Learning Frozen in time: A joint video and image encoder for end-to-end retrieval

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:22:30.767228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:b713ec2870c0c803ce0feab08a668ec8d3f50302db0b760fc4937ed5d662f5c0

Observation 01ab33df-4bcc-422b-9a4e-f1ec70c1d8ee · outbound

This paper cites Henriques, Jack Valmadre, Philip Torr, and Andrea Vedaldi.

Flamingo: a Visual Language Model for Few-Shot Learning Henriques, Jack Valmadre, Philip Torr, and Andrea Vedaldi

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:22:30.771825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:0dd2ad0bcf17e66e186913ea75680c9007479c125674fdc7c4327bf8c6f2954c

Observation fbf68780-77f1-4c04-85e5-bb99bb549a9f · outbound

This paper cites Meta-learning with differentiable closed-form solvers.

Flamingo: a Visual Language Model for Few-Shot Learning Meta-learning with differentiable closed-form solvers

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:22:30.625153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:6fb6816d2e90ed3a6068db4606c43ccafde96ac41769f7fff67c4451cac3f109

Observation 0b7577f0-4794-47cf-b8bd-8a66eedd4310 · outbound

This paper cites JAX: composable transformations of Python+NumPy programs.

Flamingo: a Visual Language Model for Few-Shot Learning JAX: composable transformations of Python+NumPy programs

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:22:30.780109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:ac21fbf2ae27b4043dd98a5d9e136b1fa5eb07d176196520c8876a54dcdf553c

Observation 857b7d0f-da3b-48ed-9540-b3e5d005af2a · outbound

This paper cites an unresolved cited work.

Flamingo: a Visual Language Model for Few-Shot Learning Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-05-12T04:22:30.789363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:5c3cf95bcce9da35e88411fd683350c385bcf3cb6c5ab6f2607ad34272b58972

Observation 359f8b36-d44c-44f7-a474-e14b426c4e43 · outbound

This paper cites High-Performance Large-Scale Image Recognition Without Normalization.

Flamingo: a Visual Language Model for Few-Shot Learning High-Performance Large-Scale Image Recognition Without Normalization

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:22:30.232180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:e8c2d676d920b6462e471ba4096af846d420d74c534e2fec37f4a23c4018ba92

Observation 3cd128ba-bb5b-49f0-89e0-a8782ef9051b · outbound

This paper cites an unresolved cited work.

Flamingo: a Visual Language Model for Few-Shot Learning Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-05-12T04:22:30.796728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:5f0417a194f22d3364853182e8fc787d232dec82bd1815f89568ff556a3d7fba

Observation 63f56a12-30c0-4d20-801a-dd6142fcc0a3 · outbound

This paper cites Gender shades: Intersectional accuracy disparities in commercial gender classification.

Flamingo: a Visual Language Model for Few-Shot Learning Gender shades: Intersectional accuracy disparities in commercial gender classification

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:22:30.800217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:80e6afe41a014a1e1dd0744f7a5573c64709b3a9ecdadd849b397c93aca0505d

Observation d7e16df7-f972-4827-9757-a25082f6efe1 · outbound

This paper cites End-to-end object detection with transformers.

Flamingo: a Visual Language Model for Few-Shot Learning End-to-end object detection with transformers

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:22:30.803262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:517579aba0ac035e3c0ec5e289904b5599d018f0f17f6ca7321151656f221885

Observation 112ee62b-5ad6-4a1a-9f39-ff3f640f861c · outbound

This paper cites Conceptual 12M: Pushing web-scale image-text pre-training to recognize long-tail visual concepts.

Flamingo: a Visual Language Model for Few-Shot Learning Conceptual 12M: Pushing web-scale image-text pre-training to recognize long-tail visual concepts

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:22:30.806769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:be459f5c7acf432da3023c618496adb2389576b8907527efb09c155b95339aa3

Observation 9bd9d5a2-0df3-4e36-a224-5533aea59dae · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

Flamingo: a Visual Language Model for Few-Shot Learning Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-12T21:38:19.742717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:1d2d8044e4af0e6c2315465d81f53982d423898b6e5add340baa9172ec2c509f

Observation e083d66e-0e2c-4ef5-add9-10b682474612 · outbound

This paper cites UNITER: Universal image-text representation learning.

Flamingo: a Visual Language Model for Few-Shot Learning UNITER: Universal image-text representation learning

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:22:30.810219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:0a56143b0ba84c5c9bd52b0818c1c55d39a11bd7adfb1a7fa8102ccec409a579

Observation 77ed2ac9-32b5-4171-8cee-b9354753f54f · outbound

This paper cites Unifying vision-and-language tasks via text generation.

Flamingo: a Visual Language Model for Few-Shot Learning Unifying vision-and-language tasks via text generation

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:22:30.814300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:63ec1b8607ed4ce6b9d11382bddd9381076c7c4221d4f717d9783c1c67591eba

Observation 5563f1ea-d7c6-4120-80aa-d8bddbddfb64 · outbound

This paper cites PaLM: Scaling Language Modeling with Pathways.

Flamingo: a Visual Language Model for Few-Shot Learning PaLM: Scaling Language Modeling with Pathways

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-12T04:22:30.327417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:9b88cb8d0e39e84901c872c8440c443ca0a697d285426af7876d8b88ffcd39f9

Observation b077227d-0ea7-4002-8dfb-fe2d8f57ffeb · outbound

This paper cites Enabling multimodal generation on clip via vision-language knowledge distillation.

Flamingo: a Visual Language Model for Few-Shot Learning Enabling multimodal generation on clip via vision-language knowledge distillation

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:22:30.818007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:11c4fa452d6603fb17f0ba2b4d0d29c305c3675c6f57d87308cba5ce1333c68c

Observation 4222e384-6851-4cd5-b95e-e13e674f6417 · outbound

This paper cites Visual dialog.

Flamingo: a Visual Language Model for Few-Shot Learning Visual dialog

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:22:30.824546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:e934b68f666b3651a178129b3c3d07ace74ac730230721e93219a4937a0962d8

Observation f1c9fa92-9171-4be0-b8ff-6f710ae209ce · outbound

This paper cites Does object recognition work for everyone? In IEEE Computer Vision and Pattern Recognition.

Flamingo: a Visual Language Model for Few-Shot Learning Does object recognition work for everyone? In IEEE Computer Vision and Pattern Recognition

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:22:30.828270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:3534c154356608853f380fae6926738b416e182f1881e8e7b451b4c5b4eea115

Observation 74c85cd4-8dff-45e0-a182-869584018a19 · outbound

This paper cites VirTex: Learning visual representations from textual annota- tions.

Flamingo: a Visual Language Model for Few-Shot Learning VirTex: Learning visual representations from textual annota- tions

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:22:30.832184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:098cb17bd184fc6464b5d2e899d509d6e206241c0e86fa5f042f98be3f9ffba5

Observation fd349d53-267c-4e14-9045-d052755e41c5 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Flamingo: a Visual Language Model for Few-Shot Learning BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-12T04:22:30.332740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:10a218cfa37219ebde4a174e5842609f270ae1ced603b22c66434d446ffc9ff6

Observation 7785852e-86c0-414a-baba-a8bbc4d678ae · outbound

This paper cites CrossTransformers: spatially-aware few-shot transfer.

Flamingo: a Visual Language Model for Few-Shot Learning CrossTransformers: spatially-aware few-shot transfer

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:22:30.835592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:97c632509584a1cd7471be2b9ba4bf6e4d843e872a987fb502ed615d6a867b3e

Observation 1273261c-bf8e-494d-9beb-cca545281b18 · outbound

This paper cites Long-term recurrent convolutional networks for visual recognition and description.

Flamingo: a Visual Language Model for Few-Shot Learning Long-term recurrent convolutional networks for visual recognition and description

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:22:30.843334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:a175c2a67b1511f55fe31de878515a958262bf05e039d97abda02a08b14a080d

Observation 6fcb7573-739f-457c-a8ce-a01256b35669 · outbound

This paper cites MAGMA -- Multimodal Augmentation of Generative Models through Adapter-based Finetuning.

Flamingo: a Visual Language Model for Few-Shot Learning MAGMA -- Multimodal Augmentation of Generative Models through Adapter-based Finetuning

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T04:22:30.180343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:cf541aeee2ce6b6795479c512aeb6525d24631122eaa8358cce6c3c0967112d7

Observation bceb97fa-76c7-446d-a422-a20fee6bb5ab · outbound

This paper cites Model-agnostic meta-learning for fast adaptation of deep networks.

Flamingo: a Visual Language Model for Few-Shot Learning Model-agnostic meta-learning for fast adaptation of deep networks

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:22:30.847039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:87d03195040f87deada5a57f15fa0fc969a2e2cd3b6ec58648b152c4ca90a7bb

Observation 066ab485-dc68-4c03-afbd-710371346d83 · outbound

This paper cites VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling.

Flamingo: a Visual Language Model for Few-Shot Learning VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:22:30.252979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:72406d412b80e396fc8638b67018ec9988b151c0adc66a7f2d06433d968ea477

Observation b5338e59-d9e3-45f1-9981-1169719ec244 · outbound

This paper cites Large-scale adversarial training for vision-and-language representation learning.

Flamingo: a Visual Language Model for Few-Shot Learning Large-scale adversarial training for vision-and-language representation learning

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:22:30.850639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:3578264feebde937b595207b0ea57a2ab7b7be09cca2f33154967c9247180772

Observation 6ab027de-991b-40ce-aea1-38a2bf372462 · outbound

This paper cites Datasheets for datasets.

Flamingo: a Visual Language Model for Few-Shot Learning Datasheets for datasets

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:22:30.853816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:6bbbbeaec5319b5f62f5c9509138387987a0151990f6715cd5e16e08977cce4e

Observation 5d9a49a5-c8a8-4c4a-8696-698f504d40c2 · outbound

This paper cites Meta-Learning Probabilistic Inference For Prediction.

Flamingo: a Visual Language Model for Few-Shot Learning Meta-Learning Probabilistic Inference For Prediction

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:22:30.415415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:62e4a8a606c891d66e7a510984399d98563747d9d2201dda826236409fafe153

Observation eae5d548-68e5-4cf8-b6cd-f45cb4bd9811 · outbound

This paper cites Generating Sequences With Recurrent Neural Networks.

Flamingo: a Visual Language Model for Few-Shot Learning Generating Sequences With Recurrent Neural Networks

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:22:30.482043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:fdd2dbc440a45e1c95782ac243eb1160edbb572e0f9221631c1af12db6c62aa9

Observation 690aab85-9388-4c11-8da9-777539c1b3da · outbound

This paper cites Griffiths, Frederick Callaway, Michael B.

Flamingo: a Visual Language Model for Few-Shot Learning Griffiths, Frederick Callaway, Michael B

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:22:30.862362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:ad2b214ee4879a0a095d7089e02cf6957040b973c486849cfa617a7f9847e118

Observation 314f70ec-9202-4977-92e7-e119a1d396f1 · outbound

This paper cites KAT: A Knowledge Augmented Transformer for Vision-and-Language.

Flamingo: a Visual Language Model for Few-Shot Learning KAT: A Knowledge Augmented Transformer for Vision-and-Language

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:22:30.515015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:61f1fa054524f8bdae92b5ab5b383dcd85a24505c7f08cbcc77653f77b254665

Observation fae816f9-52a2-4629-9202-5eda08c74b5f · outbound

This paper cites Stangl, Anhong Guo, Chi Lin, Kristen Grauman, Jiebo Luo, and Jeffrey P.

Flamingo: a Visual Language Model for Few-Shot Learning Stangl, Anhong Guo, Chi Lin, Kristen Grauman, Jiebo Luo, and Jeffrey P

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:22:30.867310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:c2c0fd6e79982d9f6412a480d3967c202bfdb9e16044bb3b453b2b6ac343db6b

Observation 8501bb37-4668-4bfc-9689-f0dd56366d5b · outbound

This paper cites Transformer Language Models without Positional Encodings Still Learn Positional Information.

Flamingo: a Visual Language Model for Few-Shot Learning Transformer Language Models without Positional Encodings Still Learn Positional Information

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T04:22:30.088404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:a2139bc395e15f3e29af780dcbc86162e1afefdf9e5d85bf0c7d22e6c1ef92a7

Observation a7263e53-1ae3-4877-84d4-63666af3292a · outbound

This paper cites Women also snowboard: Overcoming bias in captioning models.

Flamingo: a Visual Language Model for Few-Shot Learning Women also snowboard: Overcoming bias in captioning models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:22:30.872613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:c8ec027c1a4b870fad57d4f1168fef4474683b71b0a3e0f2fe402a0420ee5ae9

Observation 67a5309d-fe07-44c3-9b37-41ccd31b64fa · outbound

This paper cites Decoupling the role of data, attention, and losses in multimodal transformers.

Flamingo: a Visual Language Model for Few-Shot Learning Decoupling the role of data, attention, and losses in multimodal transformers

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:22:30.875607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:5782a08d57a94020b3da800fdc5f878c134917cc9cd3c479b9c49cb1f27845d9

Observation fbd2c84e-63ab-4759-8fd6-76ea3162830b · outbound

This paper cites Gaussian Error Linear Units (GELUs).

Flamingo: a Visual Language Model for Few-Shot Learning Gaussian Error Linear Units (GELUs)

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-12T04:22:30.238555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:3018c933924b1b45daed7e3f83d02a5cfcdf6b32644909359e56ff75990659fe

Observation 3721b229-6351-485c-a92f-d414caebec41 · outbound

This paper cites Haiku: Sonnet for JAX.

Flamingo: a Visual Language Model for Few-Shot Learning Haiku: Sonnet for JAX

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:22:30.879584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:94e4a5f1173141af4835df9b9471f81de7110d2d607b78aabee6fd72e10707f6

Observation 501f32be-e62d-4a57-b9a2-9cbf96dfbad5 · outbound

This paper cites an unresolved cited work.

Flamingo: a Visual Language Model for Few-Shot Learning Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-05-12T04:22:30.883350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:a6b1c4e93cf37368fe32d178a9c12dcb6293fdc3c8a7dfdf4497cd8e75d0feaf

Observation 9bac2006-7fc3-47a5-a7ed-45a48ac6e450 · outbound

This paper cites Long short-term memory.

Flamingo: a Visual Language Model for Few-Shot Learning Long short-term memory

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:22:30.886734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:471d3e26a25466d4377a5cbbb7bf350cf15f10ba1f74e11bfd3cf116c30d55a4

Observation dd258193-3a6c-41fc-8511-69f5ebe27dcd · outbound

This paper cites Training Compute-Optimal Large Language Models.

Flamingo: a Visual Language Model for Few-Shot Learning Training Compute-Optimal Large Language Models

Reference 43

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T04:22:30.370781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:72f36c1e5a8f5d0866ab743bdec4ff024e9605c35964187dd83d04da27539c1a

Observation 5a13d5fa-36b8-4d4a-9369-b63b5d9ea6ea · outbound

This paper cites Parameter-efficient transfer learning for NLP.

Flamingo: a Visual Language Model for Few-Shot Learning Parameter-efficient transfer learning for NLP

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:22:30.901234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:37555d86ebcd33114cab1956d0d5b9c55c045cd722165988f064fcfd1ab8f5dc

Observation 10cc3d1a-9ec3-4ba2-a696-5f5925b2e09d · outbound

This paper cites Universal Language Model Fine-tuning for Text Classification.

Flamingo: a Visual Language Model for Few-Shot Learning Universal Language Model Fine-tuning for Text Classification

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:22:30.420070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:c94842a20338003651bf8f98db46dd83b832685bee244dc5309d53401b7256d6

Observation e504dc3c-b67d-4922-bfd6-2dd887c5494d · outbound

This paper cites Scaling Up Vision-Language Pre-training for Image Captioning.

Flamingo: a Visual Language Model for Few-Shot Learning Scaling Up Vision-Language Pre-training for Image Captioning

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:22:30.459551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:a941df948fa21663801b57264d6acd65d03b453bc83b1bea4f20657301236950

Observation 978b0d15-dd70-4d47-b7e1-43f22b0eccf8 · outbound

This paper cites Attention on attention for image captioning.

Flamingo: a Visual Language Model for Few-Shot Learning Attention on attention for image captioning

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:22:30.925912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:6f3b9ceb82b0f9955bdf8212428cf930e135d09088dc75dfd8c6b88e156e4dfc

Observation 2c9c9db0-94bf-4e77-beca-211c415071a5 · outbound

This paper cites Derpanis, and Neil D.

Flamingo: a Visual Language Model for Few-Shot Learning Derpanis, and Neil D

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:22:30.929401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:f6fcca0c2632779909b3e41056e12d1ee8c1d560051dc7ab6f2cfb0c35d6d7c9

Observation 242cc7b1-b572-4c62-9d7e-11d638d04637 · outbound

This paper cites Perceiver: General perception with iterative attention.

Flamingo: a Visual Language Model for Few-Shot Learning Perceiver: General perception with iterative attention

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:22:30.935245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:f4f35a8351d7a0fb136031d607c28054978ac0fa3fe62f5884585e5f49e58ab2

Observation c22c2205-aa49-4158-8b6c-c9e0d37f0f07 · outbound

This paper cites MURAL: Multimodal, Multitask Retrieval Across Languages.

Flamingo: a Visual Language Model for Few-Shot Learning MURAL: Multimodal, Multitask Retrieval Across Languages

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:22:30.573408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:1225f1c05cd436f4bd8995a0af648ad1dd0e6509c471a0ac5f6ffe4f2bff91f0

Observation 7fd6cab6-82f1-4d96-a6d7-6e0c3bc84eb8 · outbound

This paper cites Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision.

Flamingo: a Visual Language Model for Few-Shot Learning Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:22:30.583825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:2afc0b55c11a2d2b0018037ad34738ba0d130e53b512e0579d84e56b3715376f

Observation 15f088e5-4bdb-4807-953c-29f47edcd0ee · outbound

This paper cites All in One: Exploring Unified Video-Language Pre-training.

Flamingo: a Visual Language Model for Few-Shot Learning All in One: Exploring Unified Video-Language Pre-training

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:22:30.588905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:a9e3aaf7bd23c9216b7ef5363b4fdaf9b32f9a47c28003970e2e55f63080b8dc

Observation 7018ff85-dcd2-47dd-95cb-ac00a47c8fe2 · outbound

This paper cites Exploring the Limits of Language Modeling.

Flamingo: a Visual Language Model for Few-Shot Learning Exploring the Limits of Language Modeling

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:22:30.593919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:1800fe7ab68e4c94567cb1f9152451eaba10de1cc1c562b6d46eebe7b75c0004

Observation bd706f14-42d4-4648-98d0-6da74048fa20 · outbound

This paper cites Scaling Laws for Neural Language Models.

Flamingo: a Visual Language Model for Few-Shot Learning Scaling Laws for Neural Language Models

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-05-12T04:22:30.610263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:3e6c404a6b00a7fad16b9f95508c910e4fbcd6b6e899fa9de88d6d2a69cfd597

Observation 9566ef18-588f-43db-bc06-5bbfe4536167 · outbound

This paper cites The Hateful Memes Challenge: Detecting hate speech in multimodal memes.

Flamingo: a Visual Language Model for Few-Shot Learning The Hateful Memes Challenge: Detecting hate speech in multimodal memes

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:22:30.940129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:f127fe3517ca6990f2074a75811ce2dca5be8ad0d785810ff20ca3998b480b20

Observation 0340b57d-6f89-40ff-b901-ae02952053f5 · outbound

This paper cites Few-shot classification by recycling deep learning.

Flamingo: a Visual Language Model for Few-Shot Learning Few-shot classification by recycling deep learning

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:22:30.080238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:f575fbd5b403821d1ce2e9d3be5ab026356beb5cb1e1e3e83d2730baddc45ce0

Observation f0139bf0-04ac-4adf-b773-e8d84969fd7f · outbound

This paper cites Align before fuse: Vision and language representation learning with momentum distillation.

Flamingo: a Visual Language Model for Few-Shot Learning Align before fuse: Vision and language representation learning with momentum distillation

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:22:30.946268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:c7256e0c4f4cf74afda3c69f5844a710fb6ed394676b45f640c28ccb5347454e

Observation baa19c17-51eb-4465-be33-c1137286e3f4 · outbound

This paper cites BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation.

Flamingo: a Visual Language Model for Few-Shot Learning BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:22:30.106705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:cb331fc1674634d43deec58a3e15071552e73803dd8ea7a35dd4f5274b8d976a

Observation f4092bb5-b83a-4515-add3-d7bca8a114e5 · outbound

This paper cites HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-training.

Flamingo: a Visual Language Model for Few-Shot Learning HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-training

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:22:30.162304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:788183ad82613576dc1887b2976ba2f7b4569f12531de5c2263093f66667bbe5

Observation 780ab3b3-c84f-43c0-8ee1-79ed8621e223 · outbound

This paper cites Prefix-Tuning: Optimizing Continuous Prompts for Generation.

Flamingo: a Visual Language Model for Few-Shot Learning Prefix-Tuning: Optimizing Continuous Prompts for Generation

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-05-12T04:22:30.168758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:c8dceea17a5a12de61aabafed47f19e1683a66c717d8095f1c084d8688087a3a

Observation e523e3b5-51e6-466c-b4ee-8c819c1cabf8 · outbound

This paper cites Oscar: Object-semantics aligned pre-training for vision-language tasks.

Flamingo: a Visual Language Model for Few-Shot Learning Oscar: Object-semantics aligned pre-training for vision-language tasks

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:22:30.951694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:b024cbd29c43dd40f2e4b8cfefed63a950c28ad85ea38ca25e7ae503ed25b1ab

Observation a34ccb25-8607-46f6-84e5-f53e2b0621cc · outbound

This paper cites A Multimodal Framework for the Detection of Hateful Memes.

Flamingo: a Visual Language Model for Few-Shot Learning A Multimodal Framework for the Detection of Hateful Memes

Reference 63

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T04:22:30.194481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:c279e7399d35fe2150d13f72f7b578e47c997e09facd6138688fe2bfd4e66f5d

Observation 7a530106-adc3-4ac0-87bd-ddd4d26f4c9e · outbound

This paper cites What Makes Good In-Context Examples for GPT-$3$?.

Flamingo: a Visual Language Model for Few-Shot Learning What Makes Good In-Context Examples for GPT-$3$?

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:22:30.213836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:5f1af74c7be448de8bd88984ef7ad86ae4fa7823842c6ed24f39cdd37d3b05d4

Observation 8b6c790b-2c39-4094-9837-3eb47eec5511 · outbound

This paper cites Optimization of image description metrics using policy gradient methods.

Flamingo: a Visual Language Model for Few-Shot Learning Optimization of image description metrics using policy gradient methods

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:22:30.960786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:1fb8489dd101318b2fbeb741762f8861be41f81258da1370f81467109241cef7

Observation 7498eccc-96bf-4ef4-b13a-f291b2b46933 · outbound

This paper cites Enhancing textual cues in multi-modal transformers for VQA.

Flamingo: a Visual Language Model for Few-Shot Learning Enhancing textual cues in multi-modal transformers for VQA

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:22:30.966009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:049054d8acf7780cb25149b02d6b6b9b40395bfc91398e2f8c4d8f8e010f9da8

Observation 490fb7ff-823c-4f91-9033-c03443f64e8f · outbound

This paper cites ViLBERT: Pretraining task-agnostic vi- siolinguistic representations for vision-and-language tasks.

Flamingo: a Visual Language Model for Few-Shot Learning ViLBERT: Pretraining task-agnostic vi- siolinguistic representations for vision-and-language tasks

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:22:30.975253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:80a6d97527b0bf82b4c2bdf501fe9e62fa52c8d94bd93b9bc680fef73f40f9ac

Observation 9f765688-b808-4bfd-a3a2-09216f96be46 · outbound

This paper cites UniVL: A Unified Video and Language Pre-Training Model for Multimodal Understanding and Generation.

Flamingo: a Visual Language Model for Few-Shot Learning UniVL: A Unified Video and Language Pre-Training Model for Multimodal Understanding and Generation

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:22:30.263599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:c182249d71e3671e6700c98d770165e9bdb7d14c35257c3a3f952e615a2d4254

Observation 28f6a88b-2199-4b19-af78-b15f2cf623f6 · outbound

This paper cites A Frustratingly Simple Approach for End-to-End Image Captioning.

Flamingo: a Visual Language Model for Few-Shot Learning A Frustratingly Simple Approach for End-to-End Image Captioning

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:22:30.273063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:2e5ab8a86e12ed9167ec0ac757b60a05d24c27f1fd52c9e5f2da0a5f92e97a72

Observation e8ac43fa-e81f-4dfc-8719-a01cd7c9e02f · outbound

This paper cites OK-VQA: A visual question answering benchmark requiring external knowledge.

Flamingo: a Visual Language Model for Few-Shot Learning OK-VQA: A visual question answering benchmark requiring external knowledge

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:22:30.980690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:326b05e9f0a621eecd7a7bb1a8f58bb3f6465615e0728f844fcb8cd4e49e2c24

Observation 372a54c3-18bc-472e-a84e-bda27d0259a6 · outbound

This paper cites an unresolved cited work.

Flamingo: a Visual Language Model for Few-Shot Learning Unresolved cited work

Reference 71

Resolution
unresolved
raw_fallback, observed 2026-05-12T04:22:30.983928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:098e716e326c69aa8b0e14fc9c6476cac05d99878597c2e4a86e718c0ffc4c5f

Observation 4974358d-4d46-4865-b9bd-e6c07799abb8 · outbound

This paper cites an unresolved cited work.

Flamingo: a Visual Language Model for Few-Shot Learning Unresolved cited work

Reference 72

Resolution
unresolved
raw_fallback, observed 2026-05-12T04:22:30.988885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:77a52fb7f5b2248e71a2cda9c6e59c6e190c3e2042c247c303d9a95b139bf8db

Observation 4ee3dba6-2d8b-4274-88d5-c8f6596210eb · outbound

This paper cites Teaching language models to support answers with verified quotes.

Flamingo: a Visual Language Model for Few-Shot Learning Teaching language models to support answers with verified quotes

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-17T11:45:48.341649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:16b44acb0b44b457c278cf7b81692afdc8d05dde8b0f468c86a9917ca2976d64

Observation 83fcf28b-d0b9-466c-885b-bf3d4e9e643f · outbound

This paper cites RareAct: A video dataset of unusual interactions.

Flamingo: a Visual Language Model for Few-Shot Learning RareAct: A video dataset of unusual interactions

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:22:30.390215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:171aa7a27a6116d95c6fefd0fd4e24f43115ab9f8f734ac2a87b17276618246a

Observation ff778200-0cd8-4b78-86b6-c1169808bc9b · outbound

This paper cites End-to-end learning of visual representations from uncurated instructional videos.

Flamingo: a Visual Language Model for Few-Shot Learning End-to-end learning of visual representations from uncurated instructional videos

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:22:30.993024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:e836d8bdbe90aa7d4cfd71ad6b73b96a8cc2add4493fb06a2c65333c2d60abd8

Observation 9ba6a572-e902-4f87-8f6c-d3c728fe686e · outbound

This paper cites Recurrent neural network based language model.

Flamingo: a Visual Language Model for Few-Shot Learning Recurrent neural network based language model

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:22:30.996543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:14ed4ed28da5fb4e55474609b4008508762cfffd1b9548ec8250d84c9e3831eb

Observation 7b95a539-937a-49f6-941a-d79d59df9199 · outbound

This paper cites Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?.

Flamingo: a Visual Language Model for Few-Shot Learning Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?

Reference 77

Resolution
verified exact
arxiv_id, observed 2026-05-15T09:51:47.316034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:ac9c0c1a9eba2c972d8f078fb00e05b1420cdeaf5b019db135e4ea05127f44e2

Observation 20807713-4460-43ad-b734-45868a65a50b · outbound

This paper cites Model cards for model reporting.

Flamingo: a Visual Language Model for Few-Shot Learning Model cards for model reporting

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:22:31.002677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:768ff4555a262a44746b1282127a281746ca5c6e6059dedf74e30c11d593d08f

Observation 8cfeab8e-d3b3-4a90-9f76-16506dde437b · outbound

This paper cites ClipCap: CLIP Prefix for Image Captioning.

Flamingo: a Visual Language Model for Few-Shot Learning ClipCap: CLIP Prefix for Image Captioning

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:22:30.476566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:46e13fed07fc98a7b2091b963d46f475c7503f5126ff78332efdf7428e8ef609

Observation aca7b247-230e-42f3-b3f5-67eab3c7c250 · outbound

This paper cites Large-scale pretraining for visual dialog: A simple state-of-the-art baseline.

Flamingo: a Visual Language Model for Few-Shot Learning Large-scale pretraining for visual dialog: A simple state-of-the-art baseline

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:22:31.010928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:6c45b55f744c83b1f34b6d105431e19b662dd671126158f2d500d9b17a35c97d

Observation 616e981e-8d88-481a-b490-59a984121927 · outbound

This paper cites True few-shot learning with language models.

Flamingo: a Visual Language Model for Few-Shot Learning True few-shot learning with language models

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:22:31.014259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:c2d1fd674ebe745be2c6876b7227c97116eac99ec1583957470ada2c6415594d

Observation 7a5c0f43-5227-463e-b23d-a8c257e91a83 · outbound

This paper cites Red Teaming Language Models with Language Models.

Flamingo: a Visual Language Model for Few-Shot Learning Red Teaming Language Models with Language Models

Reference 82

Resolution
verified exact
local_arxiv, observed 2026-05-12T04:22:30.504465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:13d2fd922413ad04518f34d1f223c96628b7ec868b12d358c31902d66948aa41

Observation bbaac408-67bb-4bac-a332-c9c65a5031df · outbound

This paper cites Combined Scaling for Zero-shot Transfer Learning.

Flamingo: a Visual Language Model for Few-Shot Learning Combined Scaling for Zero-shot Transfer Learning

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:22:30.510632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:d7078253fab7c697b0d95ac8a620e5bf77658090aa846a1fa575a774a61a0286

Observation c2735049-6516-4cba-b590-1e88c0e74e32 · outbound

This paper cites Train short, test long: Attention with linear biases enables input length extrapolation.

Flamingo: a Visual Language Model for Few-Shot Learning Train short, test long: Attention with linear biases enables input length extrapolation

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:22:31.022431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:a6f816797b005283355ef7dad2f215107c327c42f956d190966d8cbda5562a96

Observation 08b64eb8-6d36-4c60-922e-8fc9acc85d5b · outbound

This paper cites Winner team Mia at TextVQA challenge 2021: Vision-and-language represen- tation learning with pre-trained sequence-to-sequence model.

Flamingo: a Visual Language Model for Few-Shot Learning Winner team Mia at TextVQA challenge 2021: Vision-and-language represen- tation learning with pre-trained sequence-to-sequence model

Reference 85

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:22:30.520078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:69ef134630b36b111acaa9671b7e6b9e1473ba9c8cb6cf6387f3e767ec6013e9

Observation c3bc78d9-7dc0-4754-be2b-33cedf04a679 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

Flamingo: a Visual Language Model for Few-Shot Learning Learning Transferable Visual Models From Natural Language Supervision

Reference 86

Resolution
verified exact
local_arxiv, observed 2026-05-12T04:22:30.554692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:716eceb08048d9643cbbff8da4db36f30ca812764a85aa8ae1284d5674c96ab3

Observation 70980115-d923-4419-a225-73de2a38f884 · outbound

This paper cites Scaling Language Models: Methods, Analysis & Insights from Training Gopher.

Flamingo: a Visual Language Model for Few-Shot Learning Scaling Language Models: Methods, Analysis & Insights from Training Gopher

Reference 87

Resolution
verified exact
local_arxiv, observed 2026-05-12T04:22:30.559656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:a15c3aca6fb3acae3940dcfee3b473c05277816512beabe978589c56a52e748c

Observation 48c1ee94-b9a4-4ae8-921e-6a8ec3292d6e · outbound

This paper cites Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer.

Flamingo: a Visual Language Model for Few-Shot Learning Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer

Reference 88

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:37:56.023799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:03dea366d9ff76d9f9a01ea88143ac20ff44f2b4e9efed4c1c6bf1decb2450a8

Observation 27b79a20-df73-4d08-9117-d9277f54bf52 · outbound

This paper cites ZeRO: Memory optimizations toward training trillion parameter models.

Flamingo: a Visual Language Model for Few-Shot Learning ZeRO: Memory optimizations toward training trillion parameter models

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:22:31.031659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:6ede6b4a86dea450a0009b795d15b9395e07c80a361ac04ed0f486111481834e

Observation 27b70176-76c5-4cc4-8716-cb730171ffca · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

Flamingo: a Visual Language Model for Few-Shot Learning Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 90

Resolution
verified exact
local_arxiv, observed 2026-05-12T04:22:30.579121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:7b3d64816710e9a318d397688005378c21f0158cd7f53b89776c6ef4c0c018ab

Observation d36e3c0e-b6b9-4022-b02e-0067dfbeb058 · outbound

This paper cites Rennie, Etienne Marcheret, Youssef Mroueh, Jarret Ross, and Vaibhava Goel.

Flamingo: a Visual Language Model for Few-Shot Learning Rennie, Etienne Marcheret, Youssef Mroueh, Jarret Ross, and Vaibhava Goel

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:22:31.036099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:815523e353f2a42dbbe5af60545e7514760ed496b55d322a48a478a23d4b801a

Observation 42c5e245-4082-46ca-9972-328ce2c084a8 · outbound

This paper cites an unresolved cited work.

Flamingo: a Visual Language Model for Few-Shot Learning Unresolved cited work

Reference 92

Resolution
unresolved
raw_fallback, observed 2026-05-12T04:22:31.045820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:cb977675c740fd6dff63b7f25dc06b5e9e19f6f4a8386602ac908e40435443b3

Observation 5e2b0d6b-c673-4990-867d-118aa2d20f61 · outbound

This paper cites Prompt programming for large language models: Beyond the few-shot paradigm.

Flamingo: a Visual Language Model for Few-Shot Learning Prompt programming for large language models: Beyond the few-shot paradigm

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:22:31.053058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:bcd1ac5a250b04283b6fc4f39c5deabbff40a28d7082765bd76459fa379d6c19

Observation 731cb385-27ff-495d-bb6d-f8bcb21ffaab · outbound

This paper cites Gender Bias in Coreference Resolution.

Flamingo: a Visual Language Model for Few-Shot Learning Gender Bias in Coreference Resolution

Reference 94

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:22:30.599173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:8adf6b06b75ab613042dd7f033d7ec387d0891f3f0c48341690e71e681354fae

Observation 4934788f-9125-4afc-ac1b-fe6fffcf17f4 · outbound

This paper cites Berg, and Li Fei- Fei.

Flamingo: a Visual Language Model for Few-Shot Learning Berg, and Li Fei- Fei

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:22:31.056748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:789974c5307526941503de1df58f9a3a041fa7b08d19cc2907f5b08a20cdb4d7

Observation 22c5877c-3547-4059-a63d-91b436eec5b6 · outbound

This paper cites Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Teven Le Scao, Arun Raja, Manan Dey, M.

Flamingo: a Visual Language Model for Few-Shot Learning Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Teven Le Scao, Arun Raja, Manan Dey, M

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:22:31.061331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:2f111e862465d6167e494fe34719ec2294ad280af4f206c1d6debcef65bb7394

Observation 553f0795-8a64-42aa-b1a2-f4c213f9d25e · outbound

This paper cites LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs.

Flamingo: a Visual Language Model for Few-Shot Learning LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs

Reference 97

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:21:01.130308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:4dbf1c28d896500d608e383d2c3a54f2a327dc206e4cd8292f70442a58b54fe4

Observation b9deedba-3c00-43f8-8c8c-efcc30e8394e · outbound

This paper cites Bello-Pardo, Stan Oklobdzija, Martijn Schoonvelde, and Jeffrey W.

Flamingo: a Visual Language Model for Few-Shot Learning Bello-Pardo, Stan Oklobdzija, Martijn Schoonvelde, and Jeffrey W

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:22:31.071346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:439874e6bfe9fbf1bde64e36faa8dfdaa02f1c8fe406682f9604cc036482752d

Observation e820973e-2f32-49ed-9716-7ccaeee8a986 · outbound

This paper cites Conceptual Captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning.

Flamingo: a Visual Language Model for Few-Shot Learning Conceptual Captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:22:31.074758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:cff31f55ab7adf136077f193ed0efe989cf24c208f9e6063448f6ef2f0f4308f

Observation ceb1e885-e3ec-4b0c-b0f9-90254d56b3ee · outbound

This paper cites The Power of Scale for Parameter-Efficient Prompt Tuning.

Flamingo: a Visual Language Model for Few-Shot Learning The Power of Scale for Parameter-Efficient Prompt Tuning

Reference 100

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T04:22:30.097920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:cca9353be4de0a9d165d3cad02176b3974dcda4428ccf1432abdca4c557b57d2

Observation c6a2f3f8-e170-4d77-b0da-956c33522965 · outbound

This paper cites Towards VQA models that can read.

Flamingo: a Visual Language Model for Few-Shot Learning Towards VQA models that can read

Reference 101

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:22:31.078131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:5ae948a68fed12074872a1a6f41eb3f735cb94f6b215c92c82d6610460c8a869

Pith citing papers

Observation ceaa511f-c208-41ba-a69d-9cc06cca3cf5 · inbound

A Generalist Agent cites this paper.

A Generalist Agent Flamingo: a Visual Language Model for Few-Shot Learning

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-13T06:24:49.870511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T06:24:49.833638Z digest=sha256:1249107e48710832645fab4b7c47b0b95996ebc69881137db1d330db4b6f1369

Observation 023dd40f-411f-4039-a98b-9996e60f3fd3 · inbound

GIT: A Generative Image-to-text Transformer for Vision and Language cites this paper.

GIT: A Generative Image-to-text Transformer for Vision and Language Flamingo: a Visual Language Model for Few-Shot Learning

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-16T20:54:07.617255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T20:54:07.572136Z digest=sha256:6c150a64ed2145041de685af50a2254838ebe8078530610b1468cb505a13b8f0

Observation 0bba0898-1685-4359-ae1d-b0edafcdd314 · inbound

Emergent Abilities of Large Language Models cites this paper.

Emergent Abilities of Large Language Models Flamingo: a Visual Language Model for Few-Shot Learning

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:22:31.147880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-11T07:38:37.734402Z digest=sha256:438e6526a2c9d69c6fb7edbc8511728d47e8bb389b07494c2b9f87997ce5074c

Observation 003da045-1e68-451d-8f6a-74272615c8d9 · inbound

Language Models (Mostly) Know What They Know cites this paper.

Language Models (Mostly) Know What They Know Flamingo: a Visual Language Model for Few-Shot Learning

Reference 207

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:22:31.147880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-10T15:42:47.274448Z digest=sha256:4d2fc655ab11b6a9531ba0a4d4131bbb2c0f1d51d41da6de399a2a24e3b6d018

Observation d421aa9f-bb12-4a5e-8317-8e1d01ad052f · inbound

Inner Monologue: Embodied Reasoning through Planning with Language Models cites this paper.

Inner Monologue: Embodied Reasoning through Planning with Language Models Flamingo: a Visual Language Model for Few-Shot Learning

Reference 90

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:22:31.147880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-11T20:10:43.912935Z digest=sha256:edd11ea99ff6cb7b557ba18b092b10aa8ad7aa4da258cbd318476f72a6ad72fa

Observation f7ea1a89-21fc-4c73-993e-ee22d25cab11 · inbound

DetailCLIP: Injecting Image Details into CLIP's Feature Space cites this paper.

DetailCLIP: Injecting Image Details into CLIP's Feature Space Flamingo: a Visual Language Model for Few-Shot Learning

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-05-24T11:09:22.385769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-24T11:08:20.298043Z digest=sha256:3b70317fbf9c0220b9cec0039d2ed331c8e9ea87558ee4b0458c70fede8ebcd5

Observation fb4224aa-f985-475c-81dc-3de7d3d78a29 · inbound

PaLI: A Jointly-Scaled Multilingual Language-Image Model cites this paper.

PaLI: A Jointly-Scaled Multilingual Language-Image Model Flamingo: a Visual Language Model for Few-Shot Learning

Reference 109

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T09:29:06.199377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-16T09:29:05.956863Z digest=sha256:60530f47b26d8887a60d702279bcf3b63fb7e27a297b44807638e62b3d54940b

Observation d68a3a7f-7a8e-4d40-8baa-1530aabb2bab · inbound

LAION-5B: An open large-scale dataset for training next generation image-text models cites this paper.

LAION-5B: An open large-scale dataset for training next generation image-text models Flamingo: a Visual Language Model for Few-Shot Learning

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-13T14:22:17.143757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T14:22:16.968028Z digest=sha256:af73ef5251fa7eaa7d3940abe952c89d58db04c4c45961a9dab66541feda8d0d

Observation d2049dda-3ef6-4bac-969b-013080f9c875 · inbound

Galactica: A Large Language Model for Science cites this paper.

Galactica: A Large Language Model for Science Flamingo: a Visual Language Model for Few-Shot Learning

Reference 106

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:53:22.023486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-13T05:53:21.810346Z digest=sha256:0275c8b79a587b6986be4eaa4ad71883dc40dbb919ba29957e53ea9331dc7507

Observation a5fdc50e-d18c-40f3-a275-45883e0e1775 · inbound

Galactica: A Large Language Model for Science cites this paper.

Galactica: A Large Language Model for Science Flamingo: a Visual Language Model for Few-Shot Learning

Reference 139

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:53:22.086217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-13T05:53:21.810346Z digest=sha256:8144fc5efe9de35eb30df5120f8be5b3c4a1fc1ddf47f5edacf40fe267939a5d

Observation 49c0baeb-7464-4264-972c-23a6bdcbc89b · inbound

InternVideo: General Video Foundation Models via Generative and Discriminative Learning cites this paper.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning Flamingo: a Visual Language Model for Few-Shot Learning

Reference 72

Resolution
verified exact
local_arxiv, observed 2026-05-17T00:36:53.325661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:7ee18a62fc16fe4a124b18b6dee20a7b92547efad48a8a5d33e642a33f107bfc

Observation 01787d4e-e25d-4691-845d-50fd19f16128 · inbound

Editing Models with Task Arithmetic cites this paper.

Editing Models with Task Arithmetic Flamingo: a Visual Language Model for Few-Shot Learning

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-13T08:09:12.887999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T08:09:12.716163Z digest=sha256:5b92c480deae4f1d9ee3490323577b3fdf972520fc0bdeab91bde03c452f7673

Observation fdc8c461-3412-4c4d-900c-f3fd96a9ad4f · inbound

BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models cites this paper.

BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models Flamingo: a Visual Language Model for Few-Shot Learning

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:22:31.147880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T00:10:48.610351Z digest=sha256:6a3f11ef35727e827733c62800b5778b5e4d3ccc416179143389b94601c04add

Observation c5a07d5d-ce5f-4116-89dd-2a3ae99ee3bb · inbound

Describe, Explain, Plan and Select: Interactive Planning with Large Language Models Enables Open-World Multi-Task Agents cites this paper.

Describe, Explain, Plan and Select: Interactive Planning with Large Language Models Enables Open-World Multi-Task Agents Flamingo: a Visual Language Model for Few-Shot Learning

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-16T03:27:40.765974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T03:27:40.524895Z digest=sha256:6f21205642349da2190402d9a8dfedee093df92057064f4d2fea6cfaefc3080b

Observation 2acc12f0-f8ff-46b8-b57c-a1068ea79606 · inbound

Scaling Robot Learning with Semantically Imagined Experience cites this paper.

Scaling Robot Learning with Semantically Imagined Experience Flamingo: a Visual Language Model for Few-Shot Learning

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-17T18:59:10.403564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T18:59:10.352342Z digest=sha256:5c454e2ec7d1ce7747867b874364097782343a924ef509ddc66bb15a72d89dcf

Observation 02f15836-2537-456b-b3bf-1d45accaf8a4 · inbound

PaLM-E: An Embodied Multimodal Language Model cites this paper.

PaLM-E: An Embodied Multimodal Language Model Flamingo: a Visual Language Model for Few-Shot Learning

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:22:31.147880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T22:29:29.631351Z digest=sha256:fec46af852e5364ea2ac8441566d56b741e426ef936605aeca965f831310d758

Observation a2de8743-c579-4899-93b5-ad66d1f37fb1 · inbound

MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action cites this paper.

MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action Flamingo: a Visual Language Model for Few-Shot Learning

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-14T01:17:58.725804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-14T01:17:58.678036Z digest=sha256:4f102a7ee9214e569d40002c9c54e73fee24eeaf248d48c642f326df5aacd600

Observation b753f1ca-d525-48a2-b9c6-4cba71bdd8aa · inbound

Visual Instruction Tuning cites this paper.

Visual Instruction Tuning Flamingo: a Visual Language Model for Few-Shot Learning

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:22:31.147880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-11T08:22:03.403362Z digest=sha256:7f34903f38121581e6a4ea3b8db67e43539d3443254f9692f8d657033a750ea5

Observation c966cff0-1189-4175-b608-21e0d3cb2f22 · inbound

mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality cites this paper.

mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality Flamingo: a Visual Language Model for Few-Shot Learning

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-24T09:04:15.076622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-24T09:02:31.211260Z digest=sha256:f68bfaa8f958278b3fa80b26cc8e7e0b787a2dedb340ca4bf3751c697f12f046

Observation d0a9af47-fe4a-4aea-bbe6-a3a1611f838a · inbound

Improving Factuality and Reasoning in Language Models through Multiagent Debate cites this paper.

Improving Factuality and Reasoning in Language Models through Multiagent Debate Flamingo: a Visual Language Model for Few-Shot Learning

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:22:31.147880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T03:01:45.164412Z digest=sha256:71cd91bd05137108b888b457de638189d35901131de7232522de6ed9d4064da8

Observation 2a38ea08-2339-4521-8dc1-6d55a11228b2 · inbound

The False Promise of Imitating Proprietary LLMs cites this paper.

The False Promise of Imitating Proprietary LLMs Flamingo: a Visual Language Model for Few-Shot Learning

Reference 286

Resolution
verified exact
local_arxiv, observed 2026-05-18T06:54:31.425210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-18T06:54:31.175090Z digest=sha256:48dfb81db8d29077c765235224a807e1ed97d5b94b8f83dacbb61162b50149b2

Observation c1f02b5c-6bb4-47f6-9628-33cb6967eb42 · inbound

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation cites this paper.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Flamingo: a Visual Language Model for Few-Shot Learning

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-15T06:30:22.494018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:45bb0c66aa2320cc531874ebb1162043ea79ae3df4c641e8c4f9ba1832bdba43

Observation 507f4628-57f0-4c6e-be51-69e72b93555e · inbound

DragNUWA: Fine-grained Control in Video Generation by Integrating Text, Image, and Trajectory cites this paper.

DragNUWA: Fine-grained Control in Video Generation by Integrating Text, Image, and Trajectory Flamingo: a Visual Language Model for Few-Shot Learning

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-05-20T13:03:57.983622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-20T13:03:57.828598Z digest=sha256:12b8cf7df0637a24755065c496ce083d9cd259c2251ae7966d3f2d76da62ee62

Observation e9a6250f-adb2-4b9d-8485-354dd1714aaf · inbound

Vision Transformers Need Registers cites this paper.

Vision Transformers Need Registers Flamingo: a Visual Language Model for Few-Shot Learning

Reference 264

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T09:41:38.181572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-13T09:41:37.937046Z digest=sha256:aa98cbf427ab472c0272bf6007dae87199cd8d8d221fe5c870e2021a33e5b51d

Observation 091873d3-ecd3-4905-9638-1e72546fe89c · inbound

Demystifying CLIP Data cites this paper.

Demystifying CLIP Data Flamingo: a Visual Language Model for Few-Shot Learning

Reference 46

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T09:20:20.278429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-16T09:20:20.143143Z digest=sha256:a65d585837245ad189fa299c7d14ddcd122a1995e3bffe2b25330152d32b530d

Observation b70f5547-fd25-462f-a089-27a04a67daaf · inbound

Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution cites this paper.

Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution Flamingo: a Visual Language Model for Few-Shot Learning

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-16T08:12:31.245786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-16T08:12:30.984870Z digest=sha256:92fcbcb24403b91d85d57bd4d845b44b8f1e2d518b59fa322d37c79b1357fe5c

Observation 20a9a1d6-8938-4b6c-932b-0fa46ddf4b97 · inbound

Improved Baselines with Visual Instruction Tuning cites this paper.

Improved Baselines with Visual Instruction Tuning Flamingo: a Visual Language Model for Few-Shot Learning

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-12T19:11:33.831931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T19:11:33.783746Z digest=sha256:0c2aea6850a94fac0836d9dfa117bd90f4af5c26bd65235db76ab03fe89ec59a

Observation 47b0421e-9227-4d52-b309-3a812e2031e0 · inbound

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark cites this paper.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Flamingo: a Visual Language Model for Few-Shot Learning

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:22:35.111518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:ad0f999803bff94d38307d75fa58e8789b8b927c70054ee92a51b67c2239f4ab

Observation f1fca3cd-e7a3-4f40-ad64-1d6790b80043 · inbound

GPT-4V(ision) is a Generalist Web Agent, if Grounded cites this paper.

GPT-4V(ision) is a Generalist Web Agent, if Grounded Flamingo: a Visual Language Model for Few-Shot Learning

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T19:40:13.885466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T19:40:13.816153Z digest=sha256:26290200b349cea2dd1276529513a90d2784b7745887c9791f3c523d66be5547

Observation a41e88da-fa86-44f2-95de-921cc11b02fd · inbound

VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs cites this paper.

VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs Flamingo: a Visual Language Model for Few-Shot Learning

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:22:31.147880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-11T02:44:53.284345Z digest=sha256:ed2a05433bce92cbe1881012eeb539c5eb36995547de0485244131baecd165d9

Observation 90723d26-e518-4bec-bcb4-51b9e928d255 · inbound

PaliGemma: A versatile 3B VLM for transfer cites this paper.

PaliGemma: A versatile 3B VLM for transfer Flamingo: a Visual Language Model for Few-Shot Learning

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:22:31.147880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-11T13:10:19.972353Z digest=sha256:c962bd8a95a7476b79df2d3688e35975b86435bcc466384dc0b7f150345386d8

Observation 4301052c-2903-45d7-9e0d-cba4e2288d59 · inbound

LLaVA-Video: Video Instruction Tuning With Synthetic Data cites this paper.

LLaVA-Video: Video Instruction Tuning With Synthetic Data Flamingo: a Visual Language Model for Few-Shot Learning

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:22:31.147880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:e7753162c526beb65ce368fd57eeacbf67cf687b32c6cb96f50d362bb95f1e30

Observation abfa6e9c-6ade-47b0-9d00-6297e739ebbf · inbound

Gla-AI4BioMed at RRG24: Visual Instruction-tuned Adaptation for Radiology Report Generation cites this paper.

Gla-AI4BioMed at RRG24: Visual Instruction-tuned Adaptation for Radiology Report Generation Flamingo: a Visual Language Model for Few-Shot Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T21:10:16.529113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T21:10:16.529113Z digest=sha256:61fadacb30d205314c8eaed352a0a642812fbedc3e79fa694a82ed963582d255

Observation 675d14c4-d3a2-45a6-a40d-716825cb5614 · inbound

GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning cites this paper.

GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning Flamingo: a Visual Language Model for Few-Shot Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T18:40:21.730017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:40:21.730017Z digest=sha256:9145b8d4c75fd5211a61cc090e60e8f2792987fd7a736ac8cea543346fe6c94c

Observation 5edea55d-0b14-4aca-a234-78064cce8fe4 · inbound

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents cites this paper.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Flamingo: a Visual Language Model for Few-Shot Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:44.976022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:44.976022Z digest=sha256:5b607efd7c2aa5d0b4867515e807c9b6fb7a5496ca153398d32b0c385fc2b7d7

Observation 2fc05dbd-d25b-4d61-918b-68a7861d919a · inbound

GaussianProperty: Integrating Physical Properties to 3D Gaussians with LMMs cites this paper.

GaussianProperty: Integrating Physical Properties to 3D Gaussians with LMMs Flamingo: a Visual Language Model for Few-Shot Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T15:12:52.352469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:12:52.352469Z digest=sha256:a4697596b4785c93157a126d0e791337b233b9508b697d256a16f10d2f2120f3

Observation f692c8d6-53e4-471b-b6b3-4483b6121e35 · inbound

Improving Fine-grained Visual Understanding in VLMs through Text-Only Training cites this paper.

Improving Fine-grained Visual Understanding in VLMs through Text-Only Training Flamingo: a Visual Language Model for Few-Shot Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T13:38:48.419921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:38:48.419921Z digest=sha256:8163eb8e995dfbc6f124a9eaf6803a8c5f44743fa039f065bd1995b417887a79

Observation cd9a2d0f-4f81-454b-86d7-ca7e8b50d817 · inbound

Movie2Story: A framework for understanding videos and telling stories in the form of novel text cites this paper.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text Flamingo: a Visual Language Model for Few-Shot Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.102344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.102344Z digest=sha256:59e0b68e35747e519a585a395a45937d3d8e75e8f2f2c7aa51417c435f40965f

Observation 67733d41-6039-4a1c-98b0-ba43ace8a2d0 · inbound

RapGuard: Safeguarding Multimodal Large Language Models via Rationale-aware Defensive Prompting cites this paper.

RapGuard: Safeguarding Multimodal Large Language Models via Rationale-aware Defensive Prompting Flamingo: a Visual Language Model for Few-Shot Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T04:29:04.180247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:29:04.180247Z digest=sha256:7bbf98f7432313ac4b123547d91f96aa75924ef85f53ebcd997ccce92344adec

Observation 02129c57-bde1-4233-92de-92c5abcbf6db · inbound

Probing Visual Language Priors in VLMs cites this paper.

Probing Visual Language Priors in VLMs Flamingo: a Visual Language Model for Few-Shot Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T22:55:53.830145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:55:53.830145Z digest=sha256:2c131c3dabb8acc5ea939653d4323caa372bf54514ac9ccd0a3894854d3c3808

Observation b9c4fec2-817a-4805-9ccd-b4a3e6751c7d · inbound

Spot Risks Before Speaking! Unraveling Safety Attention Heads in Large Vision-Language Models cites this paper.

Spot Risks Before Speaking! Unraveling Safety Attention Heads in Large Vision-Language Models Flamingo: a Visual Language Model for Few-Shot Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T22:28:31.925826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:28:31.925826Z digest=sha256:eb61cb2253bdc9110ddd119b905c5a93b640ae927e5114a99a5f8e18f0617104

Observation 8fdf7032-bfc8-4664-b3fa-fdda4914506e · inbound

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends cites this paper.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends Flamingo: a Visual Language Model for Few-Shot Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:27.398595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:27.398595Z digest=sha256:7e51f9694490885ea7aadb337f9ba773e5726066f76c88f14e3b5c25b0d083b8

Observation 8c3bd792-8956-4fc2-b8c7-bf9bcec7fb30 · inbound

Visual question answering: from early developments to recent advances -- a survey cites this paper.

Visual question answering: from early developments to recent advances -- a survey Flamingo: a Visual Language Model for Few-Shot Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T21:46:28.028582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:46:28.028582Z digest=sha256:af4b10fe3c72a31066b95ca6de173a725ce65b66f376e17aa49f7d0f261dbebc

Observation 7d57b587-6a4a-43eb-b6c3-75d10307ed32 · inbound

The Quest for Visual Understanding: A Journey Through the Evolution of Visual Question Answering cites this paper.

The Quest for Visual Understanding: A Journey Through the Evolution of Visual Question Answering Flamingo: a Visual Language Model for Few-Shot Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T20:51:53.527412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:51:53.527412Z digest=sha256:0c6b7abb0dc6aec531ec6b44e0b438d7de1667d59b091811b75f4bcf28c0d84b

Observation 05431307-5a32-4082-8371-4b90293353ee · inbound

Visual Language Models as Operator Agents in the Space Domain cites this paper.

Visual Language Models as Operator Agents in the Space Domain Flamingo: a Visual Language Model for Few-Shot Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:43.912259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:43.912259Z digest=sha256:ec0703c175206c9d774d98f50bcd03b0645a96b4e31d0c933b7508fc5a8dfe87

Observation 48cbdffa-6db0-4dea-9421-54a27d1af176 · inbound

Towards Safer Social Media Platforms: Scalable and Performant Few-Shot Harmful Content Moderation Using Large Language Models cites this paper.

Towards Safer Social Media Platforms: Scalable and Performant Few-Shot Harmful Content Moderation Using Large Language Models Flamingo: a Visual Language Model for Few-Shot Learning

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-10T16:22:20.758477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:22:20.758477Z digest=sha256:37b4d2aded81edd94ea65d0d20c6418f93238fa00c2a056aae1272f48f4cc6c3

Observation 8f521cca-703d-417b-88cf-938d9f19fdc2 · inbound

PuzzleGPT: Emulating Human Puzzle-Solving Ability for Time and Location Prediction cites this paper.

PuzzleGPT: Emulating Human Puzzle-Solving Ability for Time and Location Prediction Flamingo: a Visual Language Model for Few-Shot Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T15:17:54.010963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:17:54.010963Z digest=sha256:482f3b860660ea9dafe6ffdf1672d29fb1872269e2805614b39e9450c67bf7d9

Observation 6f91db66-a518-4fdf-860f-52a5e47d5eb1 · inbound

Beyond-Labels: Advancing Open-Vocabulary Segmentation With Vision-Language Models cites this paper.

Beyond-Labels: Advancing Open-Vocabulary Segmentation With Vision-Language Models Flamingo: a Visual Language Model for Few-Shot Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T10:54:06.048139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:54:06.048139Z digest=sha256:1ce248b3587181ddb6935e935f8af39077f6a94b5b0133f774997bdfee9575da

Observation 5d2ec1ba-9818-4141-b1c4-60a0ad8879fd · inbound

Generative AI for Vision: A Comprehensive Study of Frameworks and Applications cites this paper.

Generative AI for Vision: A Comprehensive Study of Frameworks and Applications Flamingo: a Visual Language Model for Few-Shot Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T00:58:34.687621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:58:34.687621Z digest=sha256:9616cec010dd021fc8496cf3e6498ce7cc1d9fda8d1a1718b7f8c6d9fa0def03

Observation 7bf96aeb-feb5-4280-8404-bf68c1088b43 · inbound

"I am bad": Interpreting Stealthy, Universal and Robust Audio Jailbreaks in Audio-Language Models cites this paper.

"I am bad": Interpreting Stealthy, Universal and Robust Audio Jailbreaks in Audio-Language Models Flamingo: a Visual Language Model for Few-Shot Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T18:05:52.203599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:05:52.203599Z digest=sha256:29331aa83e978f57b50ccc3590140b651adf261bd37aa9149c859e29c7be3b96

Observation c1f4af97-9df2-48a2-9752-2493d90474de · inbound

Mirai: A Wearable Proactive AI "Inner-Voice" for Contextual Nudging cites this paper.

Mirai: A Wearable Proactive AI "Inner-Voice" for Contextual Nudging Flamingo: a Visual Language Model for Few-Shot Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-09T12:25:08.685367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:25:08.685367Z digest=sha256:09ffeec512b1ff1c6d6be133a14532d23e48ccc745fbfbff3d7b4c18c7384151

Observation cc374f1e-019e-4885-8ded-cd4e36732e6b · inbound

NanoVLMs: How small can we go and still make coherent Vision Language Models? cites this paper.

NanoVLMs: How small can we go and still make coherent Vision Language Models? Flamingo: a Visual Language Model for Few-Shot Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T13:35:05.044985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:35:05.044985Z digest=sha256:621e4f3e603be231d69ad060d2d2f2acfa57e2454e74da206a451d1d1ec4260b

Observation 142373a9-c858-4b13-9cd2-3eb74d7aae02 · inbound

A Real-to-Sim-to-Real Approach to Robotic Manipulation with VLM-Generated Iterative Keypoint Rewards cites this paper.

A Real-to-Sim-to-Real Approach to Robotic Manipulation with VLM-Generated Iterative Keypoint Rewards Flamingo: a Visual Language Model for Few-Shot Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T00:03:16.308743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:03:16.308743Z digest=sha256:9bf56e8f7ba5bac2abd585d5e932238e780022abb109402d938a49c77248cdf7

Observation e3a99080-5ac5-41f9-a351-04e6bdf735ed · inbound

AIDE: Agentically Improve Visual Language Model with Domain Experts cites this paper.

AIDE: Agentically Improve Visual Language Model with Domain Experts Flamingo: a Visual Language Model for Few-Shot Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T22:51:02.035005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T22:51:02.035005Z digest=sha256:1da0609b0316abddea2fef2b2965b26969adc2e16ceeb9b093be3ebf488d48d3

Observation cb625869-bd87-4cf4-8317-493e3f46aed4 · inbound

Granite Vision: a lightweight, open-source multimodal model for enterprise Intelligence cites this paper.

Granite Vision: a lightweight, open-source multimodal model for enterprise Intelligence Flamingo: a Visual Language Model for Few-Shot Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T20:06:00.672714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T20:06:00.672714Z digest=sha256:1fd1113d0a41c4825738fa19e2bcb8c05d0f9c9b8e6cd91e1d0d9a90f1666069

Observation 1fe77a1d-50ec-4d0c-b45f-31d07352f173 · inbound

GPUs, CPUs, and... NICs: Rethinking the Network's Role in Serving Complex AI Pipelines cites this paper.

GPUs, CPUs, and... NICs: Rethinking the Network's Role in Serving Complex AI Pipelines Flamingo: a Visual Language Model for Few-Shot Learning

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-10T16:33:23.448566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:33:23.448566Z digest=sha256:d173f792fe1916e196721a9611608abe172990a7075a50cf3783711ff308aa31

Observation 48a3e18a-4e49-49f5-ae0b-0819e9da3985 · inbound

When VLMs Meet Image Classification: Test Sets Renovation via Missing Label Identification cites this paper.

When VLMs Meet Image Classification: Test Sets Renovation via Missing Label Identification Flamingo: a Visual Language Model for Few-Shot Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:06.972843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:06.972843Z digest=sha256:b48faf10dcc7c457ffba7f6982db9aa23bf2325b8596dd035bedec476a730d7b

Observation 73c87c79-8445-4673-924d-e5bcdbb80462 · inbound

ChemMLLM: Chemical Multimodal Large Language Model cites this paper.

ChemMLLM: Chemical Multimodal Large Language Model Flamingo: a Visual Language Model for Few-Shot Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:04.134589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:06:04.134589Z digest=sha256:66ee20140d8ec798007ead9128fcd94440ad4791f6ad7821e055e9412eb6b3e1

Observation 796be892-b44a-4377-ac8d-ab004995c347 · inbound

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion cites this paper.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Flamingo: a Visual Language Model for Few-Shot Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:47.145455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:47.145455Z digest=sha256:f691c4a82e23924f5f21a226fccddea9411d52e96161b81a4f82898258901d1e

Observation f5bde0b2-ca05-4e48-90ce-8d2ca28116ef · inbound

LA-RCS: LLM-Agent-Based Robot Control System cites this paper.

LA-RCS: LLM-Agent-Based Robot Control System Flamingo: a Visual Language Model for Few-Shot Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:51:05.230269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:51:05.230269Z digest=sha256:a92aa14bb4f94b9e451f2a31fe7c9e2c41ed1aa109a7d2e5e8231ee4e788f505

Observation 90463c61-bf5e-443d-8851-0b8611e30f46 · inbound

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction cites this paper.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction Flamingo: a Visual Language Model for Few-Shot Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:02.224914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:02.224914Z digest=sha256:923f2d51d505d177017828121fa97c6d4b7a143f113d98a4168b67be73feb895

Observation 2d9246e3-756a-499c-88e0-38d0731381b7 · inbound

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents cites this paper.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Flamingo: a Visual Language Model for Few-Shot Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:22.058110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:22.058110Z digest=sha256:0a3f3e65a768e6e75607ed2decb04516a358d3c6e1518346a40fef42e4910348

Observation f205f64e-d814-4608-b617-3f2535bf5cc7 · inbound

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times cites this paper.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Flamingo: a Visual Language Model for Few-Shot Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.090193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.090193Z digest=sha256:27c2e3ff1782a9d18703fcf447b0416fc165f46abd8dc918796d291825270b80

Observation 43a30cec-1dad-4cb1-9574-8989f357a51c · inbound

NavBench: Probing Multimodal Large Language Models for Embodied Navigation cites this paper.

NavBench: Probing Multimodal Large Language Models for Embodied Navigation Flamingo: a Visual Language Model for Few-Shot Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:32.233009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:32.233009Z digest=sha256:d6a4f9735ae71ad57835e6a51788c582e0696df0fe6de3941ff5e2944d70920c

Observation eed9cf06-d45b-4a21-88c0-8dfa30cc481d · inbound

ANT: Adaptive Neural Temporal-Aware Text-to-Motion Model cites this paper.

ANT: Adaptive Neural Temporal-Aware Text-to-Motion Model Flamingo: a Visual Language Model for Few-Shot Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:27:47.059429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:27:47.059429Z digest=sha256:a4321485de244cc82302c9dbd561711f58d014bfde5361b22d6089ee8067af3d

Observation 910ecde8-b4f4-4ebf-8984-efea7b666151 · inbound

Towards an Explainable Comparison and Alignment of Feature Embeddings cites this paper.

Towards an Explainable Comparison and Alignment of Feature Embeddings Flamingo: a Visual Language Model for Few-Shot Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:39.399224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:39.399224Z digest=sha256:5696e78e3f95a20987345381cad2910912a863397233ad2598087865adc2f257

Observation f4858d70-7495-440f-a5b2-17793a48ed6a · inbound

Large Language Models for EEG: A Comprehensive Survey and Taxonomy cites this paper.

Large Language Models for EEG: A Comprehensive Survey and Taxonomy Flamingo: a Visual Language Model for Few-Shot Learning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T11:30:53.496195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:30:53.496195Z digest=sha256:7739352639bec6eaed560dbb30a30fc021992d1fcfdd4079d063ed3cb8fbd26d

Observation 3da569a7-e35e-43b7-831e-8a15fb0eeb92 · inbound

Hidden in plain sight: VLMs overlook their visual representations cites this paper.

Hidden in plain sight: VLMs overlook their visual representations Flamingo: a Visual Language Model for Few-Shot Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.832536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.832536Z digest=sha256:433ea1733fcb6e8f0d3f49297d063eccf8ad6b9147a597e0f6978c325667a678

Observation d3195398-5253-497d-a3c5-705797a42e3e · inbound

LLaVA-c: Continual Improved Visual Instruction Tuning cites this paper.

LLaVA-c: Continual Improved Visual Instruction Tuning Flamingo: a Visual Language Model for Few-Shot Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T05:10:48.454540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:10:48.454540Z digest=sha256:7b64117142f08c2c127b25f4297a742e353bc6cc06bca97837cad54d0e91e65d

Observation a2c6ce8e-7a85-4ee8-bff5-7d04aa538d59 · inbound

Kvasir-VQA-x1: A Multimodal Dataset for Medical Reasoning and Robust MedVQA in Gastrointestinal Endoscopy cites this paper.

Kvasir-VQA-x1: A Multimodal Dataset for Medical Reasoning and Robust MedVQA in Gastrointestinal Endoscopy Flamingo: a Visual Language Model for Few-Shot Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T04:42:02.564894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:42:02.564894Z digest=sha256:63bcfb411d554c42e74659a9b78a50e862d95e619c0aa1c016fad99f351438fb

Observation 086ef63c-0e52-4f4d-a69f-f2d391546eae · inbound

EfficientVLA: Training-Free Acceleration and Compression for Vision-Language-Action Models cites this paper.

EfficientVLA: Training-Free Acceleration and Compression for Vision-Language-Action Models Flamingo: a Visual Language Model for Few-Shot Learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T04:42:22.457907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:42:22.457907Z digest=sha256:b52ad2e0a6e87a8f1e36d8b1494b1eef85d95b55854554383a26de0e0bb82875

Observation 7a7396c1-0d9e-40bb-8825-e7a11f03ed87 · inbound

A Navigation Framework Utilizing Vision-Language Models cites this paper.

A Navigation Framework Utilizing Vision-Language Models Flamingo: a Visual Language Model for Few-Shot Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T04:35:44.061335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:35:44.061335Z digest=sha256:fde2c69e795a01292a3937b99f935509c42072c4b1601da9504d4f62502774b9

Observation 63157067-04a8-4578-be86-a399723a710e · inbound

Dynamic Modality Scheduling for Multimodal Large Models via Confidence, Uncertainty, and Semantic Consistency cites this paper.

Dynamic Modality Scheduling for Multimodal Large Models via Confidence, Uncertainty, and Semantic Consistency Flamingo: a Visual Language Model for Few-Shot Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T00:47:50.343150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:47:50.343150Z digest=sha256:9e692da387b90724e4dd673e6a61088f6a5b48907e22b7134247cfd21aa5d9e1

Observation 047da159-882b-4a99-82a3-4ae4a11b850c · inbound

Poison Once, Control Anywhere: Clean-Text Visual Backdoors in VLM-based Mobile Agents cites this paper.

Poison Once, Control Anywhere: Clean-Text Visual Backdoors in VLM-based Mobile Agents Flamingo: a Visual Language Model for Few-Shot Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T00:41:16.316084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:41:16.316084Z digest=sha256:b109360d399b06ad5246a431d6782720bfdc5534b2d56774d4a87dec559873fe

Observation 5ddc843a-52dd-4c47-bc13-381f8f0a0edb · inbound

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation cites this paper.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Flamingo: a Visual Language Model for Few-Shot Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:11.568556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:11.568556Z digest=sha256:e94ddf75d123ae800a8a3e08dde83712af40f8b0c1e6baf5e4103fb3d5c92cfc

Observation 4a1a7083-b142-43d2-96a9-2404d4e8f116 · inbound

Shape2Animal: Creative Animal Generation from Natural Silhouettes cites this paper.

Shape2Animal: Creative Animal Generation from Natural Silhouettes Flamingo: a Visual Language Model for Few-Shot Learning

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-19T07:27:08.866082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T07:25:42.121060Z digest=sha256:1a5f5c87e48936dc05f1d2dbceb08f9a23fd36409750caee328b4b1e47888383

Observation 9eb300ff-d89e-499c-9e39-8ca49f5425d9 · inbound

Symbolic or Numerical? Understanding Physics Problem Solving in Reasoning LLMs cites this paper.

Symbolic or Numerical? Understanding Physics Problem Solving in Reasoning LLMs Flamingo: a Visual Language Model for Few-Shot Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T20:59:07.191372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:59:07.191372Z digest=sha256:db44ee83db9a7df9f0d6b73585a773236cfd676d4841f9fd1ec083d01cfb099c

Observation 45eb87ab-6ff5-493e-a973-cf7ec64bea59 · inbound

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models cites this paper.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Flamingo: a Visual Language Model for Few-Shot Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:27.564674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:27.564674Z digest=sha256:c5443d794074c54b9a61bf6f7ae760aa0ae4c34de47fb943354754043de55e46

Observation aef837c9-0992-4bdb-94ae-3e0afcd318a1 · inbound

CS-VLM: Compressed Sensing Attention for Efficient Vision-Language Representation Learning cites this paper.

CS-VLM: Compressed Sensing Attention for Efficient Vision-Language Representation Learning Flamingo: a Visual Language Model for Few-Shot Learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T21:46:07.771553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:46:07.771553Z digest=sha256:38f510ea37e6ba22272a2ac484b904e025e021d57cbc1cc412750587a21eab50

Observation 0bdd19ee-a710-47ff-b531-5ce9f118ab5d · inbound

Beyond Accuracy: Metrics that Uncover What Makes a 'Good' Visual Descriptor cites this paper.

Beyond Accuracy: Metrics that Uncover What Makes a 'Good' Visual Descriptor Flamingo: a Visual Language Model for Few-Shot Learning

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T20:13:44.966493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:13:44.966493Z digest=sha256:b4b6e7d93b1b9cf2c832846fa11a9b5c97acd6932c59f04a6b8c42b3bbf0379e

Observation b47d0111-3248-4caf-b0fe-8b3b56c6c592 · inbound

Uncertainty-Driven Expert Control: Enhancing the Reliability of Medical Vision-Language Models cites this paper.

Uncertainty-Driven Expert Control: Enhancing the Reliability of Medical Vision-Language Models Flamingo: a Visual Language Model for Few-Shot Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T18:06:23.486547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:06:23.486547Z digest=sha256:7402b4c6a1321b0a9c2069fa7a3620058d3eb22cc682c047115d48a6e3c3a410

Observation 01a6443d-3f02-4fbf-947b-b3fce020aaf1 · inbound

Reframing SAR Target Recognition as Visual Reasoning: A Chain-of-Thought Dataset with Multimodal LLMs cites this paper.

Reframing SAR Target Recognition as Visual Reasoning: A Chain-of-Thought Dataset with Multimodal LLMs Flamingo: a Visual Language Model for Few-Shot Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T17:57:40.312626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:57:40.312626Z digest=sha256:f7b4e0bb50f64bda560f230a8670d4534b944f64d747735ec83523780e6648a5

Observation 73415d10-33d0-4ca3-9d97-e4ad7e0be999 · inbound

Argus: Leveraging Multiview Images for Improved 3-D Scene Understanding With Large Language Models cites this paper.

Argus: Leveraging Multiview Images for Improved 3-D Scene Understanding With Large Language Models Flamingo: a Visual Language Model for Few-Shot Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T16:42:13.216833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:42:13.216833Z digest=sha256:1b2b35341ea0887f9aa8ba147d49f9aa770cdb055723c2f5f3f2df6aff42bf0d

Observation c2067c6a-28c9-4368-b872-dce6d1a22381 · inbound

Differential Multimodal Transformers cites this paper.

Differential Multimodal Transformers Flamingo: a Visual Language Model for Few-Shot Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T16:39:19.755081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:39:19.755081Z digest=sha256:79617d815ae6c7411e654971ccede1efb889cd0ba500beaf56990a102dbdb298

Observation 9fa5cd60-95ba-42e4-871e-3961f6655280 · inbound

Visual Language Models as Zero-Shot Deepfake Detectors cites this paper.

Visual Language Models as Zero-Shot Deepfake Detectors Flamingo: a Visual Language Model for Few-Shot Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:25.682150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:45:25.682150Z digest=sha256:eb82edb51737e2793682e790cc2a70942326e9638b75de7b55302ca51743de1d

Observation d743f260-fd95-42d4-8140-e576d9fd42ed · inbound

Augmented Vision-Language Models: A Systematic Review cites this paper.

Augmented Vision-Language Models: A Systematic Review Flamingo: a Visual Language Model for Few-Shot Learning

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T14:33:29.392262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:33:29.392262Z digest=sha256:616a07b592d673b9e83bfcaf323e69fcf3192adb40bef76aa7ac18a3114adb63

Observation dc5a86d8-dd3d-48ca-a4c8-8cc0eb4a778d · inbound

Multi-TW: Benchmarking Multimodal Models on Traditional Chinese Question Answering in Taiwan cites this paper.

Multi-TW: Benchmarking Multimodal Models on Traditional Chinese Question Answering in Taiwan Flamingo: a Visual Language Model for Few-Shot Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T05:49:51.471076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:49:51.471076Z digest=sha256:fb6b624cb01c32c4dd982ccaccdba14a1f4bacefb3515216aa732ba511a2281b

Observation 2b7a9b58-17b4-4205-a330-8714e8cd6787 · inbound

LEARN: A Story-Driven Layout-to-Image Generation Framework for STEM Instruction cites this paper.

LEARN: A Story-Driven Layout-to-Image Generation Framework for STEM Instruction Flamingo: a Visual Language Model for Few-Shot Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T20:10:56.814518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:10:56.814518Z digest=sha256:880dfa77e102b2c944a03a2d07128baeee199f79a51128c7cea63a707ddb8396

Observation 7df1ea6a-cf53-4abe-8f61-2d24ae8de02a · inbound

Region-Level Context-Aware Multimodal Understanding cites this paper.

Region-Level Context-Aware Multimodal Understanding Flamingo: a Visual Language Model for Few-Shot Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T19:39:48.032858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:39:48.032858Z digest=sha256:ca85af08704a79673afa3869a8c192a1f3d60b10aa1169fda9168a1a9f86d98c

Observation 5b34e853-a4bf-445b-9064-f9c98f78ffde · inbound

LLaSO: A Foundational Framework for Reproducible Research in Large Language and Speech Model cites this paper.

LLaSO: A Foundational Framework for Reproducible Research in Large Language and Speech Model Flamingo: a Visual Language Model for Few-Shot Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T17:56:49.630598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T17:56:49.630598Z digest=sha256:77fae178471c8c54507f6624261aaff35146187a9de0cd771b1d444cdd8795a5

Observation 97c46b42-392a-4835-a400-0d62648a5eae · inbound

Time-Scaling State-Space Models for Dense Video Captioning cites this paper.

Time-Scaling State-Space Models for Dense Video Captioning Flamingo: a Visual Language Model for Few-Shot Learning

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-05T10:59:05.052003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:59:05.052003Z digest=sha256:aafc570794229dcc79383226b3dd901ace3b6f8b003370a7973010609f0ea81a

Observation 80464c7d-37da-4619-a08e-6043273e5877 · inbound

Fine-Tuning Vision-Language Models for Visual Navigation Assistance cites this paper.

Fine-Tuning Vision-Language Models for Visual Navigation Assistance Flamingo: a Visual Language Model for Few-Shot Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T22:10:23.484160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T22:10:23.484160Z digest=sha256:6c1ce8241c829518ec71f5aa0baf2424f47533199874abd4c81186add6fc30d9

Observation 6792aaef-5272-4ac6-8c45-0e485246dc4a · inbound

Toward a Vision-Language Foundation Model for Medical Data: Multimodal Dataset and Benchmarks for Vietnamese PET/CT Report Generation cites this paper.

Toward a Vision-Language Foundation Model for Medical Data: Multimodal Dataset and Benchmarks for Vietnamese PET/CT Report Generation Flamingo: a Visual Language Model for Few-Shot Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T13:51:48.855655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:51:48.855655Z digest=sha256:9f1b902394b348191fd4aab50a40cd39977b943cae9557f6a035b159a0ced919

Observation c3ce60a4-f745-44e6-921f-7c994602a1d8 · inbound

Multimodal Function Vectors for Visual Relations cites this paper.

Multimodal Function Vectors for Visual Relations Flamingo: a Visual Language Model for Few-Shot Learning

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-04T12:44:41.101436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:44:41.101436Z digest=sha256:85989d13c5f71714b846066f9ea61f644e37b0e175da62cd68fa49a5d269dd15

Observation 6a78b63a-8516-425c-af10-365c57d7375f · inbound

Adaptive Residual-Update Steering for Low-Overhead Hallucination Mitigation in Large Vision Language Models cites this paper.

Adaptive Residual-Update Steering for Low-Overhead Hallucination Mitigation in Large Vision Language Models Flamingo: a Visual Language Model for Few-Shot Learning

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-21T18:50:30.118231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-21T18:47:28.768554Z digest=sha256:ebb09a60837b08850e1c13437630d3bba0d77cd01e196945e03f51e392c46ef1

Observation 798de04f-8adf-42b6-9def-1d03258404dd · inbound

AVATAAR: Agentic Video Answering via Temporal Adaptive Alignment and Reasoning cites this paper.

AVATAAR: Agentic Video Answering via Temporal Adaptive Alignment and Reasoning Flamingo: a Visual Language Model for Few-Shot Learning

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:20:11.809708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:18:19.580156Z digest=sha256:8350077b620fe4be18b1cea276ccc7bf3fb28efd51c61d07b5be4cd946ede059

Observation 9749a4ea-f322-41c3-b0f5-35db95e60391 · inbound

SPEAR-1: Scaling Beyond Robot Demonstrations via 3D Understanding cites this paper.

SPEAR-1: Scaling Beyond Robot Demonstrations via 3D Understanding Flamingo: a Visual Language Model for Few-Shot Learning

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:22:04.564118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:21:12.375936Z digest=sha256:5512fe27ec714ab3289c6be1db6408432195aa47beed0f385db8fe35ca74cd75

Observation 912dd5cc-333f-4773-ae8a-737c921ee3a1 · inbound

Reconstructing Content with Collaborative Attention for Universal Multimodal Representation Learning cites this paper.

Reconstructing Content with Collaborative Attention for Universal Multimodal Representation Learning Flamingo: a Visual Language Model for Few-Shot Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T19:41:30.290370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:41:30.290370Z digest=sha256:47f075a55327a6d01d5b1386cb495591c3dcc674de1cf685b9775dce01aa4b02

Observation 2e0bd53e-96da-4cc8-9a63-7efccc8e4a6d · inbound

Less Data, Faster Convergence: Goal-Driven Data Optimization for Multimodal Instruction Tuning cites this paper.

Less Data, Faster Convergence: Goal-Driven Data Optimization for Multimodal Instruction Tuning Flamingo: a Visual Language Model for Few-Shot Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-14T22:18:56.115559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T22:18:56.115559Z digest=sha256:2b9d77b5f799f5b3735e9fe362f7b36e78af9d76723ed7dbd67bb1adb37ddb53

Observation 82e3e203-a7f5-4f03-96d1-453d30d21d7e · inbound

SocialOmni: Benchmarking Audio-Visual Social Interactivity in Omni Models cites this paper.

SocialOmni: Benchmarking Audio-Visual Social Interactivity in Omni Models Flamingo: a Visual Language Model for Few-Shot Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-13T23:27:58.971096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:27:58.971096Z digest=sha256:9b6573f2c26f88cd976554a09581cabb41e5422c7a332317b3793deefd47508b