Pith. sign in

Paper Citation Record · LEDGER

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey

As of 12 August 2026, this Paper Citation Record lists 100 of 229 outbound references and 0 inbound Pith citation observations for arXiv:2412.08158.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.08158 v1

Coverage vector

measured 100 of 229 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T18:11:54.632569Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 229 outbound references displayed

  • verified exact8
  • verified fuzzy0
  • unresolved92
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 465a3771-17b5-4231-846f-1cc1323618ff · outbound

This paper cites Multimodal research in vision and language: A review of current and emerging trends,.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Multimodal research in vision and language: A review of current and emerging trends,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.151310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.151310Z digest=sha256:fef99d593efc535dab1cfcbf1c45ce27b62361bde3f0b964ff5bda2999cdbb60

Observation e66d460d-6aa1-481a-a634-6b91c7030011 · outbound

This paper cites Vision+ language applications: A survey,.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Vision+ language applications: A survey,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.156570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.156570Z digest=sha256:29f226ff0708294f0b96fe48e9f607b1616aa0ec1993c1fc530ce2fc134f42dc

Observation 6b767473-ee29-4f33-b9f8-cbfc86dd314d · outbound

This paper cites Bottom-up and top-down attention for image captioning and visual question answering,.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Bottom-up and top-down attention for image captioning and visual question answering,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.161338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.161338Z digest=sha256:71d11b54142122beba2188153c9680ca1cca99691b49ed0ceab59e99664f85cf

Observation 7b637dd1-c0a8-4628-b170-8131f064ef99 · outbound

This paper cites Swinbert: End-to-end transformers with sparse attention for video captioning,.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Swinbert: End-to-end transformers with sparse attention for video captioning,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.166132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.166132Z digest=sha256:d5bcc4119e0e6a7b718e2004dde9bfc2757d79a993b0341eeda3719a7b5e64a8

Observation 0d68d81b-31f0-4f7b-8883-95817f80ffb7 · outbound

This paper cites Vqa: Visual question answering,.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Vqa: Visual question answering,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.170653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.170653Z digest=sha256:9383875a56fbb2cd2924333ab990f4edff9803e90738b050a3235b06150a22de

Observation 80602856-cd1a-40fa-baf5-3d5d9bb7d2c7 · outbound

This paper cites Movieqa: Understanding stories in movies through question- answering,.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Movieqa: Understanding stories in movies through question- answering,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.175805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.175805Z digest=sha256:94210fc53dd9ea236af89bd97b4302971786be3a79420aa0fe6a6fa2df819db2

Observation 4d63aff3-6753-4266-84fc-e51fd504f63c · outbound

This paper cites Context-aware attention network for image-text retrieval,.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Context-aware attention network for image-text retrieval,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.181811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.181811Z digest=sha256:179ad15c244b507b85924cecfa3c49174e7f6d01fdd397a6ea83e171a63a0170

Observation a0348ba6-a5ba-4010-9c44-d6577e71f2fe · outbound

This paper cites Fine-grained video-text retrieval with hierarchical graph reasoning,.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Fine-grained video-text retrieval with hierarchical graph reasoning,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.186677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.186677Z digest=sha256:0640ff9f11dad7b4f1d3ea09aeed54088b2ed6d666ecc16ccb8b076e42f17ecf

Observation 62a87f97-baf3-4fd0-af55-356f5642acd2 · outbound

This paper cites Visual classification via description from large language models,.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Visual classification via description from large language models,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.191266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.191266Z digest=sha256:df924483cccfac791bfd99e50e2ad6fd7722552584f9d6fd7acadbc37caa2c76

Observation 8d00e512-bfd5-431c-b450-887cb665d8c2 · outbound

This paper cites Learning concise and descriptive attributes for visual recognition,.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Learning concise and descriptive attributes for visual recognition,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.195864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.195864Z digest=sha256:baff0a52e562445c35729dc1e30bbc9c86f9299719e0fd68083ea4ab675c96d0

Observation 7384dc51-6dad-4064-8d14-69096b248792 · outbound

This paper cites Exploring large language models for multi-modal out-of-distribution detection,.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Exploring large language models for multi-modal out-of-distribution detection,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.201342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.201342Z digest=sha256:f59486bd36c5fd6dcfc8e589081802d5cd8b76ae49e5384991d7fb17fee06565

Observation 37c035fb-ab3d-44eb-9a01-a2425d584e56 · outbound

This paper cites Chatgpt-powered hierarchical comparisons for image classification,.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Chatgpt-powered hierarchical comparisons for image classification,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.206793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.206793Z digest=sha256:618838b10983269ad826dae298a64cea8fb77a4fd0988c4e714fc8744419db83

Observation 113f8511-064d-4e5d-9e0d-58c165307694 · outbound

This paper cites Vision-language pre-training: Basics, recent advances, and future trends,.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Vision-language pre-training: Basics, recent advances, and future trends,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.211313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.211313Z digest=sha256:5e1d3694239553c60a7902facdec85e0775f6295896db7da877910f70b131770

Observation c6d701e4-6fe3-4ef0-ae53-625cbee3e4e7 · outbound

This paper cites Show and tell: A neural image caption generator,.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Show and tell: A neural image caption generator,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.216289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.216289Z digest=sha256:18665e5d0c946f653825972d3fb341bcbf6db9bd52d58c25059baf234e3425c9

Observation ccd0f2f0-41e3-4719-a7ab-d991b54b9756 · outbound

This paper cites Deep correlation for matching images and text,.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Deep correlation for matching images and text,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.220881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.220881Z digest=sha256:899504fa7a9cd398970878f31584466d30be79da82e9d1a97f3e3e751428a976

Observation 049e467a-40f8-4420-902e-054327aa82c9 · outbound

This paper cites Draw: A recurrent neural network for image generation,.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Draw: A recurrent neural network for image generation,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.226747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.226747Z digest=sha256:7c87beb77145157020e6f81b103144028a09cfd19c6122cfa79e62cba72f1604

Observation 02f00c2e-05b1-4615-bec1-e2e2787eb9e1 · outbound

This paper cites Memory- attended recurrent network for video captioning,.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Memory- attended recurrent network for video captioning,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.231254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.231254Z digest=sha256:b589415b4a528fac5ac3684dcfaa683d5056722efc9edb50db80381425ae88db

Observation 25b134f8-93c0-45cd-b0eb-da13c31ee43a · outbound

This paper cites Heterogeneous memory enhanced multimodal attention model for video question answering,.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Heterogeneous memory enhanced multimodal attention model for video question answering,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.235684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.235684Z digest=sha256:a1ee9294e1e65fdc949dff9e94c10760716db60c25a5bfb65278ba9d2d97a0d9

Observation 441a07d7-01f5-4ab6-ba82-411924c14c60 · outbound

This paper cites Mocogan: Decompos- ing motion and content for video generation,.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Mocogan: Decompos- ing motion and content for video generation,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.240343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.240343Z digest=sha256:a166776d2ec3c0babe12b0fc000effbefe79f2a94f8fec83e82a93ced59dfa90

Observation b03707cf-9f9d-4114-b41e-61f5a28a89bb · outbound

This paper cites Rethinking the bottom-up framework for query-based video localiza- tion,.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Rethinking the bottom-up framework for query-based video localiza- tion,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.246246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.246246Z digest=sha256:89a786dec3c9f79122d969b44d4890709237a198f2844b6b63032c946416c23e

Observation 6572cfa6-72b4-4359-8923-c15656023beb · outbound

This paper cites Object detection in 20 years: A survey,.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Object detection in 20 years: A survey,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.250908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.250908Z digest=sha256:839b4146381e49e450ae6701f9979d75e8461bba382a6ae2d39f37a9e7df5392

Observation ca14498b-3398-4a7a-a899-c85219ef154f · outbound

This paper cites Neural motifs: Scene graph parsing with global context,.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Neural motifs: Scene graph parsing with global context,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.255555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.255555Z digest=sha256:f343b87cb1616ba4779748ad0ff9f56cb2a51c06780f73de5ec520a022b2e3b9

Observation f7a5829d-9b4f-4852-9d32-3b905c77a6dd · outbound

This paper cites From recognition to cognition: Visual commonsense reasoning,.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey From recognition to cognition: Visual commonsense reasoning,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.260138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.260138Z digest=sha256:bc42e1be001edfba44b56e83ddf62cef11b19a58bee271dbe76d13ba188233ed

Observation a9cf5749-a14b-49cf-a287-d85fbfe0464a · outbound

This paper cites Visual Entailment: A Novel Task for Fine-Grained Image Understanding.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Visual Entailment: A Novel Task for Fine-Grained Image Understanding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.264684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.264684Z digest=sha256:3a7132bb59995779c93329be334b3c231a6ccde16f8798db3e3b66522faa41c2

Observation 8bf45281-434a-4669-bd68-c9686711b375 · outbound

This paper cites The abduction of sherlock holmes: A dataset for visual abductive reasoning,.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey The abduction of sherlock holmes: A dataset for visual abductive reasoning,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.269433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.269433Z digest=sha256:b8bf1c27156c36f8b4289e218c8b9805c973ad1aab52d50f1c766c718d4d5835

Observation 7ae55b5e-b01a-4e42-a129-6e59542c8514 · outbound

This paper cites Breaking common sense: Whoops! a vision-and-language benchmark of synthetic and compositional im- ages,.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Breaking common sense: Whoops! a vision-and-language benchmark of synthetic and compositional im- ages,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.273497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.273497Z digest=sha256:5328b8606970fe1bf533fbe545d53be5777e6d84f892f2fa58b2dc9eeb9b93e7

Observation 1044258a-ee6d-4523-9597-e09e7efe5f26 · outbound

This paper cites Let’s think outside the box: Exploring leap-of-thought in large language models with creative humor generation,.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Let’s think outside the box: Exploring leap-of-thought in large language models with creative humor generation,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.277674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.277674Z digest=sha256:09c9c516bc21b15bdf8aba0f54d921f0d0f34ba9acf3fdd78c00921f90916fc3

Observation f2c03e37-b057-4178-9e91-6bc2f01d31d9 · outbound

This paper cites Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing,.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.281912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.281912Z digest=sha256:c581cc1712bbf5f965a4843a997473350c6f35b293d13cea82edd728628a01ea

Observation 6b439e0a-614b-4537-9e9c-7ea086b7be0d · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey LLaMA: Open and Efficient Foundation Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.285863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.285863Z digest=sha256:70b086863d9042686d66daca9b28130d3c7c269608736b7c2192a4f2f6fce371

Observation 8d8fe502-67f9-4e77-95f4-8d041814d40e · outbound

This paper cites Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality,.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.290558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.290558Z digest=sha256:1aad6b60e12bfb71fe45798007c51c79223a9c0c7c347ae94f2f0ecfd334aa5e

Observation 57c302c3-26c4-4bfa-bfea-458cca150c6f · outbound

This paper cites Qwen Technical Report.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Qwen Technical Report

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.295187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.295187Z digest=sha256:7e3cf7f9a41eef9555339e46c10dd77de4679c3436d91da679aba9c90569b1f4

Observation 3bbaebfa-4c5a-4149-805d-5941bcb5e467 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Learning transferable visual models from natural language supervision,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.300089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.300089Z digest=sha256:04726c4dad6af9e607ce6eb884a4636c61af0ce12d80672cff60e4380a539afe

Observation 209d48e2-4ca2-4878-a91a-59c43472ebf0 · outbound

This paper cites Scaling up visual and vision-language representation learning with noisy text supervision,.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Scaling up visual and vision-language representation learning with noisy text supervision,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.305441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.305441Z digest=sha256:8e6ee6dfded96f6460f27c169f9d77bf99f4e8e9040f202497d73bd6ef001dd1

Observation 019bd6a3-336d-4a83-aef0-8412d20200a9 · outbound

This paper cites Vlmo: Unified vision-language pre- training with mixture-of-modality-experts,.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Vlmo: Unified vision-language pre- training with mixture-of-modality-experts,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.310428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.310428Z digest=sha256:e6b54cb0330c0d7e0cfd2854d4362915448f76d5989ecd667048c8f5a1ef00d0

Observation 7dc86bcb-35af-4f97-9c0a-d329449e0029 · outbound

This paper cites FILIP: Fine-grained Interactive Language-Image Pre-Training.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.315913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.315913Z digest=sha256:ea95fc3f64306de7bb6e543cc428d653604e01204c531a25291629efcae5e758

Observation 094d7387-96b0-4ee9-9e74-8735de932580 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation,.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.321522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.321522Z digest=sha256:5bd4eaddde45d8b787211e4eaf2dcf3749067bf2618c5eeefdea241deefae5c8

Observation 3a3b5abe-6d1b-4b02-8605-626e79e5c783 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.326673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.326673Z digest=sha256:f4232b4563ddb59f1d5cc0b704604d4f627b34f9a771d97a5360bc1dcbf7514d

Observation d57faeb6-6edd-4ab4-8c11-7641c2b71ad5 · outbound

This paper cites Visual instruction tuning,.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Visual instruction tuning,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.331796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.331796Z digest=sha256:6d91ce062e9048e5780cdea1e314e9d30fb3ed74ae510050bf74f9b1a72e714c

Observation 5af0e801-e2e4-4aa7-91c6-e70814e66f80 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.336578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.336578Z digest=sha256:120d32df41b92a5b3354116330ae8538be90f01b4357d5023e96a79bd20772cf

Observation d99eec56-064d-4396-975f-4191cec24c0a · outbound

This paper cites A Survey of Large Language Models.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey A Survey of Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.342321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.342321Z digest=sha256:710452022464926ce4f8eafbedbcb1d10f5e3888f5b6a69fb961a99a556cf1a6

Observation 1c683867-9394-43cb-9f99-3e4b42a8682e · outbound

This paper cites A Comprehensive Overview of Large Language Models.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey A Comprehensive Overview of Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.347750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.347750Z digest=sha256:b88debe75aa1cfdb01455213c948d773b9fb356ec901f1b3e7be200687b857a7

Observation f4e04d6d-f6ac-4623-9bd1-af9457097593 · outbound

This paper cites A survey on evaluation of large language models,.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey A survey on evaluation of large language models,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.352851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.352851Z digest=sha256:1a9f5db16497837e6aadd7ee1db57d8ac5b5684e5c9411d6250ca5574f4ddf67

Observation c9093796-6795-45a6-aea2-2e049f1d89bc · outbound

This paper cites Large-scale multi-modal pre-trained models: A com- prehensive survey,.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Large-scale multi-modal pre-trained models: A com- prehensive survey,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.357550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.357550Z digest=sha256:01e74faca38a11e13b66eb9b648793b85ed7718e2f536325a4809ad848e46d8b

Observation 2f77d8d7-9b74-4845-a7fb-b5b0cd345367 · outbound

This paper cites Foundational Models Defining a New Era in Vision: A Survey and Outlook.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Foundational Models Defining a New Era in Vision: A Survey and Outlook

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.362248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.362248Z digest=sha256:9c7ba5632beadf8735dd356667d92217aafeeb8f80e8fede791decf9897ac132

Observation 937aaa72-68e8-4300-bb23-e48d3e8bfdf0 · outbound

This paper cites A Survey on Multimodal Large Language Models.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey A Survey on Multimodal Large Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.367070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.367070Z digest=sha256:ab9f98799e150f1f2880bbca11d68daf799373d0a606803b426bca07d7ce6f10

Observation f53a411d-e9ab-4795-97cd-e5dd1d4dd464 · outbound

This paper cites MM-LLMs: Recent Advances in MultiModal Large Language Models.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey MM-LLMs: Recent Advances in MultiModal Large Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.372006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.372006Z digest=sha256:3d977b0d04f5d25d0814c0859496b7ff93e970e456699c9f1aefdf4c35e010dd

Observation cf56b006-35a9-45ed-8df3-63d30f45dcc5 · outbound

This paper cites Exploring the Reasoning Abilities of Multimodal Large Language Models (MLLMs): A Comprehensive Survey on Emerging Trends in Multimodal Reasoning.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Exploring the Reasoning Abilities of Multimodal Large Language Models (MLLMs): A Comprehensive Survey on Emerging Trends in Multimodal Reasoning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.376675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.376675Z digest=sha256:25b9284602d1c6efbbddba0249be19ee5a170ba4f78ecf441eb9bf6b66f32eb1

Observation f911f90b-f4e9-483d-856f-4c2bc6699771 · outbound

This paper cites Video understanding with large language models: A survey,.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Video understanding with large language models: A survey,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.381148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.381148Z digest=sha256:37941fd94e8a1d264a939abae4ca9de485dc36e1a05770b17bf7408e820ea1f4

Observation f2612d6b-e681-4543-8935-783877968ab5 · outbound

This paper cites Vision-language models for vision tasks: A survey,.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Vision-language models for vision tasks: A survey,

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.385442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.385442Z digest=sha256:bd93a5a670fa3bbf65d4ac762e9bb01999ad93f73bf3983a221980c18decdd2c

Observation 0a3844b0-08ef-4666-96a1-3a72e29532e6 · outbound

This paper cites LLMs Meet Multimodal Generation and Editing: A Survey.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey LLMs Meet Multimodal Generation and Editing: A Survey

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.389662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.389662Z digest=sha256:7026ac5631171818c49f232fdb9f2c283dd09ea52f80f91ab95ca5ffb573b794

Observation e47f553e-4c25-427d-ba29-e9b5d4b483a4 · outbound

This paper cites Generalized Out-of-Distribution Detection and Beyond in Vision Language Model Era: A Survey.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Generalized Out-of-Distribution Detection and Beyond in Vision Language Model Era: A Survey

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.394298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.394298Z digest=sha256:891c52e1cd97ee48c9f189553aa6e5df36e689945ad7bcb4883d2c987926ba58

Observation 7f8641d0-b904-46d4-9278-50960230ec98 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.400421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.400421Z digest=sha256:3e8309805201661ebfb5fcb813ec5298c365bbed44c74f30edad944dd0478206

Observation 375c1fd8-4966-42f4-baab-4334bda14734 · outbound

This paper cites ERNIE: Enhanced Representation through Knowledge Integration.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey ERNIE: Enhanced Representation through Knowledge Integration

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.405469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.405469Z digest=sha256:df6e2c650cc7bfcadac507a34852a529fe7e4b9fa6dc46015620d653cead04cb

Observation c67356c9-fef1-474b-b6cd-a9eee42e6c1a · outbound

This paper cites Language mod- els are few-shot learners,.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Language mod- els are few-shot learners,

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.410963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.410963Z digest=sha256:0b5ac32816af6379b2a1fe62619db78eff1e2605b741c68f7436aab8416d3050

Observation 621ba948-3326-426d-b485-20f283a3945a · outbound

This paper cites Image Captioning with Very Scarce Supervised Data: Adversarial Semi-Supervised Learning Approach.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Image Captioning with Very Scarce Supervised Data: Adversarial Semi-Supervised Learning Approach

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-08-11T18:11:56.657530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T18:11:54.416016Z digest=sha256:448bc34a7301dead65383cc2639372ccd3f1d7708d0db5fd232449ba97fd2851

Observation 3c8eb5e4-e7d0-4d0a-b407-e960894fd2f2 · outbound

This paper cites Semi-supervised cross-modal retrieval with label prediction,.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Semi-supervised cross-modal retrieval with label prediction,

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.421123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.421123Z digest=sha256:5ae3b24838decca1ea541d99c2ba2ab7cba19a90d8dc7d7309fe3e6d541ca8c1

Observation b693d6ff-cbd0-4b9c-9a8e-d07f25b26752 · outbound

This paper cites Weakly supervised dense event captioning in videos,.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Weakly supervised dense event captioning in videos,

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.425724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.425724Z digest=sha256:37d5e27893ce250a2af416c08a621f78c6a6d43af3f6cdcc13935e2f596d761d

Observation 827313f4-0b87-447e-ab64-b9be562b090f · outbound

This paper cites Weakly-Supervised Visual-Retriever-Reader for Knowledge-based Question Answering.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Weakly-Supervised Visual-Retriever-Reader for Knowledge-based Question Answering

Reference 58

Resolution
verified exact
local_arxiv, observed 2026-08-11T18:11:56.635511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T18:11:54.430275Z digest=sha256:4ca457ca3b20f9dda0a766fdb8ce66a778aef4edfe8d798c190daa569b1921e4

Observation 76650929-8efa-47c4-bfce-76f3bac47aa4 · outbound

This paper cites Unsupervised image captioning,.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Unsupervised image captioning,

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.435131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.435131Z digest=sha256:d9c7674f134148acb44606c7bde2f7b0d36bcd56eb44adb7a93523edac37c08c

Observation d0adeb38-a699-458e-9157-311fff23ffc9 · outbound

This paper cites Towards unsupervised image captioning with shared multimodal embeddings,.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Towards unsupervised image captioning with shared multimodal embeddings,

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.439820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.439820Z digest=sha256:f894d5ae97e8db37f61fdc042bae4443885bdecf16ca0acc7b54a066178e0236

Observation 1bc79c32-a43a-4665-926d-0ea5ddf24a4a · outbound

This paper cites Zerocap: Zero-shot image-to-text generation for visual-semantic arithmetic,.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Zerocap: Zero-shot image-to-text generation for visual-semantic arithmetic,

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.444320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.444320Z digest=sha256:3bf06aaf7050ba71832d7dca0d7113fcf8c27db5c44ecd7cb97d21a0e04b3bfe

Observation 1de9018c-e31b-4f18-9c49-fa2fd8f7f403 · outbound

This paper cites Language Models Can See: Plugging Visual Controls in Text Generation.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Language Models Can See: Plugging Visual Controls in Text Generation

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.449739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.449739Z digest=sha256:68593702d61d27225d02dc7f763a72e08b984904ff3ce7fd90a17055ac97578f

Observation 06928aad-7ea1-4ccd-9ff3-24df7dbc80c4 · outbound

This paper cites Conzic: Controllable zero-shot image captioning by sampling-based polishing,.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Conzic: Controllable zero-shot image captioning by sampling-based polishing,

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.454925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.454925Z digest=sha256:639294ecb6df5c547e9c82a5cf113db5061bc5bdc503d24e7151f908d4900e43

Observation c4fa8a1e-202f-488d-9e21-1d809851ea6d · outbound

This paper cites MeaCap: Memory-Augmented Zero-shot Image Captioning.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey MeaCap: Memory-Augmented Zero-shot Image Captioning

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.459435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.459435Z digest=sha256:600e458297dec73946441ab9f3234a86ba818d26a7f2e09733987be8602f84e1

Observation f79e658c-c6a3-4275-b3c7-d2721a6cc561 · outbound

This paper cites Text-only training for visual storytelling,.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Text-only training for visual storytelling,

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.464285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.464285Z digest=sha256:13b4f9656932878a568ad7527b97b9e8477fe3f11be2664ed2cba557c9a80ad4

Observation 50cf5717-8a82-4a0f-a8c2-d066e766c04d · outbound

This paper cites Zero-Shot Video Captioning with Evolving Pseudo-Tokens.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Zero-Shot Video Captioning with Evolving Pseudo-Tokens

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.468557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.468557Z digest=sha256:e38143a0981a9a70ae5a7a95bb7e1907aa00ab87395675330dd53d1f4b3025e0

Observation ade0ec7a-597a-4360-b8f9-d52de4825d53 · outbound

This paper cites Zero-Shot Dense Video Captioning by Jointly Optimizing Text and Moment.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Zero-Shot Dense Video Captioning by Jointly Optimizing Text and Moment

Reference 67

Resolution
verified exact
local_arxiv, observed 2026-08-11T18:11:56.564086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T18:11:54.473486Z digest=sha256:e9c3ab4f0531e90f5f25cfe405828c1db915c00630b864d1f0e12dc09b43f923

Observation 33f3847e-a73c-4d95-b38b-ffb1e24975a7 · outbound

This paper cites CLIP Models are Few-shot Learners: Empirical Studies on VQA and Visual Entailment.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey CLIP Models are Few-shot Learners: Empirical Studies on VQA and Visual Entailment

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.478596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.478596Z digest=sha256:fd068f698a1a2a2279bbbbc0ea7db1933011426650170f4bbb155d376ec200b6

Observation 91445399-f281-4bb0-aaa1-2d80af804f73 · outbound

This paper cites Towards counterfactual image manipulation via clip,.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Towards counterfactual image manipulation via clip,

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.482923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.482923Z digest=sha256:545565e4282d2ab3a6c98adaab85af0eb51724b82ab2b16bea1b7fb81629d754

Observation 465e523a-e738-4514-8366-aefe48971188 · outbound

This paper cites An empirical study of gpt-3 for few-shot knowledge-based vqa,.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey An empirical study of gpt-3 for few-shot knowledge-based vqa,

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.487500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.487500Z digest=sha256:acea076710eb911c2e3bf6beecb7822725f5db449728ebbd583e7513dc50a317

Observation 525e4b92-0071-41d0-89d7-dcc4583a17f6 · outbound

This paper cites Plug-and-Play VQA: Zero-shot VQA by Conjoining Large Pretrained Models with Zero Training.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Plug-and-Play VQA: Zero-shot VQA by Conjoining Large Pretrained Models with Zero Training

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.491641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.491641Z digest=sha256:36fb0e7ab949b96a0f403e18e59cabfddcfb47a6d7a9485dcc10e57072824207

Observation 60c89bcf-60fc-4346-b080-90997eed73af · outbound

This paper cites Language models with image descriptors are strong few-shot video-language learners,.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Language models with image descriptors are strong few-shot video-language learners,

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.496115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.496115Z digest=sha256:f9cbae5fb4638eb6b69033fa2ec27bbdeca3c8fe49e8e1d9532798f693813153

Observation db0487ce-303f-4a8d-b77b-6d2977114b46 · outbound

This paper cites Language as the Medium: Multimodal Video Classification through text only.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Language as the Medium: Multimodal Video Classification through text only

Reference 73

Resolution
verified exact
local_arxiv, observed 2026-08-11T18:11:56.508315Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T18:11:54.500386Z digest=sha256:862385efae68d2371ef79d6cc82fe102f8dd2e5694e66bdc7d4354532de512d9

Observation 01040397-161d-470d-b440-3202c125e6b5 · outbound

This paper cites A Video Is Worth 4096 Tokens: Verbalize Videos To Understand Them In Zero Shot.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey A Video Is Worth 4096 Tokens: Verbalize Videos To Understand Them In Zero Shot

Reference 74

Resolution
verified exact
local_arxiv, observed 2026-08-11T18:11:56.486309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T18:11:54.506310Z digest=sha256:f9efe8fb24eb5c26ad4456a0baf715ea2ae59e38aa06256438ec7cb6d3ced2bf

Observation a32b45dd-6a49-4ecb-8d6a-4a7ac9906df4 · outbound

This paper cites Retrieving-to-Answer: Zero-Shot Video Question Answering with Frozen Large Language Models.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Retrieving-to-Answer: Zero-Shot Video Question Answering with Frozen Large Language Models

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.511203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.511203Z digest=sha256:1e3b00ba74b0567e40faa14c76ef1af8f64ef0d70391108c12f3340154ea7ff1

Observation 295e36b9-0db3-4fc6-8ef7-a83019bf49f5 · outbound

This paper cites Text-Only Training for Image Captioning using Noise-Injected CLIP.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Text-Only Training for Image Captioning using Noise-Injected CLIP

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.516711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.516711Z digest=sha256:94a15f4ecf3d613dc0e1a8a52b7d3edda049dab8e39ff55bf82d819f78205bb1

Observation a7fc2b48-b02b-4eb6-aade-2f88c4a4e2cc · outbound

This paper cites I Can't Believe There's No Images! Learning Visual Tasks Using only Language Supervision.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey I Can't Believe There's No Images! Learning Visual Tasks Using only Language Supervision

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.521893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.521893Z digest=sha256:d20be141359a14f7d653d52b346eeba87bf774a37b1e34ba3da7a17f436e583a

Observation a6ef1573-ff90-4843-9aa7-833e60eea867 · outbound

This paper cites DeCap: Decoding CLIP Latents for Zero-Shot Captioning via Text-Only Training.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey DeCap: Decoding CLIP Latents for Zero-Shot Captioning via Text-Only Training

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.526672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.526672Z digest=sha256:6b50d76ebda8b3d5f770c3338435805bcd1a7e310daf929df2bc333ddc35f7a1

Observation 5564bba4-21b8-4a19-9800-9fd834cb580c · outbound

This paper cites From Association to Generation: Text-only Captioning by Unsupervised Cross-modal Mapping.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey From Association to Generation: Text-only Captioning by Unsupervised Cross-modal Mapping

Reference 79

Resolution
verified exact
local_arxiv, observed 2026-08-11T18:11:56.400451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T18:11:54.531612Z digest=sha256:b1e737d6a17b182332e1bace66ad941f70f930b375947334bd514b3c1720140d

Observation e8d35d51-2e52-48f2-9935-f296ab10917f · outbound

This paper cites Zero-shot Image Captioning by Anchor-augmented Vision-Language Space Alignment.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Zero-shot Image Captioning by Anchor-augmented Vision-Language Space Alignment

Reference 80

Resolution
verified exact
local_arxiv, observed 2026-08-11T18:11:56.377916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T18:11:54.536179Z digest=sha256:c6e5042d5a5f6a0bbf44e7e112dbcf067c6994ccc0ed7262eb9574dcf63616c6

Observation 21e5e519-d460-4f18-83e6-26aa7154c4f2 · outbound

This paper cites Transferable decoding with visual entities for zero-shot image captioning,.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Transferable decoding with visual entities for zero-shot image captioning,

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.540954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.540954Z digest=sha256:de6be9614f53ee7f192d8c0886b8dfac6f427b37fe8dc7da3823e5c909b49b85

Observation 1dde8750-f1fc-4f1c-b7dc-89205ee3b4f7 · outbound

This paper cites Language-free training for zero-shot video grounding,.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Language-free training for zero-shot video grounding,

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.545994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.545994Z digest=sha256:1a3b73c237481b820139299cfe6bc2fe0b7fde2f1e770b572e03d93b533e8331

Observation 7b895ac4-d2b2-4224-9f9e-fff8f3f4666c · outbound

This paper cites CLIP-GEN: Language-Free Training of a Text-to-Image Generator with CLIP.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey CLIP-GEN: Language-Free Training of a Text-to-Image Generator with CLIP

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.551267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.551267Z digest=sha256:b6a813a68d39a4ce11e1ce3c2c5a6262035a0280c896f938290ffb2d133d460c

Observation 14f7226d-3ea3-4178-97d3-a69d0bcd57cb · outbound

This paper cites Image captioning with multi-context synthetic data,.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Image captioning with multi-context synthetic data,

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.556270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.556270Z digest=sha256:1703fbb1d0bbc1fded6519fb7ae28a168e54da31d935cac2861bf4f0e7f5ebdf

Observation 551c90a3-e164-4757-bce3-a4d6b755947b · outbound

This paper cites Improving cross-modal alignment with synthetic pairs for text-only image captioning,.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Improving cross-modal alignment with synthetic pairs for text-only image captioning,

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.560774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.560774Z digest=sha256:53db6a67f23e4ae3ecfb3fa1369e6aeae7c48c7b083e9f5b51624af10b35c7fe

Observation cb24d4eb-5337-48b7-97d5-be80ac94b343 · outbound

This paper cites FreeMask: Synthetic Images with Dense Annotations Make Stronger Segmentation Models.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey FreeMask: Synthetic Images with Dense Annotations Make Stronger Segmentation Models

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.565504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.565504Z digest=sha256:31009bf02297f735f7d8f6863196f3379b4b26fb5aece79132fb73bb081af1fc

Observation 72c2b9e4-9c40-43c3-b8f6-f4e782951e39 · outbound

This paper cites Towards language-free training for text-to-image generation,.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Towards language-free training for text-to-image generation,

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.570501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.570501Z digest=sha256:45a19e3c20007d8c39e81d5bd2c80a6aaf9decb67f512fdd1a6eaa64c382488a

Observation 1e7d941f-6c12-4b9e-90a4-30d6526632be · outbound

This paper cites See, Think, Confirm: Interactive Prompting Between Vision and Language Models for Knowledge-based Visual Reasoning.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey See, Think, Confirm: Interactive Prompting Between Vision and Language Models for Knowledge-based Visual Reasoning

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.574723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.574723Z digest=sha256:f236980b9c163c809b594e2c34b069ed2fdc75a222f5174cc50131fa5502c24c

Observation a2289860-351f-43ba-aac1-583212c2a10b · outbound

This paper cites ViCor: Bridging Visual Understanding and Commonsense Reasoning with Large Language Models.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey ViCor: Bridging Visual Understanding and Commonsense Reasoning with Large Language Models

Reference 89

Resolution
verified exact
local_arxiv, observed 2026-08-11T18:11:56.301186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T18:11:54.578985Z digest=sha256:159d19612576134f380c8cf3e569235f86f47e4ff869fa7e845584a3c2fab3df

Observation 098f0d44-7b83-40e2-8474-5d52c8782d03 · outbound

This paper cites DOMINO: A Dual-System for Multi-step Visual Language Reasoning.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey DOMINO: A Dual-System for Multi-step Visual Language Reasoning

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.583677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.583677Z digest=sha256:4b4864480897a7e7ee489665e22bc8a2f1c0ecfdd61895722cadaf32b31b49ba

Observation 37d12aea-6eab-41bc-8174-cfae4e241446 · outbound

This paper cites IdealGPT: Iteratively decomposing vision and language reasoning via large language models,.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey IdealGPT: Iteratively decomposing vision and language reasoning via large language models,

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.588732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.588732Z digest=sha256:d2601dfd25280146ba023d7c57ad6ac25fec2b3ed5993589d90e480c690621e9

Observation 9670f636-c851-4df2-a8c3-24a601df6408 · outbound

This paper cites Good Questions Help Zero-Shot Image Reasoning.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Good Questions Help Zero-Shot Image Reasoning

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.593526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.593526Z digest=sha256:d7bb88a8d7513b9fcd570461fa683a15ab62ccdcc0da11cd1c236a841ad18b36

Observation 4e8bd70e-00ea-4de7-b212-7bafd3efbe37 · outbound

This paper cites The art of SOCRATIC QUESTIONING: Recursive thinking with large language models,.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey The art of SOCRATIC QUESTIONING: Recursive thinking with large language models,

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.598411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.598411Z digest=sha256:20562fe9b0430af5c44e1ef3fbc78f9c47d2df6fdb5528deea39150a4667cb47

Observation 2b38e2a4-fbac-4dca-9f5e-b4890fd7db45 · outbound

This paper cites Filling the image information gap for VQA: Prompting large language models to proactively ask questions,.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Filling the image information gap for VQA: Prompting large language models to proactively ask questions,

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.602828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.602828Z digest=sha256:f5594d15d272346eb5a09740ffaf296b5e36a61b92fbd8458e039308bbdbaa3d

Observation 8fb3cfa1-2d26-4465-b9ad-035fa9fb644c · outbound

This paper cites Multimodal Multi-Hop Question Answering Through a Conversation Between Tools and Efficiently Finetuned Large Language Models.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Multimodal Multi-Hop Question Answering Through a Conversation Between Tools and Efficiently Finetuned Large Language Models

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.607488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.607488Z digest=sha256:16acd2a5f036c3e0d9b304890097ab0b3c4217d73a6247e66f385d9df75f808b

Observation f6e5080b-05b7-40a7-9c26-0962906a8d4a · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering,.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Learn to explain: Multimodal reasoning via thought chains for science question answering,

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.612605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.612605Z digest=sha256:ac089e390026186850f450acd4027ee15a699452cb5b662c112249edcc2e7569

Observation 8dd6dd2c-66ee-4dad-b0b4-aad6157d1847 · outbound

This paper cites Multimodal Chain-of-Thought Reasoning in Language Models.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Multimodal Chain-of-Thought Reasoning in Language Models

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.616979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.616979Z digest=sha256:d79af7ce0c584d43371cb60ad12f2daf80063ff428a88eceb0439ef74eca0952

Observation bd57b404-b707-4b78-8729-20a2c0ed20ea · outbound

This paper cites T-sciq: Teaching multimodal chain-of-thought reasoning via large language model signals for science question answering,.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey T-sciq: Teaching multimodal chain-of-thought reasoning via large language model signals for science question answering,

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.622299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.622299Z digest=sha256:60a81b31553b07798402feee9fc2b773c47061a4b5625d888bb6e2e6abe31a8c

Observation 170d0af2-a920-4136-be48-68ded2fa38b2 · outbound

This paper cites KAM-CoT: Knowledge Augmented Multimodal Chain-of-Thoughts Reasoning.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey KAM-CoT: Knowledge Augmented Multimodal Chain-of-Thoughts Reasoning

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.627119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.627119Z digest=sha256:6f787f90fad9e92c4a19026d243ca109533922fa4ffadb4979f19c633293e6f9

Observation c4cdba88-c0cf-4eb1-a42a-def771209950 · outbound

This paper cites Measuring and Improving Chain-of-Thought Reasoning in Vision-Language Models.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Measuring and Improving Chain-of-Thought Reasoning in Vision-Language Models

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.632569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.632569Z digest=sha256:a963e4f593fb32c84ccbfe0d73c37c6e7440068eaaee3535ef25b793f8b01581

Pith citing papers

No inbound Pith citation observations are available.