Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T18:11:54.632569Z
Paper Citation Record · LEDGER
As of 12 August 2026, this Paper Citation Record lists 100 of 229 outbound references and 0 inbound Pith citation observations for arXiv:2412.08158.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T18:11:54.632569Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
100 of 229 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 465a3771-17b5-4231-846f-1cc1323618ff · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Multimodal research in vision and language: A review of current and emerging trends,
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e66d460d-6aa1-481a-a634-6b91c7030011 · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Vision+ language applications: A survey,
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b767473-ee29-4f33-b9f8-cbfc86dd314d · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Bottom-up and top-down attention for image captioning and visual question answering,
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b637dd1-c0a8-4628-b170-8131f064ef99 · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Swinbert: End-to-end transformers with sparse attention for video captioning,
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d68d81b-31f0-4f7b-8883-95817f80ffb7 · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Vqa: Visual question answering,
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80602856-cd1a-40fa-baf5-3d5d9bb7d2c7 · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Movieqa: Understanding stories in movies through question- answering,
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d63aff3-6753-4266-84fc-e51fd504f63c · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Context-aware attention network for image-text retrieval,
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0348ba6-a5ba-4010-9c44-d6577e71f2fe · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Fine-grained video-text retrieval with hierarchical graph reasoning,
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62a87f97-baf3-4fd0-af55-356f5642acd2 · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Visual classification via description from large language models,
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d00e512-bfd5-431c-b450-887cb665d8c2 · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Learning concise and descriptive attributes for visual recognition,
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7384dc51-6dad-4064-8d14-69096b248792 · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Exploring large language models for multi-modal out-of-distribution detection,
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37c035fb-ab3d-44eb-9a01-a2425d584e56 · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Chatgpt-powered hierarchical comparisons for image classification,
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 113f8511-064d-4e5d-9e0d-58c165307694 · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Vision-language pre-training: Basics, recent advances, and future trends,
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6d701e4-6fe3-4ef0-ae53-625cbee3e4e7 · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Show and tell: A neural image caption generator,
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ccd0f2f0-41e3-4719-a7ab-d991b54b9756 · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Deep correlation for matching images and text,
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 049e467a-40f8-4420-902e-054327aa82c9 · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Draw: A recurrent neural network for image generation,
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02f00c2e-05b1-4615-bec1-e2e2787eb9e1 · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Memory- attended recurrent network for video captioning,
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25b134f8-93c0-45cd-b0eb-da13c31ee43a · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Heterogeneous memory enhanced multimodal attention model for video question answering,
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 441a07d7-01f5-4ab6-ba82-411924c14c60 · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Mocogan: Decompos- ing motion and content for video generation,
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b03707cf-9f9d-4114-b41e-61f5a28a89bb · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Rethinking the bottom-up framework for query-based video localiza- tion,
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6572cfa6-72b4-4359-8923-c15656023beb · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Object detection in 20 years: A survey,
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca14498b-3398-4a7a-a899-c85219ef154f · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Neural motifs: Scene graph parsing with global context,
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7a5829d-9b4f-4852-9d32-3b905c77a6dd · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey From recognition to cognition: Visual commonsense reasoning,
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9cf5749-a14b-49cf-a287-d85fbfe0464a · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Visual Entailment: A Novel Task for Fine-Grained Image Understanding
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8bf45281-434a-4669-bd68-c9686711b375 · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey The abduction of sherlock holmes: A dataset for visual abductive reasoning,
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ae55b5e-b01a-4e42-a129-6e59542c8514 · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Breaking common sense: Whoops! a vision-and-language benchmark of synthetic and compositional im- ages,
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1044258a-ee6d-4523-9597-e09e7efe5f26 · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Let’s think outside the box: Exploring leap-of-thought in large language models with creative humor generation,
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2c03e37-b057-4178-9e91-6bc2f01d31d9 · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing,
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b439e0a-614b-4537-9e9c-7ea086b7be0d · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey LLaMA: Open and Efficient Foundation Language Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d8fe502-67f9-4e77-95f4-8d041814d40e · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality,
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57c302c3-26c4-4bfa-bfea-458cca150c6f · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Qwen Technical Report
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3bbaebfa-4c5a-4149-805d-5941bcb5e467 · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Learning transferable visual models from natural language supervision,
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 209d48e2-4ca2-4878-a91a-59c43472ebf0 · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Scaling up visual and vision-language representation learning with noisy text supervision,
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 019bd6a3-336d-4a83-aef0-8412d20200a9 · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Vlmo: Unified vision-language pre- training with mixture-of-modality-experts,
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7dc86bcb-35af-4f97-9c0a-d329449e0029 · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 094d7387-96b0-4ee9-9e74-8735de932580 · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation,
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a3b5abe-6d1b-4b02-8605-626e79e5c783 · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d57faeb6-6edd-4ab4-8c11-7641c2b71ad5 · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Visual instruction tuning,
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5af0e801-e2e4-4aa7-91c6-e70814e66f80 · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d99eec56-064d-4396-975f-4191cec24c0a · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey A Survey of Large Language Models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c683867-9394-43cb-9f99-3e4b42a8682e · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey A Comprehensive Overview of Large Language Models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4e04d6d-f6ac-4623-9bd1-af9457097593 · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey A survey on evaluation of large language models,
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9093796-6795-45a6-aea2-2e049f1d89bc · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Large-scale multi-modal pre-trained models: A com- prehensive survey,
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f77d8d7-9b74-4845-a7fb-b5b0cd345367 · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Foundational Models Defining a New Era in Vision: A Survey and Outlook
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 937aaa72-68e8-4300-bb23-e48d3e8bfdf0 · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey A Survey on Multimodal Large Language Models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f53a411d-e9ab-4795-97cd-e5dd1d4dd464 · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey MM-LLMs: Recent Advances in MultiModal Large Language Models
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf56b006-35a9-45ed-8df3-63d30f45dcc5 · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Exploring the Reasoning Abilities of Multimodal Large Language Models (MLLMs): A Comprehensive Survey on Emerging Trends in Multimodal Reasoning
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f911f90b-f4e9-483d-856f-4c2bc6699771 · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Video understanding with large language models: A survey,
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2612d6b-e681-4543-8935-783877968ab5 · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Vision-language models for vision tasks: A survey,
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a3844b0-08ef-4666-96a1-3a72e29532e6 · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey LLMs Meet Multimodal Generation and Editing: A Survey
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e47f553e-4c25-427d-ba29-e9b5d4b483a4 · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Generalized Out-of-Distribution Detection and Beyond in Vision Language Model Era: A Survey
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f8641d0-b904-46d4-9278-50960230ec98 · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 375c1fd8-4966-42f4-baab-4334bda14734 · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey ERNIE: Enhanced Representation through Knowledge Integration
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c67356c9-fef1-474b-b6cd-a9eee42e6c1a · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Language mod- els are few-shot learners,
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 621ba948-3326-426d-b485-20f283a3945a · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Image Captioning with Very Scarce Supervised Data: Adversarial Semi-Supervised Learning Approach
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 3c8eb5e4-e7d0-4d0a-b407-e960894fd2f2 · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Semi-supervised cross-modal retrieval with label prediction,
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b693d6ff-cbd0-4b9c-9a8e-d07f25b26752 · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Weakly supervised dense event captioning in videos,
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 827313f4-0b87-447e-ab64-b9be562b090f · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Weakly-Supervised Visual-Retriever-Reader for Knowledge-based Question Answering
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 76650929-8efa-47c4-bfce-76f3bac47aa4 · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Unsupervised image captioning,
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0adeb38-a699-458e-9157-311fff23ffc9 · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Towards unsupervised image captioning with shared multimodal embeddings,
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1bc79c32-a43a-4665-926d-0ea5ddf24a4a · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Zerocap: Zero-shot image-to-text generation for visual-semantic arithmetic,
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1de9018c-e31b-4f18-9c49-fa2fd8f7f403 · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Language Models Can See: Plugging Visual Controls in Text Generation
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06928aad-7ea1-4ccd-9ff3-24df7dbc80c4 · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Conzic: Controllable zero-shot image captioning by sampling-based polishing,
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4fa8a1e-202f-488d-9e21-1d809851ea6d · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey MeaCap: Memory-Augmented Zero-shot Image Captioning
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f79e658c-c6a3-4275-b3c7-d2721a6cc561 · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Text-only training for visual storytelling,
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50cf5717-8a82-4a0f-a8c2-d066e766c04d · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Zero-Shot Video Captioning with Evolving Pseudo-Tokens
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ade0ec7a-597a-4360-b8f9-d52de4825d53 · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Zero-Shot Dense Video Captioning by Jointly Optimizing Text and Moment
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 33f3847e-a73c-4d95-b38b-ffb1e24975a7 · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey CLIP Models are Few-shot Learners: Empirical Studies on VQA and Visual Entailment
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91445399-f281-4bb0-aaa1-2d80af804f73 · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Towards counterfactual image manipulation via clip,
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 465e523a-e738-4514-8366-aefe48971188 · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey An empirical study of gpt-3 for few-shot knowledge-based vqa,
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 525e4b92-0071-41d0-89d7-dcc4583a17f6 · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Plug-and-Play VQA: Zero-shot VQA by Conjoining Large Pretrained Models with Zero Training
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60c89bcf-60fc-4346-b080-90997eed73af · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Language models with image descriptors are strong few-shot video-language learners,
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db0487ce-303f-4a8d-b77b-6d2977114b46 · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Language as the Medium: Multimodal Video Classification through text only
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 01040397-161d-470d-b440-3202c125e6b5 · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey A Video Is Worth 4096 Tokens: Verbalize Videos To Understand Them In Zero Shot
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation a32b45dd-6a49-4ecb-8d6a-4a7ac9906df4 · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Retrieving-to-Answer: Zero-Shot Video Question Answering with Frozen Large Language Models
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 295e36b9-0db3-4fc6-8ef7-a83019bf49f5 · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Text-Only Training for Image Captioning using Noise-Injected CLIP
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7fc2b48-b02b-4eb6-aade-2f88c4a4e2cc · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey I Can't Believe There's No Images! Learning Visual Tasks Using only Language Supervision
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6ef1573-ff90-4843-9aa7-833e60eea867 · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey DeCap: Decoding CLIP Latents for Zero-Shot Captioning via Text-Only Training
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5564bba4-21b8-4a19-9800-9fd834cb580c · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey From Association to Generation: Text-only Captioning by Unsupervised Cross-modal Mapping
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e8d35d51-2e52-48f2-9935-f296ab10917f · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Zero-shot Image Captioning by Anchor-augmented Vision-Language Space Alignment
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 21e5e519-d460-4f18-83e6-26aa7154c4f2 · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Transferable decoding with visual entities for zero-shot image captioning,
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1dde8750-f1fc-4f1c-b7dc-89205ee3b4f7 · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Language-free training for zero-shot video grounding,
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b895ac4-d2b2-4224-9f9e-fff8f3f4666c · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey CLIP-GEN: Language-Free Training of a Text-to-Image Generator with CLIP
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14f7226d-3ea3-4178-97d3-a69d0bcd57cb · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Image captioning with multi-context synthetic data,
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 551c90a3-e164-4757-bce3-a4d6b755947b · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Improving cross-modal alignment with synthetic pairs for text-only image captioning,
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb24d4eb-5337-48b7-97d5-be80ac94b343 · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey FreeMask: Synthetic Images with Dense Annotations Make Stronger Segmentation Models
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72c2b9e4-9c40-43c3-b8f6-f4e782951e39 · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Towards language-free training for text-to-image generation,
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e7d941f-6c12-4b9e-90a4-30d6526632be · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey See, Think, Confirm: Interactive Prompting Between Vision and Language Models for Knowledge-based Visual Reasoning
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2289860-351f-43ba-aac1-583212c2a10b · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey ViCor: Bridging Visual Understanding and Commonsense Reasoning with Large Language Models
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 098f0d44-7b83-40e2-8474-5d52c8782d03 · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey DOMINO: A Dual-System for Multi-step Visual Language Reasoning
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37d12aea-6eab-41bc-8174-cfae4e241446 · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey IdealGPT: Iteratively decomposing vision and language reasoning via large language models,
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9670f636-c851-4df2-a8c3-24a601df6408 · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Good Questions Help Zero-Shot Image Reasoning
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e8bd70e-00ea-4de7-b212-7bafd3efbe37 · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey The art of SOCRATIC QUESTIONING: Recursive thinking with large language models,
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b38e2a4-fbac-4dca-9f5e-b4890fd7db45 · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Filling the image information gap for VQA: Prompting large language models to proactively ask questions,
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8fb3cfa1-2d26-4465-b9ad-035fa9fb644c · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Multimodal Multi-Hop Question Answering Through a Conversation Between Tools and Efficiently Finetuned Large Language Models
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6e5080b-05b7-40a7-9c26-0962906a8d4a · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Learn to explain: Multimodal reasoning via thought chains for science question answering,
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8dd6dd2c-66ee-4dad-b0b4-aad6157d1847 · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Multimodal Chain-of-Thought Reasoning in Language Models
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd57b404-b707-4b78-8729-20a2c0ed20ea · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey T-sciq: Teaching multimodal chain-of-thought reasoning via large language model signals for science question answering,
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 170d0af2-a920-4136-be48-68ded2fa38b2 · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey KAM-CoT: Knowledge Augmented Multimodal Chain-of-Thoughts Reasoning
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4cdba88-c0cf-4eb1-a42a-def771209950 · outbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey Measuring and Improving Chain-of-Thought Reasoning in Vision-Language Models
Reference 100
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.