Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-15T15:46:06.334088Z
Paper Citation Record · LEDGER
As of 12 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 100 inbound Pith citation observations for arXiv:2311.03079.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-15T15:46:06.334088Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-12T16:27:47.920413Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
33 of 33 outbound references displayed
External citation measurements
77
pith, observed 2026-08-05T02:28:24.338817Z
Observation 4b86f9a7-ec6b-47c0-9a91-4452105c5311 · outbound
CogVLM: Visual Expert for Pretrained Language Models OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 625d3bbf-dadb-4393-9b83-582458ed4974 · outbound
CogVLM: Visual Expert for Pretrained Language Models Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 3f6d95b5-5760-4aa0-9141-bd09618eae37 · outbound
CogVLM: Visual Expert for Pretrained Language Models Murel: Multimodal relational reasoning for visual ques- tion answering
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9da9aa74-5e99-41d6-a622-b84d30f38b24 · outbound
CogVLM: Visual Expert for Pretrained Language Models Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 09cd8215-1680-48de-993a-e4103b1646ff · outbound
CogVLM: Visual Expert for Pretrained Language Models Generating More Pertinent Captions by Leveraging Semantics and Style on Multi-Source Datasets
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 2dc7da0d-0149-4e62-8740-615d6cc017fc · outbound
CogVLM: Visual Expert for Pretrained Language Models DreamLLM: Synergistic Multimodal Comprehension and Creation
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation adf8a310-159f-45f3-8dcb-db23fc9b56e4 · outbound
CogVLM: Visual Expert for Pretrained Language Models PaLM-E: An Embodied Multimodal Language Model
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation bb095439-4e28-4314-ac2f-8f2a54d0f649 · outbound
CogVLM: Visual Expert for Pretrained Language Models Measuring Massive Multitask Language Understanding
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 62779542-b21e-4d98-b8bd-ba319fb66631 · outbound
CogVLM: Visual Expert for Pretrained Language Models and Johnson, M
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9f85b72b-94a1-453a-ab5e-017259b938ef · outbound
CogVLM: Visual Expert for Pretrained Language Models LoRA: Low-Rank Adaptation of Large Language Models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f71bd79c-755e-439f-a1b6-e5e671d9ada2 · outbound
CogVLM: Visual Expert for Pretrained Language Models and Kanan, C
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f20819c4-a886-473b-a30d-2a371eeec989 · outbound
CogVLM: Visual Expert for Pretrained Language Models Referitgame: Referring to objects in photographs of natu- ral scenes
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 2fa0e400-3725-4159-9c6c-bf2f7eee87b1 · outbound
CogVLM: Visual Expert for Pretrained Language Models SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 1dcdb279-1b6e-4ad5-bd79-483278e7dfd3 · outbound
CogVLM: Visual Expert for Pretrained Language Models SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation da9ffbd5-379e-4e7d-bd97-b5a39e506ed9 · outbound
CogVLM: Visual Expert for Pretrained Language Models Prismer: A Vision-Language Model with Multi-Task Experts
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 663d4eb7-e7ee-4c9c-8322-5813b578a765 · outbound
CogVLM: Visual Expert for Pretrained Language Models MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c1af65c4-c446-46a8-82b2-cccc93584cb9 · outbound
CogVLM: Visual Expert for Pretrained Language Models K., and Chakraborty, A
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9c2b806b-a009-4a4a-9b98-df6822e90a3c · outbound
CogVLM: Visual Expert for Pretrained Language Models Kosmos-2: Grounding Multimodal Large Language Models to the World
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8a0c2fc0-effc-4960-b243-1333364587e4 · outbound
CogVLM: Visual Expert for Pretrained Language Models GLU Variants Improve Transformer
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6c139398-96d6-49f9-83b4-9413b0aeb3e2 · outbound
CogVLM: Visual Expert for Pretrained Language Models Textcaps: a dataset for image captioning with reading comprehen- sion
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ae3d4092-7407-449d-8aa4-9e8916771767 · outbound
CogVLM: Visual Expert for Pretrained Language Models Generative Multimodal Models are In-Context Learners
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 19a88fac-d386-40a9-b7cd-50e9e5f1e1df · outbound
CogVLM: Visual Expert for Pretrained Language Models GIT: A Generative Image-to-text Transformer for Vision and Language
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 176fbb6b-ee30-4a1c-80a2-e6bf6a6721d7 · outbound
CogVLM: Visual Expert for Pretrained Language Models mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 5afb32c7-08f4-4794-a333-92b4451453ff · outbound
CogVLM: Visual Expert for Pretrained Language Models Ferret: Refer and Ground Anything Anywhere at Any Granularity
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 195d2684-3568-427c-b238-e9559c78c3cc · outbound
CogVLM: Visual Expert for Pretrained Language Models CoCa: Contrastive Captioners are Image-Text Foundation Models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ebbba877-31a4-46ba-b8a4-fbe87650d643 · outbound
CogVLM: Visual Expert for Pretrained Language Models C., and Berg, T
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 2f5fda1f-dd3d-477c-8ba8-0b0bc7c61300 · outbound
CogVLM: Visual Expert for Pretrained Language Models MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation a4051937-4f6c-4242-acda-184b081f658c · outbound
CogVLM: Visual Expert for Pretrained Language Models MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c4541158-a550-4c90-addd-f9b7f61169f1 · outbound
CogVLM: Visual Expert for Pretrained Language Models MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8a8cde92-4dde-4d6d-860a-4c83055990f1 · outbound
CogVLM: Visual Expert for Pretrained Language Models Unresolved cited work
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 57e1ef59-ae86-4564-ad34-1cd7213156dd · outbound
CogVLM: Visual Expert for Pretrained Language Models in”, “near
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 76fe0e15-6009-4abf-9cb9-e4dbc990f6c5 · outbound
CogVLM: Visual Expert for Pretrained Language Models which”-type 15 CogVLM: Visual Expert for Pretrained Language Models question, such as “Which is the small computer in the corner?
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 7bffdc91-e0ab-4720-91dc-f2bd7d3198ee · outbound
CogVLM: Visual Expert for Pretrained Language Models Unresolved cited work
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation bb0b95b1-382f-4360-be3b-63d8e10b47e8 · inbound
A Survey on Multimodal Large Language Models CogVLM: Visual Expert for Pretrained Language Models
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f01eb9ad-d63b-4779-aacf-59f9eded3c5a · inbound
MMBench: Is Your Multi-modal Model an All-around Player? CogVLM: Visual Expert for Pretrained Language Models
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation cd2b4c3f-a9fb-4032-8ada-7fea5f961cd6 · inbound
MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI CogVLM: Visual Expert for Pretrained Language Models
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c3857ff0-9581-4aae-be51-8c33e9374e42 · inbound
SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents CogVLM: Visual Expert for Pretrained Language Models
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 563dea3e-45a6-4451-a552-83786cce7ae1 · inbound
DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models CogVLM: Visual Expert for Pretrained Language Models
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 74c5dc15-11c8-4ce0-aa77-d1029d41eb18 · inbound
Scaling Rectified Flow Transformers for High-Resolution Image Synthesis CogVLM: Visual Expert for Pretrained Language Models
Reference 192
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c60aef8f-086a-4728-90b2-499a09754b02 · inbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training CogVLM: Visual Expert for Pretrained Language Models
Reference 114
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation cde6715c-4def-4681-9d01-777b8d6b1e43 · inbound
Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models CogVLM: Visual Expert for Pretrained Language Models
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 91d94187-9a12-40b2-9836-1c4a7b0ea5a5 · inbound
Are We on the Right Way for Evaluating Large Vision-Language Models? CogVLM: Visual Expert for Pretrained Language Models
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ed3cac7c-96ca-47b9-8f86-8590ce4ce26e · inbound
How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites CogVLM: Visual Expert for Pretrained Language Models
Reference 117
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 2728c380-e409-4570-ba97-9d949b7b1da1 · inbound
Hallucination of Multimodal Large Language Models: A Survey CogVLM: Visual Expert for Pretrained Language Models
Reference 169
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c02d76f9-d739-43c7-88aa-ba2ab3e72a24 · inbound
Cracking the Code of Juxtaposition: Can AI Models Understand the Humorous Contradictions CogVLM: Visual Expert for Pretrained Language Models
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 37fccd2a-a5a8-4c60-a750-8f363b616856 · inbound
LVBench: An Extreme Long Video Understanding Benchmark CogVLM: Visual Expert for Pretrained Language Models
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 91b19249-652c-4607-900e-656d843fa259 · inbound
PaliGemma: A versatile 3B VLM for transfer CogVLM: Visual Expert for Pretrained Language Models
Reference 143
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8388e13e-22c3-4078-8a3b-d33e97485d23 · inbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone CogVLM: Visual Expert for Pretrained Language Models
Reference 102
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 7f99b24f-61a3-40f1-b2db-3c1b961f7800 · inbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models CogVLM: Visual Expert for Pretrained Language Models
Reference 250
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d94def45-28f3-4cc1-9511-1acdca22261a · inbound
CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer CogVLM: Visual Expert for Pretrained Language Models
Reference 103
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation defb3e9e-c668-49ce-91df-cbdba7e44a1d · inbound
CogVLM2: Visual Language Models for Image and Video Understanding CogVLM: Visual Expert for Pretrained Language Models
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 724ddaa9-7e39-4b2a-adf6-984c04aa1b4f · inbound
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference CogVLM: Visual Expert for Pretrained Language Models
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation dedad807-91d7-4d41-ad1f-0185068474a2 · inbound
ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection CogVLM: Visual Expert for Pretrained Language Models
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 74d654de-be48-4838-aaa6-d1f869bc91ca · inbound
VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality Documents CogVLM: Visual Expert for Pretrained Language Models
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ab197ee2-0f26-4732-8964-92cdb419fb80 · inbound
LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding CogVLM: Visual Expert for Pretrained Language Models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 61c960ff-14dc-47dc-a679-324aae4881cc · inbound
Senna: Bridging Large Vision-Language Models and End-to-End Autonomous Driving CogVLM: Visual Expert for Pretrained Language Models
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9622adbb-a55c-4c06-a685-10946707da83 · inbound
Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models CogVLM: Visual Expert for Pretrained Language Models
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation a7edf753-b562-4d43-b73e-52dc00117bd9 · inbound
Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization CogVLM: Visual Expert for Pretrained Language Models
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 1767b3c0-e5e6-4c2c-bfd6-305adca70dc8 · inbound
AdaptAgent: Adapting Multimodal Web Agents with Few-Shot Learning from Human Demonstrations CogVLM: Visual Expert for Pretrained Language Models
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ff5cd4d-0220-46b9-8588-730cdad390dd · inbound
FoPru: Focal Pruning for Efficient Large Vision-Language Models CogVLM: Visual Expert for Pretrained Language Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21a86877-63f8-4aef-98fe-69ff3cccdbea · inbound
De-biased Multimodal Electrocardiogram Analysis CogVLM: Visual Expert for Pretrained Language Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0a32663-3ec7-471b-90af-7515e6269bf9 · inbound
Continual SFT Matches Multimodal RLHF with Negative Supervision CogVLM: Visual Expert for Pretrained Language Models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64d55457-6c01-4971-a84b-5ff9ad6c3059 · inbound
MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs CogVLM: Visual Expert for Pretrained Language Models
Reference 184
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9b05143-1491-4162-b3e6-62c7a56d6240 · inbound
Pathways on the Image Manifold: Image Editing via Video Generation CogVLM: Visual Expert for Pretrained Language Models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation edea6d44-21f1-4b73-93e6-bf0f9cdef135 · inbound
Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach CogVLM: Visual Expert for Pretrained Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 585288b8-058f-4bc6-a95a-b5a3a7a589b7 · inbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding CogVLM: Visual Expert for Pretrained Language Models
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dfa154cd-9702-4351-b68a-78497c026d30 · inbound
Improving Medical Diagnostics with Vision-Language Models: Convex Hull-Based Uncertainty Analysis CogVLM: Visual Expert for Pretrained Language Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8332cbc5-faa8-4c16-95d2-3ed4313155c7 · inbound
Perturb and Recover: Fine-tuning for Effective Backdoor Removal from CLIP CogVLM: Visual Expert for Pretrained Language Models
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6342a59d-883b-4146-85e9-a195894c8a59 · inbound
Multi-View Incongruity Learning for Multimodal Sarcasm Detection CogVLM: Visual Expert for Pretrained Language Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52e08cb8-b7fd-4948-8414-c6623e9a8449 · inbound
Beyond Pixels: Text Enhances Generalization in Real-World Image Restoration CogVLM: Visual Expert for Pretrained Language Models
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49e0f7b4-3685-4d8d-8ba2-841c10fe3838 · inbound
PKRD-CoT: A Unified Chain-of-thought Prompting for Multi-Modal Large Language Models in Autonomous Driving CogVLM: Visual Expert for Pretrained Language Models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2caa988-2a4a-4ab2-8623-9eebfe452aa3 · inbound
Agri-LLaVA: Knowledge-Infused Large Multimodal Assistant on Agricultural Pests and Diseases CogVLM: Visual Expert for Pretrained Language Models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4ba772b-d8ff-4974-9980-796228dceb1a · inbound
AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information? CogVLM: Visual Expert for Pretrained Language Models
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96a96334-1a74-4b8c-946b-61c0d9d3a5e4 · inbound
AdvDreamer Unveils: Are Vision-Language Models Truly Ready for Real-World 3D Variations? CogVLM: Visual Expert for Pretrained Language Models
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6146e7bc-49c0-40c8-8f60-38f6987a6fd9 · inbound
Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis CogVLM: Visual Expert for Pretrained Language Models
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39a38205-86a0-41d5-8e18-fbdee6d4621a · inbound
Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension CogVLM: Visual Expert for Pretrained Language Models
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a50e6e9-38ca-415a-a5a2-167b8703414a · inbound
AnyDressing: Customizable Multi-Garment Virtual Dressing via Latent Diffusion Models CogVLM: Visual Expert for Pretrained Language Models
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 630d2cc9-02a1-4e40-8ccb-e6e1e968d272 · inbound
VladVA: Discriminative Fine-tuning of LVLMs CogVLM: Visual Expert for Pretrained Language Models
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 927b86c0-83fc-4ee1-8fa7-88eb074a9566 · inbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling CogVLM: Visual Expert for Pretrained Language Models
Reference 249
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 5c68940f-5fda-4306-97e0-a664b6ef4eea · inbound
RSUniVLM: A Unified Vision Language Model for Remote Sensing via Granularity-oriented Mixture of Experts CogVLM: Visual Expert for Pretrained Language Models
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b17676b0-df9e-4568-9076-ca336662f59a · inbound
[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs CogVLM: Visual Expert for Pretrained Language Models
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0eaa8d3-4c83-47aa-8133-c6cca7ca74b2 · inbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding CogVLM: Visual Expert for Pretrained Language Models
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e68e8702-e122-46c0-9500-451de0b6bcb1 · inbound
Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation CogVLM: Visual Expert for Pretrained Language Models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ea7d8b6-87ae-4ddc-ab77-d7f697758fd0 · inbound
Enhancing Nursing and Elderly Care with Large Language Models: An AI-Driven Framework CogVLM: Visual Expert for Pretrained Language Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 877ec80c-6122-4bc2-ac7d-1dc5a236e381 · inbound
Attention-driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models without Fine-Tuning CogVLM: Visual Expert for Pretrained Language Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2749b36c-99e9-4b98-b6ed-a938fe68861a · inbound
From 2D CAD Drawings to 3D Parametric Models: A Vision-Language Approach CogVLM: Visual Expert for Pretrained Language Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ecb7134a-b6a2-4d1b-af30-570c73754c14 · inbound
PunchBench: Benchmarking MLLMs in Multimodal Punchline Comprehension CogVLM: Visual Expert for Pretrained Language Models
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f706b0dc-398a-4973-92fb-f12b81ed17fb · inbound
Advancing Comprehensive Aesthetic Insight with Multi-Scale Text-Guided Self-Supervised Learning CogVLM: Visual Expert for Pretrained Language Models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ca56343-1589-4d85-a10a-5f224ddf4070 · inbound
Rethinking Comprehensive Benchmark for Chart Understanding: A Perspective from Scientific Literature CogVLM: Visual Expert for Pretrained Language Models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dadbc503-a094-43af-b542-441b1a132c0b · inbound
Multi-Dimensional Insights: Benchmarking Real-World Personalization in Large Multimodal Models CogVLM: Visual Expert for Pretrained Language Models
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86015391-7127-4d78-811c-616777739bfe · inbound
Deploying Foundation Model Powered Agent Services: A Survey CogVLM: Visual Expert for Pretrained Language Models
Reference 157
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bfe81834-379f-4856-a0b2-c8471692022d · inbound
Unlocking the Potential of Weakly Labeled Data: A Co-Evolutionary Learning Framework for Abnormality Detection and Report Generation CogVLM: Visual Expert for Pretrained Language Models
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4dce76a-0d41-4b56-a2ee-5d023895c247 · inbound
Consistency of Compositional Generalization across Multiple Levels CogVLM: Visual Expert for Pretrained Language Models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0dfd2b33-9f51-48f3-9f88-6e7ca5665c2e · inbound
HumanVBench: Probing Human-Centric Video Understanding in MLLMs with Automatically Synthesized Benchmarks CogVLM: Visual Expert for Pretrained Language Models
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 5aed9335-5f5f-4087-81fb-1af4fb7cf83a · inbound
DreamFit: Garment-Centric Human Generation via a Lightweight Anything-Dressing Encoder CogVLM: Visual Expert for Pretrained Language Models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 252471fb-a11c-4262-80f7-2a748906c09d · inbound
MMFactory: A Universal Solution Search Engine for Vision-Language Tasks CogVLM: Visual Expert for Pretrained Language Models
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fdc05b06-be11-43dd-bac4-348323dcb897 · inbound
ICM-Assistant: Instruction-tuning Multimodal Large Language Models for Rule-based Explainable Image Content Moderation CogVLM: Visual Expert for Pretrained Language Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36c3196c-d36c-4d9c-9ba2-9839c7216a56 · inbound
RapGuard: Safeguarding Multimodal Large Language Models via Rationale-aware Defensive Prompting CogVLM: Visual Expert for Pretrained Language Models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4514247-2ebd-4ed2-bc74-e7053b154cfd · inbound
ViPCap: Retrieval Text-Based Visual Prompts for Lightweight Image Captioning CogVLM: Visual Expert for Pretrained Language Models
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b1ceba9-f1e3-476b-9189-44c649c5b49d · inbound
ETTA: Elucidating the Design Space of Text-to-Audio Models CogVLM: Visual Expert for Pretrained Language Models
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 497cdf64-dd86-408d-a9bf-31c6b0a0e194 · inbound
Is Your Text-to-Image Model Robust to Caption Noise? CogVLM: Visual Expert for Pretrained Language Models
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30bed91f-ec62-4b7e-8fde-d8312b5b6016 · inbound
ST$^3$: Accelerating Multimodal Large Language Model by Spatial-Temporal Visual Token Trimming CogVLM: Visual Expert for Pretrained Language Models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33b1bf55-a471-450e-9439-51434881b217 · inbound
MM-MoralBench: A MultiModal Moral Evaluation Benchmark for Large Vision-Language Models CogVLM: Visual Expert for Pretrained Language Models
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation bc2621c8-ed9f-4f68-b995-4d1f4e3459b5 · inbound
VisionReward: Fine-Grained Multi-Dimensional Human Preference Learning for Image and Video Generation CogVLM: Visual Expert for Pretrained Language Models
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9bcf1ecd-5a72-4ea7-95d3-87d3b26dae30 · inbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning CogVLM: Visual Expert for Pretrained Language Models
Reference 144
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6673995a-2184-4aa0-92c2-0e47c10a8f6f · inbound
IllusionBench+: A Large-scale and Comprehensive Benchmark for Visual Illusion Understanding in Vision-Language Models CogVLM: Visual Expert for Pretrained Language Models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f57e492e-a729-44b6-ab68-593985511094 · inbound
GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models CogVLM: Visual Expert for Pretrained Language Models
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 926dd163-b180-4a9b-8484-4d248918d0a2 · inbound
MoColl: Agent-Based Specific and General Model Collaboration for Image Captioning CogVLM: Visual Expert for Pretrained Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc59dfea-9a5f-4d55-acb7-7be0fffb2415 · inbound
Visual Large Language Models for Generalized and Specialized Applications CogVLM: Visual Expert for Pretrained Language Models
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dcb7a881-5241-43e7-84da-9e2fcf6fb8b1 · inbound
Foundations of GenIR CogVLM: Visual Expert for Pretrained Language Models
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c20b871e-8556-44df-8fee-58b08fa38ff4 · inbound
MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models CogVLM: Visual Expert for Pretrained Language Models
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 1dff1d06-f1c9-4ff2-9dd9-654a89816006 · inbound
InfiGUIAgent: A Multimodal Generalist GUI Agent with Native Reasoning and Reflection CogVLM: Visual Expert for Pretrained Language Models
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ad8e6e5-a33a-4a77-a1b2-694f80493fcc · inbound
GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing CogVLM: Visual Expert for Pretrained Language Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c7a0eb7-ceb2-47c2-a88a-fabe790fdcc5 · inbound
Zero-Shot Scene Understanding for Automatic Target Recognition Using Large Vision-Language Models CogVLM: Visual Expert for Pretrained Language Models
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08684fc7-abda-46a9-b442-365b65683e12 · inbound
LeapVAD: A Leap in Autonomous Driving via Cognitive Perception and Dual-Process Thinking CogVLM: Visual Expert for Pretrained Language Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d873e978-b25b-4177-9e5f-b885631f7cf5 · inbound
Parameter-Efficient Fine-Tuning for Foundation Models CogVLM: Visual Expert for Pretrained Language Models
Reference 264
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c94e11a-902c-4632-b71e-6a9746b15428 · inbound
Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models CogVLM: Visual Expert for Pretrained Language Models
Reference 191
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fdede704-6ae9-4a90-97d6-a00146a2845a · inbound
A Causality-aware Paradigm for Evaluating Creativity of Multimodal Large Language Models CogVLM: Visual Expert for Pretrained Language Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50448d07-38c6-45dc-997b-dde94c40c5b2 · inbound
MME-Industry: A Cross-Industry Multimodal Evaluation Benchmark CogVLM: Visual Expert for Pretrained Language Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 552eab35-d5bf-4d44-9fa3-3c97d50748a0 · inbound
Membership Inference Attacks Against Vision-Language Models CogVLM: Visual Expert for Pretrained Language Models
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae49b89e-fb72-4cfa-9e34-b119b4cf6da8 · inbound
RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs CogVLM: Visual Expert for Pretrained Language Models
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 255b3b4f-9f35-4f6c-8fc1-341bcef9839a · inbound
MQuant: Unleashing the Inference Potential of Multimodal Large Language Models via Full Static Quantization CogVLM: Visual Expert for Pretrained Language Models
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f630ceb-7b5a-4204-923b-fb43ffb6d6e5 · inbound
Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective CogVLM: Visual Expert for Pretrained Language Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f8917dd-0a2a-4a17-889c-e93499148733 · inbound
SKI Models: Skeleton Induced Vision-Language Embeddings for Understanding Activities of Daily Living CogVLM: Visual Expert for Pretrained Language Models
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 762dcc6a-bb1d-4db7-9f9b-e93c91d3ba83 · inbound
Multitwine: Multi-Object Compositing with Text and Layout Control CogVLM: Visual Expert for Pretrained Language Models
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb2464dd-4628-4751-a494-465575e88afc · inbound
Animate Anyone 2: High-Fidelity Character Image Animation with Environment Affordance CogVLM: Visual Expert for Pretrained Language Models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0dce2cd-fd45-4e5c-8037-27e7602e7d54 · inbound
EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering CogVLM: Visual Expert for Pretrained Language Models
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e12ed85-88bc-4133-92d1-5561f8b37bc8 · inbound
When 'YES' Meets 'BUT': Can Large Models Comprehend Contradictory Humor Through Comparative Reasoning? CogVLM: Visual Expert for Pretrained Language Models
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c078eb9a-f62e-4e13-9c20-2097f5044aab · inbound
AdaMMS: Model Merging for Heterogeneous Multimodal Large Language Models with Unsupervised Coefficient Optimization CogVLM: Visual Expert for Pretrained Language Models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 41640757-fd0e-440a-b384-afa4376651ec · inbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models CogVLM: Visual Expert for Pretrained Language Models
Reference 124
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation acb0fee3-f39e-4074-945b-701c840f2d67 · inbound
InfiGUI-R1: Advancing Multimodal GUI Agents from Reactive Actors to Deliberative Reasoners CogVLM: Visual Expert for Pretrained Language Models
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 68b6a288-7828-4199-8f15-ee9f452b8a6a · inbound
From Street View to Visual Network: Mapping the Visibility of Urban Landmarks with Vision-Language Models CogVLM: Visual Expert for Pretrained Language Models
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6e0f3a88-6f01-436d-83d3-2e7ed9ae359b · inbound
Efficient Multi-modal Long Context Learning for Training-free Adaptation CogVLM: Visual Expert for Pretrained Language Models
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.