Pith. sign in

Paper Citation Record · LEDGER

CogVLM: Visual Expert for Pretrained Language Models

As of 12 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 100 inbound Pith citation observations for arXiv:2311.03079.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2311.03079 v2

Coverage vector

measured 33 of 33 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-15T15:46:06.334088Z

measured 133 of 133 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 100 of 165 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T16:27:47.920413Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

33 of 33 outbound references displayed

  • verified exact18
  • verified fuzzy9
  • unresolved2
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch4

External citation measurements

77
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 4b86f9a7-ec6b-47c0-9a91-4452105c5311 · outbound

This paper cites OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models.

CogVLM: Visual Expert for Pretrained Language Models OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T15:46:06.464567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-15T15:46:06.334088Z digest=sha256:74763f3b034766ce4f6f0c1d4399e4d674c2cfc5d3b336f96b9b9d0c3532997f

Observation 625d3bbf-dadb-4393-9b83-582458ed4974 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

CogVLM: Visual Expert for Pretrained Language Models Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-15T15:46:06.438703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-15T15:46:06.334088Z digest=sha256:9988d1536dfa928de5a3b858cf39f56e32a9dc643f5575c84ed67923381a8498

Observation 3f6d95b5-5760-4aa0-9141-bd09618eae37 · outbound

This paper cites Murel: Multimodal relational reasoning for visual ques- tion answering.

CogVLM: Visual Expert for Pretrained Language Models Murel: Multimodal relational reasoning for visual ques- tion answering

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:46:06.547274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-15T15:46:06.334088Z digest=sha256:2e4a25b1307c3a08a5034ed8ca3d0170a494b6f785dc62e5a51aa207b071d94b

Observation 9da9aa74-5e99-41d6-a622-b84d30f38b24 · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

CogVLM: Visual Expert for Pretrained Language Models Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-15T15:46:06.457279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-15T15:46:06.334088Z digest=sha256:aaa45d695b6e5e48b3e356d6059bf1e4e0ea959593aad1aa15e0bee2f66c14cc

Observation 09cd8215-1680-48de-993a-e4103b1646ff · outbound

This paper cites Generating More Pertinent Captions by Leveraging Semantics and Style on Multi-Source Datasets.

CogVLM: Visual Expert for Pretrained Language Models Generating More Pertinent Captions by Leveraging Semantics and Style on Multi-Source Datasets

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-15T15:46:06.472683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-15T15:46:06.334088Z digest=sha256:4e16a862d8024f517037969d998a1338c1a08628b27471ba632cb3fe0ddd1a60

Observation 2dc7da0d-0149-4e62-8740-615d6cc017fc · outbound

This paper cites DreamLLM: Synergistic Multimodal Comprehension and Creation.

CogVLM: Visual Expert for Pretrained Language Models DreamLLM: Synergistic Multimodal Comprehension and Creation

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-15T15:46:06.504847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-15T15:46:06.334088Z digest=sha256:a9b0ee1d4f7ea72f2c3ae6e6ba7587889ad8cd95fbaf2583334d5ac783d7e2fa

Observation adf8a310-159f-45f3-8dcb-db23fc9b56e4 · outbound

This paper cites PaLM-E: An Embodied Multimodal Language Model.

CogVLM: Visual Expert for Pretrained Language Models PaLM-E: An Embodied Multimodal Language Model

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T15:46:06.522386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-15T15:46:06.334088Z digest=sha256:9b3821705ce99c26ed9743586186ff1618f70e9e5d4fecc48e1e66bca0731c80

Observation bb095439-4e28-4314-ac2f-8f2a54d0f649 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

CogVLM: Visual Expert for Pretrained Language Models Measuring Massive Multitask Language Understanding

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-15T15:46:06.405707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-15T15:46:06.334088Z digest=sha256:132cdf5d39d61962acf6fb7b8e5a034c8e72d4ade0d9559a8b5c3819ad5be9fb

Observation 62779542-b21e-4d98-b8bd-ba319fb66631 · outbound

This paper cites and Johnson, M.

CogVLM: Visual Expert for Pretrained Language Models and Johnson, M

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:46:06.551835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-15T15:46:06.334088Z digest=sha256:bc6e398e89e6470df2c1deb8d1950d37a1b5f09f9e8885fa511efeb74b95dcf9

Observation 9f85b72b-94a1-453a-ab5e-017259b938ef · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

CogVLM: Visual Expert for Pretrained Language Models LoRA: Low-Rank Adaptation of Large Language Models

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T15:46:06.447809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-15T15:46:06.334088Z digest=sha256:42b5084f32e234de55e40e509e2dcd6686cbb2d51c7bf01d04d18520c99f1945

Observation f71bd79c-755e-439f-a1b6-e5e671d9ada2 · outbound

This paper cites and Kanan, C.

CogVLM: Visual Expert for Pretrained Language Models and Kanan, C

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:46:06.556366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-15T15:46:06.334088Z digest=sha256:c56debe84463eb6f440afe4dc2a3847d5ce9536595eb260ba1e47554cd9a601e

Observation f20819c4-a886-473b-a30d-2a371eeec989 · outbound

This paper cites Referitgame: Referring to objects in photographs of natu- ral scenes.

CogVLM: Visual Expert for Pretrained Language Models Referitgame: Referring to objects in photographs of natu- ral scenes

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:46:06.560300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-15T15:46:06.334088Z digest=sha256:bd9f5a0ff6db4f69fe3e60673ed9ba62bd8dfe8ffe3a0affbc5dbf9ec17e6bb8

Observation 2fa0e400-3725-4159-9c6c-bf2f7eee87b1 · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

CogVLM: Visual Expert for Pretrained Language Models SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-15T15:46:06.480520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-15T15:46:06.334088Z digest=sha256:2344ca67d387d8ba74b3159b880908b5df4a00b35e9385d23d8415206c5bed7f

Observation 1dcdb279-1b6e-4ad5-bd79-483278e7dfd3 · outbound

This paper cites SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models.

CogVLM: Visual Expert for Pretrained Language Models SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:03:27.047113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-15T15:46:06.334088Z digest=sha256:984a92e0acac513edeb97e96156b87daa1c2eda932153bacec44b1be0df8d5a2

Observation da9ffbd5-379e-4e7d-bd97-b5a39e506ed9 · outbound

This paper cites Prismer: A Vision-Language Model with Multi-Task Experts.

CogVLM: Visual Expert for Pretrained Language Models Prismer: A Vision-Language Model with Multi-Task Experts

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T15:46:06.492169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-15T15:46:06.334088Z digest=sha256:1da7b12548fd73052ffbfb21b4d9046125d65767ececb2db68a2de37347b2558

Observation 663d4eb7-e7ee-4c9c-8322-5813b578a765 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

CogVLM: Visual Expert for Pretrained Language Models MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-15T15:46:06.498574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-15T15:46:06.334088Z digest=sha256:46b7607962bc3e46b80553833d3978769d9a475751c9f9a54aaf6c664f3d261e

Observation c1af65c4-c446-46a8-82b2-cccc93584cb9 · outbound

This paper cites K., and Chakraborty, A.

CogVLM: Visual Expert for Pretrained Language Models K., and Chakraborty, A

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:46:06.565138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-15T15:46:06.334088Z digest=sha256:8f5ebaf976b8ee407631b89df9bbd48943d26868a4643abbee3168bd56909927

Observation 9c2b806b-a009-4a4a-9b98-df6822e90a3c · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

CogVLM: Visual Expert for Pretrained Language Models Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-15T15:46:06.510686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-15T15:46:06.334088Z digest=sha256:6b7d0e915f84b6df8f607555a3ef9c157a62b1e33a2789d644ff6b107461b0e7

Observation 8a0c2fc0-effc-4960-b243-1333364587e4 · outbound

This paper cites GLU Variants Improve Transformer.

CogVLM: Visual Expert for Pretrained Language Models GLU Variants Improve Transformer

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-15T15:46:06.516264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-15T15:46:06.334088Z digest=sha256:f2bde5449b61e2683a1dbe7565b365420956dc4968ca6def9b88e2267a2c02d1

Observation 6c139398-96d6-49f9-83b4-9413b0aeb3e2 · outbound

This paper cites Textcaps: a dataset for image captioning with reading comprehen- sion.

CogVLM: Visual Expert for Pretrained Language Models Textcaps: a dataset for image captioning with reading comprehen- sion

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:46:06.570077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-15T15:46:06.334088Z digest=sha256:5bf4e56b5260c184783cbff84b8766ae6dfe52e6e51ff9051a7cfe1ce603af31

Observation ae3d4092-7407-449d-8aa4-9e8916771767 · outbound

This paper cites Generative Multimodal Models are In-Context Learners.

CogVLM: Visual Expert for Pretrained Language Models Generative Multimodal Models are In-Context Learners

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-15T15:46:06.530109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-15T15:46:06.334088Z digest=sha256:e37abf87374da74e662fcfa11c5bd4c8fd8aa4124b8ffd008cbe9233d7e10e15

Observation 19a88fac-d386-40a9-b7cd-50e9e5f1e1df · outbound

This paper cites GIT: A Generative Image-to-text Transformer for Vision and Language.

CogVLM: Visual Expert for Pretrained Language Models GIT: A Generative Image-to-text Transformer for Vision and Language

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-16T20:54:07.740206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-15T15:46:06.334088Z digest=sha256:ec3a0399761d0d20ef6d88af5ee323e1c861222ce44e606db88d8e003f672b7b

Observation 176fbb6b-ee30-4a1c-80a2-e6bf6a6721d7 · outbound

This paper cites mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration.

CogVLM: Visual Expert for Pretrained Language Models mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-18T03:18:51.946750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-15T15:46:06.334088Z digest=sha256:42000a0458566dde195d413f4d0bc61dad9a28c4a1458380428fb562637811d0

Observation 5afb32c7-08f4-4794-a333-92b4451453ff · outbound

This paper cites Ferret: Refer and Ground Anything Anywhere at Any Granularity.

CogVLM: Visual Expert for Pretrained Language Models Ferret: Refer and Ground Anything Anywhere at Any Granularity

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-15T15:46:06.379349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-15T15:46:06.334088Z digest=sha256:e53fd39f7b87c5ba2fe6c0c61228cbcea71f0a2fdd3ef51022f2081fad19a0da

Observation 195d2684-3568-427c-b238-e9559c78c3cc · outbound

This paper cites CoCa: Contrastive Captioners are Image-Text Foundation Models.

CogVLM: Visual Expert for Pretrained Language Models CoCa: Contrastive Captioners are Image-Text Foundation Models

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-15T15:46:06.394093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-15T15:46:06.334088Z digest=sha256:3b48e565c507ee063ae4f6b3318292cf88b9a2ec7c2332504684bc791da9d8f8

Observation ebbba877-31a4-46ba-b8a4-fbe87650d643 · outbound

This paper cites C., and Berg, T.

CogVLM: Visual Expert for Pretrained Language Models C., and Berg, T

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:46:06.574371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-15T15:46:06.334088Z digest=sha256:2362b8fdfe155ab613cd4e7ef28fcf9bb065161e5826e7f7b8f7d5064f3080d9

Observation 2f5fda1f-dd3d-477c-8ba8-0b0bc7c61300 · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

CogVLM: Visual Expert for Pretrained Language Models MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-15T15:46:06.413968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-15T15:46:06.334088Z digest=sha256:c319e3daa14f5edad64ffde84ad55b96dca253b1ba5541498a637cb982ccb68c

Observation a4051937-4f6c-4242-acda-184b081f658c · outbound

This paper cites MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI.

CogVLM: Visual Expert for Pretrained Language Models MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-15T15:46:06.421184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-15T15:46:06.334088Z digest=sha256:0eea104b2f735af9a94f712d7fcc3d6d753bf098e4b72ac4623800fa55ffc79b

Observation c4541158-a550-4c90-addd-f9b7f61169f1 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

CogVLM: Visual Expert for Pretrained Language Models MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-15T15:46:06.430717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-15T15:46:06.334088Z digest=sha256:0274f4c3c5aa12c62f9033713699df4f740d27088422fc7958a57c4b7af8e511

Observation 8a8cde92-4dde-4d6d-860a-4c83055990f1 · outbound

This paper cites an unresolved cited work.

CogVLM: Visual Expert for Pretrained Language Models Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-05-15T15:46:06.579214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-15T15:46:06.334088Z digest=sha256:b6a1301b4a04ca57673d7205fc04b6c7968898d4e2bbbccdc8f4605d61707aac

Observation 57e1ef59-ae86-4564-ad34-1cd7213156dd · outbound

This paper cites in”, “near.

CogVLM: Visual Expert for Pretrained Language Models in”, “near

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:46:06.583337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-15T15:46:06.334088Z digest=sha256:819fd0d7e5fb0ce963a6c9eb9b101e26597b5886a4da8f30cc12484bef479e56

Observation 76fe0e15-6009-4abf-9cb9-e4dbc990f6c5 · outbound

This paper cites which”-type 15 CogVLM: Visual Expert for Pretrained Language Models question, such as “Which is the small computer in the corner?.

CogVLM: Visual Expert for Pretrained Language Models which”-type 15 CogVLM: Visual Expert for Pretrained Language Models question, such as “Which is the small computer in the corner?

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:46:06.588349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-15T15:46:06.334088Z digest=sha256:7e80e5a02015ffe7962e8b2f58214d716a3a5df27d572878e58782bb9ce3497e

Observation 7bffdc91-e0ab-4720-91dc-f2bd7d3198ee · outbound

This paper cites an unresolved cited work.

CogVLM: Visual Expert for Pretrained Language Models Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-05-15T15:46:06.542791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-15T15:46:06.334088Z digest=sha256:3b9c567cc8dd9472e5fb7874285e967987aff0cb5976642e62073e094fe86194

Pith citing papers

Observation bb0b95b1-382f-4360-be3b-63d8e10b47e8 · inbound

A Survey on Multimodal Large Language Models cites this paper.

A Survey on Multimodal Large Language Models CogVLM: Visual Expert for Pretrained Language Models

Reference 77

Resolution
verified exact
local_arxiv, observed 2026-05-16T02:56:42.577157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T02:56:41.658658Z digest=sha256:13df58ae2bfc93ac987e8425463b1ba3f3464c6db509200112a2d90a7f49afa5

Observation f01eb9ad-d63b-4779-aacf-59f9eded3c5a · inbound

MMBench: Is Your Multi-modal Model an All-around Player? cites this paper.

MMBench: Is Your Multi-modal Model an All-around Player? CogVLM: Visual Expert for Pretrained Language Models

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-15T15:46:06.589810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-12T17:20:53.687692Z digest=sha256:cc09b8a08d4faa8c32e3bbde06601bb12da27b1ae85bf6d7f968044366effae6

Observation cd2b4c3f-a9fb-4032-8ada-7fea5f961cd6 · inbound

MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI cites this paper.

MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI CogVLM: Visual Expert for Pretrained Language Models

Reference 78

Resolution
verified exact
arxiv_id, observed 2026-05-15T15:46:06.589810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-15T05:37:41.401736Z digest=sha256:4c1174f1ee5debb4882aa62b2e3c9d1d3daf972337023e95e7cbae492ea721c3

Observation c3857ff0-9581-4aae-be51-8c33e9374e42 · inbound

SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents cites this paper.

SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents CogVLM: Visual Expert for Pretrained Language Models

Reference 34

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T10:09:46.664149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-17T10:09:46.447508Z digest=sha256:2c237efc3965cf4798cf151549fac9c98dc7150d0708894256853c9525083800

Observation 563dea3e-45a6-4451-a552-83786cce7ae1 · inbound

DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models cites this paper.

DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models CogVLM: Visual Expert for Pretrained Language Models

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-15T15:46:06.589810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-12T19:22:35.305220Z digest=sha256:3b2ff1a35af98a411abf328c3ca8e58f1d090a62dd231d95a05af468971f49f0

Observation 74c5dc15-11c8-4ce0-aa77-d1029d41eb18 · inbound

Scaling Rectified Flow Transformers for High-Resolution Image Synthesis cites this paper.

Scaling Rectified Flow Transformers for High-Resolution Image Synthesis CogVLM: Visual Expert for Pretrained Language Models

Reference 192

Resolution
verified exact
arxiv_id, observed 2026-05-15T15:46:06.589810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-12T08:27:53.446686Z digest=sha256:eac3f96ae3ca3aa93445d5e7c7478d34e9b333d402d37cbe2a06e34d629f9b5b

Observation c60aef8f-086a-4728-90b2-499a09754b02 · inbound

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training cites this paper.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training CogVLM: Visual Expert for Pretrained Language Models

Reference 114

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T04:09:36.305301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:ddaf1135b5efd9c71cd807b20b468ba5c717bd6f837f49238049294c9942ed2f

Observation cde6715c-4def-4681-9d01-777b8d6b1e43 · inbound

Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models cites this paper.

Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models CogVLM: Visual Expert for Pretrained Language Models

Reference 66

Resolution
verified exact
local_arxiv, observed 2026-05-17T07:44:47.542649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T07:44:47.355960Z digest=sha256:7e5b2416f24836b67132b0abdd0e09bbc92a65c817d44b3d5ef4ba02c784e38a

Observation 91d94187-9a12-40b2-9836-1c4a7b0ea5a5 · inbound

Are We on the Right Way for Evaluating Large Vision-Language Models? cites this paper.

Are We on the Right Way for Evaluating Large Vision-Language Models? CogVLM: Visual Expert for Pretrained Language Models

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-15T15:46:06.589810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-12T19:41:44.263663Z digest=sha256:59fc354d50bb0ba25f42aa0caf193dcdf926979de9726ee8fdfefd9ba7a471c1

Observation ed3cac7c-96ca-47b9-8f86-8590ce4ce26e · inbound

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites cites this paper.

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites CogVLM: Visual Expert for Pretrained Language Models

Reference 117

Resolution
verified exact
arxiv_id, observed 2026-05-15T15:46:06.589810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-12T20:58:58.849040Z digest=sha256:23e8270e9eb296ba789edfb8bb3845dc34e609a0065f46aaa5676f1040a9db33

Observation 2728c380-e409-4570-ba97-9d949b7b1da1 · inbound

Hallucination of Multimodal Large Language Models: A Survey cites this paper.

Hallucination of Multimodal Large Language Models: A Survey CogVLM: Visual Expert for Pretrained Language Models

Reference 169

Resolution
verified exact
arxiv_id, observed 2026-05-15T15:46:06.589810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-11T12:33:32.631346Z digest=sha256:e4b349d0a9290d8623fb733173cab7a26c7f5d6615f23082a95a2a3144a7a00b

Observation c02d76f9-d739-43c7-88aa-ba2ab3e72a24 · inbound

Cracking the Code of Juxtaposition: Can AI Models Understand the Humorous Contradictions cites this paper.

Cracking the Code of Juxtaposition: Can AI Models Understand the Humorous Contradictions CogVLM: Visual Expert for Pretrained Language Models

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-05-24T00:53:40.687462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-24T00:52:52.056076Z digest=sha256:07301ce030578de6e82591c95451bb7b488f475c519105be7e8ac0e77af02466

Observation 37fccd2a-a5a8-4c60-a750-8f363b616856 · inbound

LVBench: An Extreme Long Video Understanding Benchmark cites this paper.

LVBench: An Extreme Long Video Understanding Benchmark CogVLM: Visual Expert for Pretrained Language Models

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-19T11:55:30.076307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-19T11:55:30.048525Z digest=sha256:91d61fa3f549d3a5e4d86a85f74cdc208fda3acba25c422029522932fdb9d6e0

Observation 91b19249-652c-4607-900e-656d843fa259 · inbound

PaliGemma: A versatile 3B VLM for transfer cites this paper.

PaliGemma: A versatile 3B VLM for transfer CogVLM: Visual Expert for Pretrained Language Models

Reference 143

Resolution
verified exact
arxiv_id, observed 2026-05-15T15:46:06.589810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-11T13:10:19.972353Z digest=sha256:13b4616c886ac9c1055b85b19d07b9df97840bf7250dbfc6177fa566d845d804

Observation 8388e13e-22c3-4078-8a3b-d33e97485d23 · inbound

MiniCPM-V: A GPT-4V Level MLLM on Your Phone cites this paper.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone CogVLM: Visual Expert for Pretrained Language Models

Reference 102

Resolution
verified exact
arxiv_id, observed 2026-05-15T15:46:06.589810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:4a25d634dc8091125c94f2cfc37c1827e8fa32df9eca0fe1ccd0f55ad1f3e7be

Observation 7f99b24f-61a3-40f1-b2db-3c1b961f7800 · inbound

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models cites this paper.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models CogVLM: Visual Expert for Pretrained Language Models

Reference 250

Resolution
verified exact
local_arxiv, observed 2026-05-20T06:20:36.522342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:8a5b27462f5feb7c4fb376ed0b7dca8c69571f782d1e3618084356acc55ec5b9

Observation d94def45-28f3-4cc1-9511-1acdca22261a · inbound

CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer cites this paper.

CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer CogVLM: Visual Expert for Pretrained Language Models

Reference 103

Resolution
verified exact
arxiv_id, observed 2026-05-15T15:46:06.589810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-10T18:26:22.224924Z digest=sha256:4eb95925bdfee137b82ebc9f789a8e7dcd58ce3c191ef2ebdbf083e753c4168a

Observation defb3e9e-c668-49ce-91df-cbdba7e44a1d · inbound

CogVLM2: Visual Language Models for Image and Video Understanding cites this paper.

CogVLM2: Visual Language Models for Image and Video Understanding CogVLM: Visual Expert for Pretrained Language Models

Reference 78

Resolution
verified exact
local_arxiv, observed 2026-05-16T20:10:27.751296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:dfe00063975f0aad7624a744cc4217d181600c54d98df354257b9b3019c08689

Observation 724ddaa9-7e39-4b2a-adf6-984c04aa1b4f · inbound

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference cites this paper.

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference CogVLM: Visual Expert for Pretrained Language Models

Reference 74

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T15:46:06.589810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-15T14:58:32.303101Z digest=sha256:65373fddfec8a0f7516e1cabb18f2798b897509d0cd97226dcd1dd5a478174d0

Observation dedad807-91d7-4d41-ad1f-0185068474a2 · inbound

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection cites this paper.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection CogVLM: Visual Expert for Pretrained Language Models

Reference 62

Resolution
verified exact
local_arxiv, observed 2026-05-23T20:13:24.656773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:c87772906344a956ad3e4589cf4018ed784a4848530916facfa386bb44f1f968

Observation 74d654de-be48-4838-aaa6-d1f869bc91ca · inbound

VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality Documents cites this paper.

VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality Documents CogVLM: Visual Expert for Pretrained Language Models

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-16T15:37:25.839749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T15:37:25.781240Z digest=sha256:fdb660ca69fe8ea87f6cdc829f778d6301035d0a443f1e0d7e5a5ee646b2ac4f

Observation ab197ee2-0f26-4732-8964-92cdb419fb80 · inbound

LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding cites this paper.

LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding CogVLM: Visual Expert for Pretrained Language Models

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-16T13:53:33.703016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T13:53:33.585035Z digest=sha256:77a8c001ea5b1a5f96b4e78974daf9f5f05b75f741c32f98bb50035888a2ddec

Observation 61c960ff-14dc-47dc-a679-324aae4881cc · inbound

Senna: Bridging Large Vision-Language Models and End-to-End Autonomous Driving cites this paper.

Senna: Bridging Large Vision-Language Models and End-to-End Autonomous Driving CogVLM: Visual Expert for Pretrained Language Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-15T15:46:06.589810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-15T15:24:23.756052Z digest=sha256:de644ae286be747f2ecda18b97993c8f4d00b7817228fcc7e08030b703ba8142

Observation 9622adbb-a55c-4c06-a685-10946707da83 · inbound

Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models cites this paper.

Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models CogVLM: Visual Expert for Pretrained Language Models

Reference 37

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T02:48:44.950606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:3f127007d5eab23f26b2becdc39462ab19ea1335dc285982c88290472cb39370

Observation a7edf753-b562-4d43-b73e-52dc00117bd9 · inbound

Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization cites this paper.

Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization CogVLM: Visual Expert for Pretrained Language Models

Reference 100

Resolution
verified exact
local_arxiv, observed 2026-05-16T09:16:17.496406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T09:16:17.150383Z digest=sha256:23bd6421ac44f98778874ac78520d2f375e3271ffecd2423a9cf8ddfb2d67ea1

Observation 1767b3c0-e5e6-4c2c-bfd6-305adca70dc8 · inbound

AdaptAgent: Adapting Multimodal Web Agents with Few-Shot Learning from Human Demonstrations cites this paper.

AdaptAgent: Adapting Multimodal Web Agents with Few-Shot Learning from Human Demonstrations CogVLM: Visual Expert for Pretrained Language Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-12T16:27:47.920413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:27:47.920413Z digest=sha256:544e8a4e851acd20f35b5fe6826d75642d751b9747f70c809a53ebc275742f8c

Observation 5ff5cd4d-0220-46b9-8588-730cdad390dd · inbound

FoPru: Focal Pruning for Efficient Large Vision-Language Models cites this paper.

FoPru: Focal Pruning for Efficient Large Vision-Language Models CogVLM: Visual Expert for Pretrained Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T15:31:59.606386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:31:59.606386Z digest=sha256:707fa7535b9ade2644b1379e6936e56e61bfb6e27b473248425a15c6aa798a77

Observation 21a86877-63f8-4aef-98fe-69ff3cccdbea · inbound

De-biased Multimodal Electrocardiogram Analysis cites this paper.

De-biased Multimodal Electrocardiogram Analysis CogVLM: Visual Expert for Pretrained Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T14:58:10.044146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:58:10.044146Z digest=sha256:6e48b33989d60a88ed3a3013dad55103ae776b1b856021c14474d0fa30c2cb09

Observation a0a32663-3ec7-471b-90af-7515e6269bf9 · inbound

Continual SFT Matches Multimodal RLHF with Negative Supervision cites this paper.

Continual SFT Matches Multimodal RLHF with Negative Supervision CogVLM: Visual Expert for Pretrained Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T14:58:43.106019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:58:43.106019Z digest=sha256:6bba03d675a65dc399f77c960f10b14d12ce220ceee0bf5ffa6ebb7a3146515c

Observation 64d55457-6c01-4971-a84b-5ff9ad6c3059 · inbound

MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs cites this paper.

MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs CogVLM: Visual Expert for Pretrained Language Models

Reference 184

Resolution
unresolved
no resolver link, observed 2026-08-12T14:31:37.456232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:31:37.456232Z digest=sha256:d08c9e9128f9ad8fc6e75c05a0a7ca79ca6c05675b78e34ca3c298a984296ad5

Observation c9b05143-1491-4162-b3e6-62c7a56d6240 · inbound

Pathways on the Image Manifold: Image Editing via Video Generation cites this paper.

Pathways on the Image Manifold: Image Editing via Video Generation CogVLM: Visual Expert for Pretrained Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T13:03:42.587146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:03:42.587146Z digest=sha256:a6b173d48932ec604a19b9317748fb81ff015aee426c379472469034da74dd4a

Observation edea6d44-21f1-4b73-93e6-bf0f9cdef135 · inbound

Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach cites this paper.

Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach CogVLM: Visual Expert for Pretrained Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T12:42:49.268335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:42:49.268335Z digest=sha256:b2df440bd403af91f326c459155fc04f4cf0f5d375967e0bec10100cbb94ef78

Observation 585288b8-058f-4bc6-a95a-b5a3a7a589b7 · inbound

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding cites this paper.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding CogVLM: Visual Expert for Pretrained Language Models

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-12T11:19:33.850035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:19:33.850035Z digest=sha256:26a4c63cbef7c5286ff5b5b3d0ec9504e7b32d300879c6128224dc92038faa4b

Observation dfa154cd-9702-4351-b68a-78497c026d30 · inbound

Improving Medical Diagnostics with Vision-Language Models: Convex Hull-Based Uncertainty Analysis cites this paper.

Improving Medical Diagnostics with Vision-Language Models: Convex Hull-Based Uncertainty Analysis CogVLM: Visual Expert for Pretrained Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T13:46:28.198600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:46:28.198600Z digest=sha256:0d119a26d1db24484989ad26f1719caec66708bfbbef72db026b6375f59e3f5f

Observation 8332cbc5-faa8-4c16-95d2-3ed4313155c7 · inbound

Perturb and Recover: Fine-tuning for Effective Backdoor Removal from CLIP cites this paper.

Perturb and Recover: Fine-tuning for Effective Backdoor Removal from CLIP CogVLM: Visual Expert for Pretrained Language Models

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-05-23T08:17:43.774916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-23T08:16:53.133986Z digest=sha256:f674cd5beabae4f520966e4130655ff2db620815f9a13cc69e90e10b2dd68d7f

Observation 6342a59d-883b-4146-85e9-a195894c8a59 · inbound

Multi-View Incongruity Learning for Multimodal Sarcasm Detection cites this paper.

Multi-View Incongruity Learning for Multimodal Sarcasm Detection CogVLM: Visual Expert for Pretrained Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T05:06:45.305321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:06:45.305321Z digest=sha256:b7b68541e2bf7cfd9f1ffc27da664da2023161a34718363d7108ec7ab0aced0b

Observation 52e08cb8-b7fd-4948-8414-c6623e9a8449 · inbound

Beyond Pixels: Text Enhances Generalization in Real-World Image Restoration cites this paper.

Beyond Pixels: Text Enhances Generalization in Real-World Image Restoration CogVLM: Visual Expert for Pretrained Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:58.648038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:57:58.648038Z digest=sha256:0f26c9368ac637b7f5250135d4d2e6e601318a35b23a49b78198267aeddc0de4

Observation 49e0f7b4-3685-4d8d-8ba2-841c10fe3838 · inbound

PKRD-CoT: A Unified Chain-of-thought Prompting for Multi-Modal Large Language Models in Autonomous Driving cites this paper.

PKRD-CoT: A Unified Chain-of-thought Prompting for Multi-Modal Large Language Models in Autonomous Driving CogVLM: Visual Expert for Pretrained Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T23:56:46.488520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:56:46.488520Z digest=sha256:8a0a2780fbed351c8498f7063e5f970799d8dd1912442f4a2c651458003e6aa5

Observation e2caa988-2a4a-4ab2-8623-9eebfe452aa3 · inbound

Agri-LLaVA: Knowledge-Infused Large Multimodal Assistant on Agricultural Pests and Diseases cites this paper.

Agri-LLaVA: Knowledge-Infused Large Multimodal Assistant on Agricultural Pests and Diseases CogVLM: Visual Expert for Pretrained Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T23:50:37.093016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:50:37.093016Z digest=sha256:bbf97e7a918a6d6ceb0b49886f7138992e0d42e07b42901092b0f59bdf6bcb7b

Observation d4ba772b-d8ff-4974-9980-796228dceb1a · inbound

AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information? cites this paper.

AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information? CogVLM: Visual Expert for Pretrained Language Models

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-11T23:19:11.668665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:19:11.668665Z digest=sha256:2a1248f40cefbad9b73a9c1a90c621b2d6b88c88f7e30fa7110430d111b3486d

Observation 96a96334-1a74-4b8c-946b-61c0d9d3a5e4 · inbound

AdvDreamer Unveils: Are Vision-Language Models Truly Ready for Real-World 3D Variations? cites this paper.

AdvDreamer Unveils: Are Vision-Language Models Truly Ready for Real-World 3D Variations? CogVLM: Visual Expert for Pretrained Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T22:56:44.666623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:56:44.666623Z digest=sha256:c5e11a91311e42b615e51f87dec97f13af73be551b3b01186bc0c640206d694c

Observation 6146e7bc-49c0-40c8-8f60-38f6987a6fd9 · inbound

Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis cites this paper.

Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis CogVLM: Visual Expert for Pretrained Language Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-11T22:16:58.741333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:16:58.741333Z digest=sha256:9e8e80dff256bc91886c8a170096df66c625c306bab05f32fa45296b45fed823

Observation 39a38205-86a0-41d5-8e18-fbdee6d4621a · inbound

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension cites this paper.

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension CogVLM: Visual Expert for Pretrained Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T22:18:01.649362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:18:01.649362Z digest=sha256:a8478a8de482bf3437a6730d2bd6361cf43f1038a39a257863ae2d2956390824

Observation 9a50e6e9-38ca-415a-a5a2-167b8703414a · inbound

AnyDressing: Customizable Multi-Garment Virtual Dressing via Latent Diffusion Models cites this paper.

AnyDressing: Customizable Multi-Garment Virtual Dressing via Latent Diffusion Models CogVLM: Visual Expert for Pretrained Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T21:45:35.942543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:45:35.942543Z digest=sha256:3fe0e78d8d903b77272b2adabdad44081052df2c7d9488cf45c16f623ed1a695

Observation 630d2cc9-02a1-4e40-8ccb-e6e1e968d272 · inbound

VladVA: Discriminative Fine-tuning of LVLMs cites this paper.

VladVA: Discriminative Fine-tuning of LVLMs CogVLM: Visual Expert for Pretrained Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T21:29:26.508759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:29:26.508759Z digest=sha256:dd71befa167bf7e9178f4849cb26223a3a8c7461f2c3dd3864aa9245b5176bff

Observation 927b86c0-83fc-4ee1-8fa7-88eb074a9566 · inbound

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling cites this paper.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling CogVLM: Visual Expert for Pretrained Language Models

Reference 249

Resolution
verified exact
arxiv_id, observed 2026-05-15T15:46:06.589810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:1c93dce9e6535cafb4645850b8711c89a0ff823144411ecbc3eb36c71f6dfbc2

Observation 5c68940f-5fda-4306-97e0-a664b6ef4eea · inbound

RSUniVLM: A Unified Vision Language Model for Remote Sensing via Granularity-oriented Mixture of Experts cites this paper.

RSUniVLM: A Unified Vision Language Model for Remote Sensing via Granularity-oriented Mixture of Experts CogVLM: Visual Expert for Pretrained Language Models

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-11T20:32:31.756956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:32:31.756956Z digest=sha256:518686eca2593b5eec55d1162ad869d5299b509e553bd2ed9eb179fcc571843b

Observation b17676b0-df9e-4568-9076-ca336662f59a · inbound

[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs cites this paper.

[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs CogVLM: Visual Expert for Pretrained Language Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T20:23:18.691345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:23:18.691345Z digest=sha256:e6a10c4a0578a09a3d6d71f6567f0113730dba95dd6d6a78643048965bb81f27

Observation d0eaa8d3-4c83-47aa-8133-c6cca7ca74b2 · inbound

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding cites this paper.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding CogVLM: Visual Expert for Pretrained Language Models

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-11T17:00:09.024739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:00:09.024739Z digest=sha256:61788e555ba0ae9a36c0272f6fd95776254507752f4bc98c7b95ce9fd73bdb9e

Observation e68e8702-e122-46c0-9500-451de0b6bcb1 · inbound

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation cites this paper.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation CogVLM: Visual Expert for Pretrained Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T16:45:44.816068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:45:44.816068Z digest=sha256:2a634175d843f46dd49b3e2c3d798db089318919325ff80349a4eba4c64de8e5

Observation 4ea7d8b6-87ae-4ddc-ab77-d7f697758fd0 · inbound

Enhancing Nursing and Elderly Care with Large Language Models: An AI-Driven Framework cites this paper.

Enhancing Nursing and Elderly Care with Large Language Models: An AI-Driven Framework CogVLM: Visual Expert for Pretrained Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T16:35:39.968009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:35:39.968009Z digest=sha256:8b652deabbf946bd25b56c9ee4dc102dca7b3a2ab05990d52401a34dd44d2a10

Observation 877ec80c-6122-4bc2-ac7d-1dc5a236e381 · inbound

Attention-driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models without Fine-Tuning cites this paper.

Attention-driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models without Fine-Tuning CogVLM: Visual Expert for Pretrained Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T15:38:09.327899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:38:09.327899Z digest=sha256:ed604e097a06a6f0172ce99b0c473745913aef0ab81d2da39f319d32f1852ce1

Observation 2749b36c-99e9-4b98-b6ed-a938fe68861a · inbound

From 2D CAD Drawings to 3D Parametric Models: A Vision-Language Approach cites this paper.

From 2D CAD Drawings to 3D Parametric Models: A Vision-Language Approach CogVLM: Visual Expert for Pretrained Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T14:34:00.743480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:34:00.743480Z digest=sha256:2e92177fe605b69656813ba23aa6f4bee41c88a3570f1a679c1a13e1befb9c20

Observation ecb7134a-b6a2-4d1b-af30-570c73754c14 · inbound

PunchBench: Benchmarking MLLMs in Multimodal Punchline Comprehension cites this paper.

PunchBench: Benchmarking MLLMs in Multimodal Punchline Comprehension CogVLM: Visual Expert for Pretrained Language Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T14:32:24.249886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:32:24.249886Z digest=sha256:e0a44f6aaf0ae26c013e0341c7125972c1431d305c3ca292d47f6c9f719cba6b

Observation f706b0dc-398a-4973-92fb-f12b81ed17fb · inbound

Advancing Comprehensive Aesthetic Insight with Multi-Scale Text-Guided Self-Supervised Learning cites this paper.

Advancing Comprehensive Aesthetic Insight with Multi-Scale Text-Guided Self-Supervised Learning CogVLM: Visual Expert for Pretrained Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T14:29:04.408030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:29:04.408030Z digest=sha256:a7e4a4d7f8997e91cf1dd65c9c2b5725e21b0cac6bf872ad607c9b50b4507d65

Observation 4ca56343-1589-4d85-a10a-5f224ddf4070 · inbound

Rethinking Comprehensive Benchmark for Chart Understanding: A Perspective from Scientific Literature cites this paper.

Rethinking Comprehensive Benchmark for Chart Understanding: A Perspective from Scientific Literature CogVLM: Visual Expert for Pretrained Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T18:19:03.957978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T18:19:03.957978Z digest=sha256:8ca7ed2ce870db67d84438d79c5556805ef177a2611ce4b68be864f656126eb1

Observation dadbc503-a094-43af-b542-441b1a132c0b · inbound

Multi-Dimensional Insights: Benchmarking Real-World Personalization in Large Multimodal Models cites this paper.

Multi-Dimensional Insights: Benchmarking Real-World Personalization in Large Multimodal Models CogVLM: Visual Expert for Pretrained Language Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T13:58:43.716199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:58:43.716199Z digest=sha256:594672187c53edf904b41dda2f388e057a2f1f1d69db12c24cba242dd352588d

Observation 86015391-7127-4d78-811c-616777739bfe · inbound

Deploying Foundation Model Powered Agent Services: A Survey cites this paper.

Deploying Foundation Model Powered Agent Services: A Survey CogVLM: Visual Expert for Pretrained Language Models

Reference 157

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:46.253907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:46.253907Z digest=sha256:64cd576e96938bb6f5c6bcf89beafcfc1479e0540cc0626c6300da9e4e4c7e32

Observation bfe81834-379f-4856-a0b2-c8471692022d · inbound

Unlocking the Potential of Weakly Labeled Data: A Co-Evolutionary Learning Framework for Abnormality Detection and Report Generation cites this paper.

Unlocking the Potential of Weakly Labeled Data: A Co-Evolutionary Learning Framework for Abnormality Detection and Report Generation CogVLM: Visual Expert for Pretrained Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T13:03:42.330170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:03:42.330170Z digest=sha256:29d3275880eb9f967c4cebf07072ebd3d1b960ba56bf488c098ed2d26227cb58

Observation b4dce76a-0d41-4b56-a2ee-5d023895c247 · inbound

Consistency of Compositional Generalization across Multiple Levels cites this paper.

Consistency of Compositional Generalization across Multiple Levels CogVLM: Visual Expert for Pretrained Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T13:02:11.018923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:02:11.018923Z digest=sha256:f2a65d5b9c4c28a5edb08d1d14e68b3cb4ceb55f471cc628efa9a30ca2d27855

Observation 0dfd2b33-9f51-48f3-9f88-6e7ca5665c2e · inbound

HumanVBench: Probing Human-Centric Video Understanding in MLLMs with Automatically Synthesized Benchmarks cites this paper.

HumanVBench: Probing Human-Centric Video Understanding in MLLMs with Automatically Synthesized Benchmarks CogVLM: Visual Expert for Pretrained Language Models

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-05-23T07:05:29.263612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-23T07:05:08.716223Z digest=sha256:0c87312cb52a32a6cd7b122457419ee7796fb0118b4bf13076d17fef2a5b7750

Observation 5aed9335-5f5f-4087-81fb-1af4fb7cf83a · inbound

DreamFit: Garment-Centric Human Generation via a Lightweight Anything-Dressing Encoder cites this paper.

DreamFit: Garment-Centric Human Generation via a Lightweight Anything-Dressing Encoder CogVLM: Visual Expert for Pretrained Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T05:23:02.126767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:23:02.126767Z digest=sha256:dcc96d50d8ac863dc97dd9f28cb905eff42ca2b4fdb16246062b649b8732573d

Observation 252471fb-a11c-4262-80f7-2a748906c09d · inbound

MMFactory: A Universal Solution Search Engine for Vision-Language Tasks cites this paper.

MMFactory: A Universal Solution Search Engine for Vision-Language Tasks CogVLM: Visual Expert for Pretrained Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T05:08:18.392280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:08:18.392280Z digest=sha256:04ce195bc676af9a9f49867134c3ef001da76187c7e3f44f3c266b36ca08307d

Observation fdc05b06-be11-43dd-bac4-348323dcb897 · inbound

ICM-Assistant: Instruction-tuning Multimodal Large Language Models for Rule-based Explainable Image Content Moderation cites this paper.

ICM-Assistant: Instruction-tuning Multimodal Large Language Models for Rule-based Explainable Image Content Moderation CogVLM: Visual Expert for Pretrained Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T05:01:12.255236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:01:12.255236Z digest=sha256:1eb8b5b954004f292f63841b8d3f54a88c04a0b46a382c0669fff2a9b08e9dde

Observation 36c3196c-d36c-4d9c-9ba2-9839c7216a56 · inbound

RapGuard: Safeguarding Multimodal Large Language Models via Rationale-aware Defensive Prompting cites this paper.

RapGuard: Safeguarding Multimodal Large Language Models via Rationale-aware Defensive Prompting CogVLM: Visual Expert for Pretrained Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T04:29:04.330178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:29:04.330178Z digest=sha256:3f4e0de857db658ae23c8500c0982e1fa67a744acb98d08294bb9dcfdf7daf0f

Observation e4514247-2ebd-4ed2-bc74-e7053b154cfd · inbound

ViPCap: Retrieval Text-Based Visual Prompts for Lightweight Image Captioning cites this paper.

ViPCap: Retrieval Text-Based Visual Prompts for Lightweight Image Captioning CogVLM: Visual Expert for Pretrained Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T00:47:52.509718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:47:52.509718Z digest=sha256:a760f19b98ec5396244222ceb391397eb6420827631efec7aec25ec210ca92e0

Observation 1b1ceba9-f1e3-476b-9189-44c649c5b49d · inbound

ETTA: Elucidating the Design Space of Text-to-Audio Models cites this paper.

ETTA: Elucidating the Design Space of Text-to-Audio Models CogVLM: Visual Expert for Pretrained Language Models

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-11T00:45:19.903369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:45:19.903369Z digest=sha256:be230bacacf0037d0dd200252e8746b39f3eec687ce6a6da9641c893887d161b

Observation 497cdf64-dd86-408d-a9bf-31c6b0a0e194 · inbound

Is Your Text-to-Image Model Robust to Caption Noise? cites this paper.

Is Your Text-to-Image Model Robust to Caption Noise? CogVLM: Visual Expert for Pretrained Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T00:18:36.839423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:18:36.839423Z digest=sha256:55f239a0117632f479594649f77d22c0dc963656ae514a9bc50522ea2884cf56

Observation 30bed91f-ec62-4b7e-8fde-d8312b5b6016 · inbound

ST$^3$: Accelerating Multimodal Large Language Model by Spatial-Temporal Visual Token Trimming cites this paper.

ST$^3$: Accelerating Multimodal Large Language Model by Spatial-Temporal Visual Token Trimming CogVLM: Visual Expert for Pretrained Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T23:39:18.175993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:39:18.175993Z digest=sha256:7982323c1ce64f89baa0f65b8a05128531fc09235400a5ed977b5acac94ae1a9

Observation 33b1bf55-a471-450e-9439-51434881b217 · inbound

MM-MoralBench: A MultiModal Moral Evaluation Benchmark for Large Vision-Language Models cites this paper.

MM-MoralBench: A MultiModal Moral Evaluation Benchmark for Large Vision-Language Models CogVLM: Visual Expert for Pretrained Language Models

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-23T07:15:28.516075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-23T07:14:27.786918Z digest=sha256:3e450081031a22631c2409e2b3f9a9757c6b41f6758a162e49cf2e73c526431a

Observation bc2621c8-ed9f-4f68-b995-4d1f4e3459b5 · inbound

VisionReward: Fine-Grained Multi-Dimensional Human Preference Learning for Image and Video Generation cites this paper.

VisionReward: Fine-Grained Multi-Dimensional Human Preference Learning for Image and Video Generation CogVLM: Visual Expert for Pretrained Language Models

Reference 36

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T11:49:14.800686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-16T11:49:14.698249Z digest=sha256:5e891ca3862f1778c7bf06f94993605579d11561071361adfc8b3f325f2eefca

Observation 9bcf1ecd-5a72-4ea7-95d3-87d3b26dae30 · inbound

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning cites this paper.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning CogVLM: Visual Expert for Pretrained Language Models

Reference 144

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:33:26.906583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:1a42bea36362de6239640abac5b5050b2972b5351da6aa1f849ec91f522c6c24

Observation 6673995a-2184-4aa0-92c2-0e47c10a8f6f · inbound

IllusionBench+: A Large-scale and Comprehensive Benchmark for Visual Illusion Understanding in Vision-Language Models cites this paper.

IllusionBench+: A Large-scale and Comprehensive Benchmark for Visual Illusion Understanding in Vision-Language Models CogVLM: Visual Expert for Pretrained Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T22:46:41.994091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:46:41.994091Z digest=sha256:3af6a9cac33b90cf975ff81b4c31eacd4358b093bab03653959dea5bbe28f85f

Observation f57e492e-a729-44b6-ab68-593985511094 · inbound

GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models cites this paper.

GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models CogVLM: Visual Expert for Pretrained Language Models

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-10T22:32:55.600402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:32:55.600402Z digest=sha256:bde72161853faa210b171b2c8e7cce82527cbc674fd623187af02f23e42f5509

Observation 926dd163-b180-4a9b-8484-4d248918d0a2 · inbound

MoColl: Agent-Based Specific and General Model Collaboration for Image Captioning cites this paper.

MoColl: Agent-Based Specific and General Model Collaboration for Image Captioning CogVLM: Visual Expert for Pretrained Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T22:23:27.822294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:23:27.822294Z digest=sha256:e892703ec75fe5ddee0873ab50abc768c3cbd212252718fc15a58c211891fc8a

Observation dc59dfea-9a5f-4d55-acb7-7be0fffb2415 · inbound

Visual Large Language Models for Generalized and Specialized Applications cites this paper.

Visual Large Language Models for Generalized and Specialized Applications CogVLM: Visual Expert for Pretrained Language Models

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:09.260850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:09.260850Z digest=sha256:943cd725d00a0b84c2db056f728e4961209b717f40b808119daaeeab5f505a63

Observation dcb7a881-5241-43e7-84da-9e2fcf6fb8b1 · inbound

Foundations of GenIR cites this paper.

Foundations of GenIR CogVLM: Visual Expert for Pretrained Language Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-10T22:06:04.907190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:06:04.907190Z digest=sha256:9087c349033ad4de3f25bb5bcd83150118f04ba1da64f59b2f222a8c25d10a17

Observation c20b871e-8556-44df-8fee-58b08fa38ff4 · inbound

MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models cites this paper.

MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models CogVLM: Visual Expert for Pretrained Language Models

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-23T05:45:28.373708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-23T05:44:31.546843Z digest=sha256:a174b2932b849e3f9edf437c3e3e099a4565fa4b54a4cddcc493cbf2c28fbd3c

Observation 1dff1d06-f1c9-4ff2-9dd9-654a89816006 · inbound

InfiGUIAgent: A Multimodal Generalist GUI Agent with Native Reasoning and Reflection cites this paper.

InfiGUIAgent: A Multimodal Generalist GUI Agent with Native Reasoning and Reflection CogVLM: Visual Expert for Pretrained Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-10T21:33:21.243962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:33:21.243962Z digest=sha256:13fce0748f3e326cd1cc1356f4432e4c39a4e5839fd36259fade0a2d1d278665

Observation 2ad8e6e5-a33a-4a77-a1b2-694f80493fcc · inbound

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing cites this paper.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing CogVLM: Visual Expert for Pretrained Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:48.427182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:53:48.427182Z digest=sha256:bd70e685b0119e939a5d399525e49d3903ebc91aa7e7024868c7e75601b09dfe

Observation 0c7a0eb7-ceb2-47c2-a88a-fabe790fdcc5 · inbound

Zero-Shot Scene Understanding for Automatic Target Recognition Using Large Vision-Language Models cites this paper.

Zero-Shot Scene Understanding for Automatic Target Recognition Using Large Vision-Language Models CogVLM: Visual Expert for Pretrained Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-10T20:46:17.674626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:46:17.674626Z digest=sha256:e34b3cea8bd45f7a654c325b1b69543391efdcc88ed9ec180c2dbf90135896a7

Observation 08684fc7-abda-46a9-b442-365b65683e12 · inbound

LeapVAD: A Leap in Autonomous Driving via Cognitive Perception and Dual-Process Thinking cites this paper.

LeapVAD: A Leap in Autonomous Driving via Cognitive Perception and Dual-Process Thinking CogVLM: Visual Expert for Pretrained Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T20:34:11.747261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:34:11.747261Z digest=sha256:d4c2665bf2a3747e5efeb28881b2bdaca548b68f8cb6738538856b8df0a5b458

Observation d873e978-b25b-4177-9e5f-b885631f7cf5 · inbound

Parameter-Efficient Fine-Tuning for Foundation Models cites this paper.

Parameter-Efficient Fine-Tuning for Foundation Models CogVLM: Visual Expert for Pretrained Language Models

Reference 264

Resolution
unresolved
no resolver link, observed 2026-08-10T15:38:03.671026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:38:03.671026Z digest=sha256:e8487668360b4c3b21fde46527011bec2b040ecfd5f25d78216d788e250f2f6f

Observation 3c94e11a-902c-4632-b71e-6a9746b15428 · inbound

Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models cites this paper.

Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models CogVLM: Visual Expert for Pretrained Language Models

Reference 191

Resolution
unresolved
no resolver link, observed 2026-08-10T18:04:34.830855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:04:34.830855Z digest=sha256:4946d99f9fc758faa389819d53358e06ad562fbeb760ee8fb362382c8259e6d1

Observation fdede704-6ae9-4a90-97d6-a00146a2845a · inbound

A Causality-aware Paradigm for Evaluating Creativity of Multimodal Large Language Models cites this paper.

A Causality-aware Paradigm for Evaluating Creativity of Multimodal Large Language Models CogVLM: Visual Expert for Pretrained Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T14:38:24.799578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:38:24.799578Z digest=sha256:3c9da66b987889ebd1d73145baf372c54155ec18c33a4b69ba862c730d551313

Observation 50448d07-38c6-45dc-997b-dde94c40c5b2 · inbound

MME-Industry: A Cross-Industry Multimodal Evaluation Benchmark cites this paper.

MME-Industry: A Cross-Industry Multimodal Evaluation Benchmark CogVLM: Visual Expert for Pretrained Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T11:26:29.215489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:26:29.215489Z digest=sha256:acc506802ea50226e43f28653f240daf63027df573fbb542a20b27add87cae22

Observation 552eab35-d5bf-4d44-9fa3-3c97d50748a0 · inbound

Membership Inference Attacks Against Vision-Language Models cites this paper.

Membership Inference Attacks Against Vision-Language Models CogVLM: Visual Expert for Pretrained Language Models

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-10T14:01:11.094765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:01:11.094765Z digest=sha256:6d4092f2e75c0f1a616edd62d972a2c787e84a6d785ee1daac881242e1704661

Observation ae49b89e-fb72-4cfa-9e34-b119b4cf6da8 · inbound

RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs cites this paper.

RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs CogVLM: Visual Expert for Pretrained Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-09T21:39:21.810068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T21:39:21.810068Z digest=sha256:87284f91b9e5b7f4378b5b082569f0aa1e8c3979139f117ad3a776bb7227718d

Observation 255b3b4f-9f35-4f6c-8fc1-341bcef9839a · inbound

MQuant: Unleashing the Inference Potential of Multimodal Large Language Models via Full Static Quantization cites this paper.

MQuant: Unleashing the Inference Potential of Multimodal Large Language Models via Full Static Quantization CogVLM: Visual Expert for Pretrained Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-09T19:12:56.249039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:12:56.249039Z digest=sha256:280753b39b7b87e9597577eda2da7bdd07d9f31a9582f3917a1c529130a9a82c

Observation 7f630ceb-7b5a-4204-923b-fb43ffb6d6e5 · inbound

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective cites this paper.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective CogVLM: Visual Expert for Pretrained Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:39.937725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:39.937725Z digest=sha256:cf0e5b03dd9b29dc20a626b372d2a01d22089279bff3a5eb0ccc959f1a982e6a

Observation 9f8917dd-0a2a-4a17-889c-e93499148733 · inbound

SKI Models: Skeleton Induced Vision-Language Embeddings for Understanding Activities of Daily Living cites this paper.

SKI Models: Skeleton Induced Vision-Language Embeddings for Understanding Activities of Daily Living CogVLM: Visual Expert for Pretrained Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-09T04:44:58.880265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:44:58.880265Z digest=sha256:6fa8b58543b7d9d9b2e6d86a2ce1cb670378a977ebd793e116d54aa88fbf40a2

Observation 762dcc6a-bb1d-4db7-9f9b-e93c91d3ba83 · inbound

Multitwine: Multi-Object Compositing with Text and Layout Control cites this paper.

Multitwine: Multi-Object Compositing with Text and Layout Control CogVLM: Visual Expert for Pretrained Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-08T20:05:16.194518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:05:16.194518Z digest=sha256:b894c4523a1b514e884ea35fbd346d60140d32dcc4c8313a7fa0a8c40feb05fa

Observation bb2464dd-4628-4751-a494-465575e88afc · inbound

Animate Anyone 2: High-Fidelity Character Image Animation with Environment Affordance cites this paper.

Animate Anyone 2: High-Fidelity Character Image Animation with Environment Affordance CogVLM: Visual Expert for Pretrained Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-08T16:38:09.840436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:38:09.840436Z digest=sha256:8f3739f6a6d5a2775abd239bd765db4e2e5cfbd51136aac6da412b1febad76f3

Observation e0dce2cd-fd45-4e5c-8037-27e7602e7d54 · inbound

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering cites this paper.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering CogVLM: Visual Expert for Pretrained Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.325636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.325636Z digest=sha256:c5311ab4b0d8d42eeb3b384c002360414fe104516898ce84d2f394007f576a61

Observation 3e12ed85-88bc-4133-92d1-5561f8b37bc8 · inbound

When 'YES' Meets 'BUT': Can Large Models Comprehend Contradictory Humor Through Comparative Reasoning? cites this paper.

When 'YES' Meets 'BUT': Can Large Models Comprehend Contradictory Humor Through Comparative Reasoning? CogVLM: Visual Expert for Pretrained Language Models

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-05-22T22:42:13.544301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-22T22:38:35.969273Z digest=sha256:fc0e21d5d4e9edf6e4555d7cf60faf7b1f3685d51972a6ab470f1bac28354233

Observation c078eb9a-f62e-4e13-9c20-2097f5044aab · inbound

AdaMMS: Model Merging for Heterogeneous Multimodal Large Language Models with Unsupervised Coefficient Optimization cites this paper.

AdaMMS: Model Merging for Heterogeneous Multimodal Large Language Models with Unsupervised Coefficient Optimization CogVLM: Visual Expert for Pretrained Language Models

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-22T22:47:13.045306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-22T22:47:09.500229Z digest=sha256:f7ed8ff73aa56622767aab8459dc19329485cd8345d233eab9cfe12cdd63b62b

Observation 41640757-fd0e-440a-b384-afa4376651ec · inbound

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models cites this paper.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models CogVLM: Visual Expert for Pretrained Language Models

Reference 124

Resolution
verified exact
arxiv_id, observed 2026-05-15T15:46:06.589810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:6d7fa60b8a161c1faf7d7bde1dfadae73cb052bfd2ab6a75bab9c952a457b49c

Observation acb0fee3-f39e-4074-945b-701c840f2d67 · inbound

InfiGUI-R1: Advancing Multimodal GUI Agents from Reactive Actors to Deliberative Reasoners cites this paper.

InfiGUI-R1: Advancing Multimodal GUI Agents from Reactive Actors to Deliberative Reasoners CogVLM: Visual Expert for Pretrained Language Models

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-05-18T13:54:44.270579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-18T13:54:44.011048Z digest=sha256:bfca0e74beec2b9d0d7e39680795e32c07525d34589746db7c3dbf309383e402

Observation 68b6a288-7828-4199-8f15-ee9f452b8a6a · inbound

From Street View to Visual Network: Mapping the Visibility of Urban Landmarks with Vision-Language Models cites this paper.

From Street View to Visual Network: Mapping the Visibility of Urban Landmarks with Vision-Language Models CogVLM: Visual Expert for Pretrained Language Models

Reference 13

Resolution
metadata mismatch
local_arxiv, observed 2026-05-22T15:01:42.194759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-22T15:00:45.315015Z digest=sha256:de391bbf5b5a187689ab48bb442128b35edce6a9e87681ab09e5a04a6baff3cd

Observation 6e0f3a88-6f01-436d-83d3-2e7ed9ae359b · inbound

Efficient Multi-modal Long Context Learning for Training-free Adaptation cites this paper.

Efficient Multi-modal Long Context Learning for Training-free Adaptation CogVLM: Visual Expert for Pretrained Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:11:34.762926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:11:34.762926Z digest=sha256:da0b65769119baaafae88e525c20d2413f8f639f728cadcd497188cdc82cf462