Pith. sign in

Paper Citation Record · LEDGER

CogVLM2: Visual Language Models for Image and Video Understanding

As of 12 August 2026, this Paper Citation Record lists 94 of 94 outbound references and 100 inbound Pith citation observations for arXiv:2408.16500.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2408.16500 v1

Coverage vector

measured 94 of 94 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-16T20:10:27.633010Z

measured 194 of 194 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 100 of 100 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T14:51:13.200029Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T05:34:32.402425Z

Reference resolution

94 of 94 outbound references displayed

  • verified exact28
  • verified fuzzy29
  • unresolved35
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9a54e822-4ab9-45b7-a39e-79850111f800 · outbound

This paper cites Acharya, K.

CogVLM2: Visual Language Models for Image and Video Understanding Acharya, K

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T20:10:27.960763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:2ed6ea477087fe42e1448a0f926cc28ffb5b8a3cc0f6d7526132fe39567130be

Observation f519b59e-a69c-4472-b0bb-d9b17202b92b · outbound

This paper cites GPT-4 Technical Report.

CogVLM2: Visual Language Models for Image and Video Understanding GPT-4 Technical Report

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-16T20:10:27.836941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:c70855c41d547c9dff6f47d343d14cb103091b93d8f73f710841f0e6e17b0b9a

Observation 12e69b45-834a-484c-aabc-ec199a801209 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

CogVLM2: Visual Language Models for Image and Video Understanding An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-16T20:10:27.784533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:094b9254d6270e1614df5e9253a20d8c9b3b4e7311351dccdf719e24c1a28d4b

Observation fac6f531-5b74-4cec-bd42-bba3ee2772cb · outbound

This paper cites Antol, A.

CogVLM2: Visual Language Models for Image and Video Understanding Antol, A

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T20:10:27.966570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:72017c03c90ee048b479ee4618448cd0d47dbe5d6ea015edc6b57981203a8e62

Observation f42fe827-7dae-465d-b5df-dddf0df14e2f · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

CogVLM2: Visual Language Models for Image and Video Understanding Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-16T20:10:27.696978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:ddea808275af20134905c95ea66d263a5fef56cba3bd3f102d3b6033418a40c0

Observation db0ce25f-ab37-496b-971a-64edebd88ffb · outbound

This paper cites an unresolved cited work.

CogVLM2: Visual Language Models for Image and Video Understanding Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-05-16T20:10:27.969757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:c950183414f66adcff2fa2cbe7c029cb62e93d1a059d3e4ff862feccbb4375f9

Observation 399385a3-d04a-451c-a72f-118c1319ef66 · outbound

This paper cites Nougat: Neural Optical Understanding for Academic Documents.

CogVLM2: Visual Language Models for Image and Video Understanding Nougat: Neural Optical Understanding for Academic Documents

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-16T20:10:27.832094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:a637365bf95e4001382f143662a5fda42e431f37cd3f8009fa179d6cd7e691bd

Observation 46bc0e43-3c59-4e0a-b2be-c0b73962a863 · outbound

This paper cites Byeon, B.

CogVLM2: Visual Language Models for Image and Video Understanding Byeon, B

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T20:10:27.972978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:12715496200971c2197d9322cc2c06176ad98743449d8345560817545d8c18c3

Observation b7f05ec2-497f-48cd-866b-68e5f0bb5ddc · outbound

This paper cites an unresolved cited work.

CogVLM2: Visual Language Models for Image and Video Understanding Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-05-16T20:10:27.975701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:2d6dea5a680d2db186efed204fae0cfb6cd640b50a19a85f8779c4e215f8999b

Observation be43286c-aaa6-4f05-ac68-2b86a40dd27a · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

CogVLM2: Visual Language Models for Image and Video Understanding Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-16T20:10:27.718078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:ca4e6b5d1aca4b51bc53af1c935abd1f83e74429552dee782460ab0b2c011e29

Observation adafa0e3-6f81-4449-ad51-6dc21c67c500 · outbound

This paper cites ShareGPT4Video: Improving Video Understanding and Generation with Better Captions.

CogVLM2: Visual Language Models for Image and Video Understanding ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-16T20:10:27.726864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:d224e4f629e76b1826d9b2f7960e074f8726fe52ea317a3ce3bc3b4c7d3f3d8a

Observation 62580b45-2181-4231-958f-393f3d04b395 · outbound

This paper cites PaLI: A Jointly-Scaled Multilingual Language-Image Model.

CogVLM2: Visual Language Models for Image and Video Understanding PaLI: A Jointly-Scaled Multilingual Language-Image Model

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-16T20:10:27.740978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:66bcacbbebc3708e961c42185b392cc2ba22f7f24b48ae5854075a9428aa893a

Observation 51873788-83fa-429c-b8c5-18d16cb5aa6c · outbound

This paper cites InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks.

CogVLM2: Visual Language Models for Image and Video Understanding InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-16T20:10:27.762682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:9b5c04c3ed7dd1087f8ebef9af34feb222b87b151428593519664c363455c2d8

Observation 8a0fe6c1-7d03-4073-8176-4821a424e7e4 · outbound

This paper cites an unresolved cited work.

CogVLM2: Visual Language Models for Image and Video Understanding Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-05-16T20:10:27.978720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:205df7255b5233237cf1dc4466ef7e48a9687712595b1d0833adb9edcc027cbf

Observation 4949e172-7d11-4456-b104-8a87ba7f15eb · outbound

This paper cites G-LLaVA: Solving Geometric Problem with Multi-Modal Large Language Model.

CogVLM2: Visual Language Models for Image and Video Understanding G-LLaVA: Solving Geometric Problem with Multi-Modal Large Language Model

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-16T20:10:27.800589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:aa868a2df894c45d8d794778dc1aa57be78092bf71c8379e197a3006affdeb7a

Observation d2271787-305b-4483-ac57-d267a4f73ea6 · outbound

This paper cites ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools.

CogVLM2: Visual Language Models for Image and Video Understanding ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-16T20:10:27.806236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:a260befefc3809dc6eb3bc39c41d7313791fa1a73c8657a811177723400836ad

Observation 37a4fbc5-5967-4bd6-9e07-fbfb8378c8bf · outbound

This paper cites something something.

CogVLM2: Visual Language Models for Image and Video Understanding something something

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T20:10:27.981730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:63b61a51a11ff95a3f87c763464b2dcead1c2c00d08e418e9bd67ac8ca2a227e

Observation 1b6a4e1b-ab6b-4c72-866b-0d04e57b18b4 · outbound

This paper cites Grauman, A.

CogVLM2: Visual Language Models for Image and Video Understanding Grauman, A

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T20:10:27.984758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:ce992aced17576196ce4081d3f1ba3672799c8035146e7f3561c3bf146e1c894

Observation 0a3bfad9-14b4-4706-9b86-3de7b1caf828 · outbound

This paper cites an unresolved cited work.

CogVLM2: Visual Language Models for Image and Video Understanding Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-05-16T20:10:27.987696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:ad82ceec249c2a307b6429b7a91079e96d78127d3c81b869f91c6fa5afc6d851

Observation 11bb57e1-7c3f-4c79-845c-c6eff8d4c976 · outbound

This paper cites an unresolved cited work.

CogVLM2: Visual Language Models for Image and Video Understanding Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-05-16T20:10:27.991068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:5c3bf4df15d0ddf9a774c0209a3a6fd4925c8a5ceb622d2418751f63782855a3

Observation 48d09d39-17d0-4168-9d62-508ad1e94839 · outbound

This paper cites Kafle, S.

CogVLM2: Visual Language Models for Image and Video Understanding Kafle, S

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T20:10:27.994523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:d5cb7143a70d982b521e2d971bc3f351fdde607f54b7e3827ae28d8611606ae4

Observation 0bd75f35-dd08-4032-b41e-c260c900e94f · outbound

This paper cites Kafle and C.

CogVLM2: Visual Language Models for Image and Video Understanding Kafle and C

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T20:10:27.997490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:df645b29259434775326855d7ff1911a16c8d845045d61f9e04ba6f81d54034c

Observation afefd87b-f583-4aab-9f92-0a9981b0e610 · outbound

This paper cites an unresolved cited work.

CogVLM2: Visual Language Models for Image and Video Understanding Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-05-16T20:10:28.000979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:29eb9e21ab7fb5384074144ade133e24a009b7b2be6a7f88c52177439d8feaac

Observation 5d35e18a-09b7-4db3-bfd2-bc7f5ee8c95b · outbound

This paper cites GeomVerse: A Systematic Evaluation of Large Models for Geometric Reasoning.

CogVLM2: Visual Language Models for Image and Video Understanding GeomVerse: A Systematic Evaluation of Large Models for Geometric Reasoning

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-16T20:10:27.768097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:b1567938a8d2752504ecf6e3a7029dcdf3aba7465322f689e316a40267786073

Observation e1baaf04-8de4-4025-a39c-71e9b8e2ab62 · outbound

This paper cites Kembhavi, M.

CogVLM2: Visual Language Models for Image and Video Understanding Kembhavi, M

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T20:10:28.004022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:8119c37c10e656376064ff4657ba58df55aa852aa7753e1006f5b3ff404eadcd

Observation e26de3c2-270c-4c6a-86b2-75d766afb424 · outbound

This paper cites Kembhavi, M.

CogVLM2: Visual Language Models for Image and Video Understanding Kembhavi, M

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T20:10:28.007065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:dba98da914c9e80eb2b1fe06223101eb6df677c5980c04f1e0a8cbc641852306

Observation da9fac4f-310f-4c21-8981-f7ab8b46e639 · outbound

This paper cites Kembhavi, M.

CogVLM2: Visual Language Models for Image and Video Understanding Kembhavi, M

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T20:10:28.009882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:31512fb754dcecc1a7a3ea88c29335fbcdc82cfc26f7b21e7238ec9adb27e372

Observation 89f11a57-515c-414d-9240-cf1375e20207 · outbound

This paper cites an unresolved cited work.

CogVLM2: Visual Language Models for Image and Video Understanding Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-05-16T20:10:28.012864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:683238dcff67a399a2d27f9aacd745690562cc2c5990a2073f17e67852a60c92

Observation a7a21ed4-e12f-4fc1-965f-0791d835de4e · outbound

This paper cites Krishna, Y.

CogVLM2: Visual Language Models for Image and Video Understanding Krishna, Y

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T20:10:28.015504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:bee1fe9dcb8982ab3a0ed07c0ac2e2ac859f73fcb4af61a422982af302181a0b

Observation 9885d6ee-3a90-4f91-a337-566fd07ff03e · outbound

This paper cites an unresolved cited work.

CogVLM2: Visual Language Models for Image and Video Understanding Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-05-16T20:10:28.018224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:aff416ea38a04f9b8bf5d15d05afa702c900e0733f6d1d022def7c9b4b75995f

Observation 4607e1dd-2765-48e0-a96f-340c3d8c9cf5 · outbound

This paper cites PP-OCRv3: More Attempts for the Improvement of Ultra Lightweight OCR System.

CogVLM2: Visual Language Models for Image and Video Understanding PP-OCRv3: More Attempts for the Improvement of Ultra Lightweight OCR System

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-16T20:10:27.712122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:9643928316965659d79c86c48159ce5975508d4feda24cfa017e6a05efa02b12

Observation e0d47ba8-940f-4113-95a6-4d73cdb8ecbc · outbound

This paper cites an unresolved cited work.

CogVLM2: Visual Language Models for Image and Video Understanding Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-05-16T20:10:28.020838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:12d28465b696b31c144a1547b20c1b0d8539cf4e3ce1ba5c023a1e26da589602

Observation 0652cd27-6b42-445f-8066-cb5879896a28 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

CogVLM2: Visual Language Models for Image and Video Understanding VideoChat: Chat-Centric Video Understanding

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-16T20:10:27.722354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:231e8f82e6b4b200bb04de4490fa2cac99bd1e4a912c7f0459b943c39401938a

Observation 08d38f54-f572-4214-ba13-000243b31a84 · outbound

This paper cites an unresolved cited work.

CogVLM2: Visual Language Models for Image and Video Understanding Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-05-16T20:10:28.023954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:8c9832c82e2d20a357fc27903d13975359d0a09bf0c357bc4d73a32223cf9c41

Observation 215434de-9888-454e-8479-a6ce6fd57807 · outbound

This paper cites UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer.

CogVLM2: Visual Language Models for Image and Video Understanding UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-16T20:10:27.735657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:8d63e8b41ccffb5ed879f6613908fbf3a64569656703f19fe251980a72b897d0

Observation 4752dbaa-481f-4104-8ffd-fc14d2ab159a · outbound

This paper cites an unresolved cited work.

CogVLM2: Visual Language Models for Image and Video Understanding Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-05-16T20:10:28.026867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:3fdf155ec3076d00dd57085c40effa39993e33ef11514af3d9b0eadd7a86befe

Observation b6a8c910-c016-4550-a1a3-c16656d56357 · outbound

This paper cites Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models.

CogVLM2: Visual Language Models for Image and Video Understanding Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-17T07:44:47.683285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:e494324cdb570c43dab53f89157a047bb29e5035c090de0e9534c66686713cd9

Observation b56541ca-3f59-498b-8d37-80c9589f0766 · outbound

This paper cites an unresolved cited work.

CogVLM2: Visual Language Models for Image and Video Understanding Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-05-16T20:10:28.029941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:9efcdf74263ee412a3ab4b13cea1549f228354bece983eee77f723ce85e3aad1

Observation 933aeecf-ff3e-412e-a5d0-0bcdedb3c0d7 · outbound

This paper cites an unresolved cited work.

CogVLM2: Visual Language Models for Image and Video Understanding Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-05-16T20:10:28.033478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:88ea0f7f4ba3a080bfd3976687768e676f9e397877875bc63689be40de83aa23

Observation 9628919e-6ccf-4368-99b2-e7c76070419a · outbound

This paper cites an unresolved cited work.

CogVLM2: Visual Language Models for Image and Video Understanding Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-05-16T20:10:28.037335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:7dc804a3d6c90cf9a4b61cdcd7a3d35b2da4d2558bf96fe3207b2de380a7d662

Observation 006b469e-8057-48f0-993d-12705df4fd29 · outbound

This paper cites an unresolved cited work.

CogVLM2: Visual Language Models for Image and Video Understanding Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-05-16T20:10:28.041078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:ce039f2dbc30871fe6840ddf4506fa3608a0180480418c68fe4b3c4a92ef3127

Observation fba93d6d-aade-463d-ac2e-5d114150ca90 · outbound

This paper cites an unresolved cited work.

CogVLM2: Visual Language Models for Image and Video Understanding Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-05-16T20:10:28.044513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:78557f24ec5cb70d6df769b1bb04ae644ff6335052a7cf2c1ec5c3d287c9a017

Observation ed258d1e-0ec7-43b8-9f8e-ec97490736e6 · outbound

This paper cites ST-LLM: Large Language Models Are Effective Temporal Learners.

CogVLM2: Visual Language Models for Image and Video Understanding ST-LLM: Large Language Models Are Effective Temporal Learners

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-16T20:10:27.812133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:bd8f8e82b7208afbc0a487864bcc5f78363aaa23ef913ee3495e8ac3d3514a20

Observation d22a74fe-19eb-4142-9be5-aade3379bff5 · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

CogVLM2: Visual Language Models for Image and Video Understanding MMBench: Is Your Multi-modal Model an All-around Player?

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-05-16T20:10:27.822433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:36bf67f9b86c4b598f5445cdef0298d1d2bf38c6cb73ea34bf003989f63f9234

Observation dda4b2f0-c538-47a5-bac8-b6bf4feddb93 · outbound

This paper cites an unresolved cited work.

CogVLM2: Visual Language Models for Image and Video Understanding Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-05-16T20:10:28.047689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:8d079446b273bde13693a55892f8f282261fcc04734bb1d6f94f9f4a315a2e2a

Observation 3fadd1fa-3df3-4adc-a7d5-5021ba705c83 · outbound

This paper cites an unresolved cited work.

CogVLM2: Visual Language Models for Image and Video Understanding Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-05-16T20:10:28.050733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:3645564e0e0031a21029ab618c5cebf766b0d8f22280179fc736efc9f2d53544

Observation 8378d1cc-3b1a-4cea-825a-131a12ef1827 · outbound

This paper cites an unresolved cited work.

CogVLM2: Visual Language Models for Image and Video Understanding Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-05-16T20:10:27.845245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:7183e4f1bfeb900fb42946278c3cce9b613212c60833f81257975717d9bbb3c6

Observation 8c5f4bca-301a-409b-b94e-fbcf5869b7da · outbound

This paper cites an unresolved cited work.

CogVLM2: Visual Language Models for Image and Video Understanding Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-05-16T20:10:27.848655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:3662023eec0a8404ba358b4ee831de2dc299a6a9478dc6922bdaff0e7a1052fd

Observation 2e7da8ad-e0e9-4990-b680-b7a544d73a61 · outbound

This paper cites an unresolved cited work.

CogVLM2: Visual Language Models for Image and Video Understanding Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-05-16T20:10:27.852603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:84cc0786fa33a82bd0eeb87ba1050043dd649c36a81181b23b6cb6c6d1a1a448

Observation bc8c901a-6038-4a53-bf88-5c8e21cfa6ee · outbound

This paper cites an unresolved cited work.

CogVLM2: Visual Language Models for Image and Video Understanding Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-05-16T20:10:27.856463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:38e865be2fcb7a899e8952332173ea0a0e1b17a8e7e8902ec72ad77dbf75d527

Observation 67d4e8bc-6a3f-4ec2-976b-43b6e236dfc6 · outbound

This paper cites an unresolved cited work.

CogVLM2: Visual Language Models for Image and Video Understanding Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-05-16T20:10:27.860240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:fcdd24bad1ea5407e507ff58dfc774b12e6ed6c9394f42a3e866110d316f1b22

Observation 83c4514c-9187-4a43-8e84-18741ebd59e5 · outbound

This paper cites an unresolved cited work.

CogVLM2: Visual Language Models for Image and Video Understanding Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-05-16T20:10:27.863551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:2b06e7c738ec59b07a8a3a7e096a2a717d1e743b1807e49c0e01129ccf465119

Observation e8f9307d-1efa-4220-8899-e3a489d2c08b · outbound

This paper cites an unresolved cited work.

CogVLM2: Visual Language Models for Image and Video Understanding Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-05-16T20:10:27.867046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:d793dd02ef23134326b069c27fa1c8ac96575e25a62aa8d03c36ea1a7f6b5c62

Observation 49ecbadf-39ca-4f01-ae6c-126297641938 · outbound

This paper cites Marino, M.

CogVLM2: Visual Language Models for Image and Video Understanding Marino, M

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T20:10:27.870344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:07e1a0a6a18b6086aa792a4c91530bea9d2b3712a89cc88548ba69985668e5a3

Observation ab2f312b-a3dd-4ac5-8edb-6a4dcf3ce4a3 · outbound

This paper cites Marti and H.

CogVLM2: Visual Language Models for Image and Video Understanding Marti and H

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T20:10:27.873786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:c26271b57e1425e44f521af3eee5945d9dd6bbcd8a659d4770d7ef8f5b7fcd24

Observation d620c82a-f797-4f38-99b7-ea1df47a204a · outbound

This paper cites Masry, D.

CogVLM2: Visual Language Models for Image and Video Understanding Masry, D

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T20:10:27.877201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:7942f63ed78f71cd80d26fe9d6d7bdb3bcb5384d3e162734cc2fd5c6eeb05833

Observation c0ac1a4a-60ca-4b97-b35b-6498f7642f97 · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

CogVLM2: Visual Language Models for Image and Video Understanding ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 58

Resolution
verified exact
local_arxiv, observed 2026-05-16T20:10:27.773421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:d0cb84b697095eb7534fde16760d7c039b5f555e422eef1742de3b3c873c2d69

Observation 2be31272-5224-4446-b428-9ab5ec0606c1 · outbound

This paper cites Masry, D.

CogVLM2: Visual Language Models for Image and Video Understanding Masry, D

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T20:10:27.880648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:adc36e6a91673cc9a09a493671fe6fdf79b813d2c48a4b5f7e5cc863c539a591

Observation c93e4f5f-46a4-48ae-a803-e880fa3cccdf · outbound

This paper cites Mathew, V.

CogVLM2: Visual Language Models for Image and Video Understanding Mathew, V

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T20:10:27.884006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:cf36b11cf987e0f5a39755ffb590c594955ebf40544760608b7ba4273fbc4847

Observation 5eb98273-7f29-4c13-902c-a3e918dcdebc · outbound

This paper cites Mathew, D.

CogVLM2: Visual Language Models for Image and Video Understanding Mathew, D

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T20:10:27.887727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:c2eaadc522701d18998b62a1e5390fe5cc1b27506455446bef8ca225bbc81e1b

Observation 60611023-95f5-44f8-a7f2-64c1d6a7e560 · outbound

This paper cites Mathew, D.

CogVLM2: Visual Language Models for Image and Video Understanding Mathew, D

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T20:10:27.890525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:5fa43d108613dbe9a890cb5d179032f16d0b7d4fc3ec49559b5f20dcae579739

Observation 4bc8b320-4139-4516-970f-837028af3720 · outbound

This paper cites Mathew, D.

CogVLM2: Visual Language Models for Image and Video Understanding Mathew, D

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T20:10:27.893362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:8601430b00195d8e6f53ca1cc62151502e5b540ad093ebc0722af5ec43f0fe96

Observation 105879cc-09d2-4e7c-8ba8-8b74d15ede99 · outbound

This paper cites Mishra, S.

CogVLM2: Visual Language Models for Image and Video Understanding Mishra, S

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T20:10:27.896528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:56ff4399c892f0b3002500a3ba3994f5895fde6a178a1859ac59f21a825cee47

Observation 1da3e238-20a2-4283-8d52-77202e7b22ee · outbound

This paper cites an unresolved cited work.

CogVLM2: Visual Language Models for Image and Video Understanding Unresolved cited work

Reference 65

Resolution
parse uncertain
raw_fallback, observed 2026-05-16T20:10:27.899897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:fb55dc3beda387e7c2edecde1042884aa0dd45db1a9a607749e5c96d91feaf57

Observation 67e47033-07ba-4894-852d-4d263c54316e · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

CogVLM2: Visual Language Models for Image and Video Understanding Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 66

Resolution
verified exact
local_arxiv, observed 2026-05-16T20:10:27.841615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:a40dc3f27c1d8b2116d495ce33d4a1f3735e8a0c53d5de41d2c0c939d73dd1c2

Observation ca324a57-d643-4590-a8bb-7f477c6345fe · outbound

This paper cites TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding.

CogVLM2: Visual Language Models for Image and Video Understanding TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-16T20:10:27.687282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:d5f44bcfc53bb72eeb48d1404c3bd1c9b3067ab4fb096594827e789a36194de1

Observation 4ee1ecd9-8719-4618-9e4e-fa49f8b6f344 · outbound

This paper cites Schuhmann, R.

CogVLM2: Visual Language Models for Image and Video Understanding Schuhmann, R

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T20:10:27.903327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:9366c210320f4e14d223b1527253de7e33e627156cf38ec95db948c091765f70

Observation 431a3330-2eca-4b36-a522-366b88b7377b · outbound

This paper cites LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs.

CogVLM2: Visual Language Models for Image and Video Understanding LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs

Reference 69

Resolution
verified exact
local_arxiv, observed 2026-05-16T20:10:27.705585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:84b20947f40c70af9ef96c09ef509e56d538fa7c2271f9e122015b11e24539ad

Observation 010ea674-078b-429a-8d10-5aa2a183eb75 · outbound

This paper cites Schwenk, A.

CogVLM2: Visual Language Models for Image and Video Understanding Schwenk, A

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T20:10:27.906584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:0c8a682d1a2006fc54f5e8382bd4b1d8d534a8e6b951e73ed1cce9601e2b4119

Observation 6b4bff8d-8b16-4a8c-a3a9-be4adb8b7413 · outbound

This paper cites an unresolved cited work.

CogVLM2: Visual Language Models for Image and Video Understanding Unresolved cited work

Reference 71

Resolution
unresolved
raw_fallback, observed 2026-05-16T20:10:27.909919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:d4e5e17027702342fb196c618c088fab98acdcb7b6296509914994bbb4a5d0a7

Observation de1d903c-78f6-49bc-a3f4-57328560d13a · outbound

This paper cites Singh, V.

CogVLM2: Visual Language Models for Image and Video Understanding Singh, V

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T20:10:27.913661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:d3e7c4bb448743cd0f06912b4957bcb6e001b36e0d17f3fb0a956f277a0223c6

Observation d58a588d-378d-4b28-b029-c1e75ef45aa2 · outbound

This paper cites Singh, V.

CogVLM2: Visual Language Models for Image and Video Understanding Singh, V

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T20:10:27.917169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:0c2af019cb1cc983aa0821e70636046df8d0bda78661596c4c1f3e7dd606724a

Observation a75f5f2b-8ca2-4d28-8161-68348ad4b46a · outbound

This paper cites EVA-CLIP: Improved Training Techniques for CLIP at Scale.

CogVLM2: Visual Language Models for Image and Video Understanding EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 74

Resolution
verified exact
local_arxiv, observed 2026-05-16T20:10:27.731179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:e558e291112e182530fcec558cbdceed249b16a9b886d06a67e378e5432d6f0f

Observation aa30b319-8c08-4bd8-b180-873ab01e7054 · outbound

This paper cites an unresolved cited work.

CogVLM2: Visual Language Models for Image and Video Understanding Unresolved cited work

Reference 75

Resolution
unresolved
raw_fallback, observed 2026-05-16T20:10:27.920352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:5682bdfccfabf9478e9fbafc5ad886c983b65355468473bde9f1faeff92a71a7

Observation 8acc3d6d-4f6e-4ccb-88be-df59bc93c15b · outbound

This paper cites an unresolved cited work.

CogVLM2: Visual Language Models for Image and Video Understanding Unresolved cited work

Reference 76

Resolution
unresolved
raw_fallback, observed 2026-05-16T20:10:27.923301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:6f50e7522eff450f8837b7f806c305e3c5e79fbdf2d5b1a037dcf1af6187705e

Observation a09a5b4c-0772-46d4-9811-b8b2837694c9 · outbound

This paper cites LVBench: An Extreme Long Video Understanding Benchmark.

CogVLM2: Visual Language Models for Image and Video Understanding LVBench: An Extreme Long Video Understanding Benchmark

Reference 77

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:55:30.239348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:7dbe202b1a2a7a5832aebb5175945d8dd39f093de57a897f51bab7323a0bc06c

Observation defb3e9e-c668-49ce-91df-cbdba7e44a1d · outbound

This paper cites CogVLM: Visual Expert for Pretrained Language Models.

CogVLM2: Visual Language Models for Image and Video Understanding CogVLM: Visual Expert for Pretrained Language Models

Reference 78

Resolution
verified exact
local_arxiv, observed 2026-05-16T20:10:27.751296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:dfe00063975f0aad7624a744cc4217d181600c54d98df354257b9b3019c08689

Observation 8c54c02a-278b-4aa7-a14d-fae33148d94d · outbound

This paper cites an unresolved cited work.

CogVLM2: Visual Language Models for Image and Video Understanding Unresolved cited work

Reference 79

Resolution
unresolved
raw_fallback, observed 2026-05-16T20:10:27.926322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:0809445c5e0563a7fe257456c0db1c32e7fbbe430f25ca6f06f0d7168a4f8dc8

Observation 38dd3c32-185c-47ea-9826-8967c0e5108a · outbound

This paper cites an unresolved cited work.

CogVLM2: Visual Language Models for Image and Video Understanding Unresolved cited work

Reference 80

Resolution
unresolved
raw_fallback, observed 2026-05-16T20:10:27.929608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:171ebfdd3f3952b1d1a00a2739688367c7e5cc95ef8dbab201a4537b390333f0

Observation d2d8a7ec-ac15-4a91-8c1d-9bbc1e9bd6b4 · outbound

This paper cites an unresolved cited work.

CogVLM2: Visual Language Models for Image and Video Understanding Unresolved cited work

Reference 81

Resolution
unresolved
raw_fallback, observed 2026-05-16T20:10:27.932605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:96967eeca7f2e3c42fa9e472f728e3cfff8b64683196853f2bc7d85d8768ad5c

Observation 5e0aa4f6-4677-4cb0-bb0a-a5f35afd971d · outbound

This paper cites an unresolved cited work.

CogVLM2: Visual Language Models for Image and Video Understanding Unresolved cited work

Reference 82

Resolution
unresolved
raw_fallback, observed 2026-05-16T20:10:27.936695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:d60173d023cf595540159bf2fac2e24f991ceb1515d6b52cdeca25919ddf02e6

Observation cf4a475c-3692-42f6-8ae1-c03796870c41 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

CogVLM2: Visual Language Models for Image and Video Understanding MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 83

Resolution
verified exact
local_arxiv, observed 2026-05-16T20:10:27.778951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:d58eb08b7b27af9fd36786b0912ef0b2efe1100ad79d26e1354314cb1871f056

Observation 7bfaa28d-75c4-4274-8a50-bab90447ddc4 · outbound

This paper cites an unresolved cited work.

CogVLM2: Visual Language Models for Image and Video Understanding Unresolved cited work

Reference 84

Resolution
unresolved
raw_fallback, observed 2026-05-16T20:10:27.939735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:99d931ea754ed43057d9e44f91d063d7494bb04935f30d798ec259a1ef58b7d1

Observation e8073d0e-b871-4d18-8e33-d78f836ab8aa · outbound

This paper cites A Survey on Multimodal Large Language Models.

CogVLM2: Visual Language Models for Image and Video Understanding A Survey on Multimodal Large Language Models

Reference 85

Resolution
verified exact
local_arxiv, observed 2026-05-16T20:10:27.789856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:7ab037c720a58c456750aa3ef96c59fab0f64b596c5a0c7ffc0eb999c2a8a280

Observation 6e246c8e-b37c-4fd6-b63b-d90deedb94c1 · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

CogVLM2: Visual Language Models for Image and Video Understanding MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 86

Resolution
verified exact
local_arxiv, observed 2026-05-16T20:10:27.795148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:e102661f194ed2cff28f93bb7edca0fa491d298ec15922fbb877ba009bf9c15b

Observation 5f9bf804-442e-4c99-b2d8-6d45bb7631a1 · outbound

This paper cites an unresolved cited work.

CogVLM2: Visual Language Models for Image and Video Understanding Unresolved cited work

Reference 87

Resolution
unresolved
raw_fallback, observed 2026-05-16T20:10:27.942389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:2a3b6373bc958d8bd47b89c4cbe0703321603e6420ce421ca7393809a2d3e2bb

Observation d2753ffe-2fff-48c6-ad30-819a8c2e8589 · outbound

This paper cites an unresolved cited work.

CogVLM2: Visual Language Models for Image and Video Understanding Unresolved cited work

Reference 88

Resolution
unresolved
raw_fallback, observed 2026-05-16T20:10:27.945129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:e58daf441b6a89dc828bdb279ce59668607899a6cf20a87770f5f7117f1ce56d

Observation 1e3d03fb-344b-4f66-8db2-158f0e57127a · outbound

This paper cites Zhang, F.

CogVLM2: Visual Language Models for Image and Video Understanding Zhang, F

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T20:10:27.947867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:ae101d2ea22d6adfde92f06bc4c466a6a45b2bacb0e49604a30901eeaac8fe50

Observation 2027a0c2-8aa6-414c-8543-c10731b1bfbc · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

CogVLM2: Visual Language Models for Image and Video Understanding Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 90

Resolution
verified exact
local_arxiv, observed 2026-05-16T20:10:27.817218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:64ca5925f22e8b80743ed7fbdca64e566e2741552135f0ca2272b80dd9b5e999

Observation 4eef3165-d815-461c-9d58-baed95544be4 · outbound

This paper cites Zhang, P.

CogVLM2: Visual Language Models for Image and Video Understanding Zhang, P

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T20:10:27.950812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:2ed8e07183bd158fb3662fd6cf33e380fad75d5f2b20ca89ce046d0077930541

Observation 7a856c60-8a1e-498a-be79-9ab924c70cf5 · outbound

This paper cites VCR: A Task for Pixel-Level Complex Reasoning in Vision Language Models via Restoring Occluded Text.

CogVLM2: Visual Language Models for Image and Video Understanding VCR: A Task for Pixel-Level Complex Reasoning in Vision Language Models via Restoring Occluded Text

Reference 92

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T20:10:27.827632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:d6c0d1107990e4aee821edeb4669c484e2c86f6952be09734eea35dffa6e7018

Observation 426d7758-4284-4e31-a0cb-05e65d850e1f · outbound

This paper cites It features a blue and white circular design with a black border.

CogVLM2: Visual Language Models for Image and Video Understanding It features a blue and white circular design with a black border

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T20:10:27.953910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:1872313b22d68071846aa1d9a758ad13ac246040010f10132abb187ce6fd9a1e

Observation e51ea41b-eae4-428d-8390-f18a89c42103 · outbound

This paper cites It consists of a stylized letter "Q" enclosed within a circular shape.

CogVLM2: Visual Language Models for Image and Video Understanding It consists of a stylized letter "Q" enclosed within a circular shape

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T20:10:27.957571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:16d3182daacb71d93bbbcea789daaf7eae50747750b4fc3a77c67bd62203dcb0

Observation 36b71b83-e388-4da3-884a-785ad178df4d · outbound

This paper cites N/A" instead).{.

CogVLM2: Visual Language Models for Image and Video Understanding N/A" instead).{

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T20:10:27.963491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:4f38ec8ea60c02718db2bfa96ff2bbddc49dc8bd4ae1a1d2beab97a06e48e6a0

Pith citing papers

Observation ae03e68c-0b35-47c4-8a43-06f9a44a98fb · inbound

LVBench: An Extreme Long Video Understanding Benchmark cites this paper.

LVBench: An Extreme Long Video Understanding Benchmark CogVLM2: Visual Language Models for Image and Video Understanding

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-19T11:55:30.175570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-19T11:55:30.048525Z digest=sha256:81921b9cc2d05b6eb8ce71c168664b9b7608754925b55cfec13eade86ae1fbf0

Observation bfb2e576-3c5b-4b29-9add-7a0e9196be4d · inbound

CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer cites this paper.

CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer CogVLM2: Visual Language Models for Image and Video Understanding

Reference 81

Resolution
verified exact
arxiv_id, observed 2026-05-16T20:10:28.051835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-10T18:26:22.224924Z digest=sha256:e96825d4b52007cf47bdaa4982dd8416dbdbc8b957d61dea689a29d04340ddaf

Observation 3ed25625-ffcc-4338-b771-77e891f93c84 · inbound

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models cites this paper.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models CogVLM2: Visual Language Models for Image and Video Understanding

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-16T20:10:28.051835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:b72a4265ee5c998ccb40092259fa99c5843ff16d95fe116644dd855d9ca24d76

Observation eb35f914-a198-4dbf-9193-7bd2b7f97aef · inbound

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling cites this paper.

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling CogVLM2: Visual Language Models for Image and Video Understanding

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-05-23T19:43:23.734159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:1bafd34e31907f24a177a2dbad8fb7004328d2f8fef0781cff70d304f7e671cc

Observation 80bab832-fcb6-443e-9f08-8096876b0a1b · inbound

S$^4$ST: A Strong, Self-transferable, faSt, and Simple Scale Transformation for Transferable Targeted Attack cites this paper.

S$^4$ST: A Strong, Self-transferable, faSt, and Simple Scale Transformation for Transferable Targeted Attack CogVLM2: Visual Language Models for Image and Video Understanding

Reference 76

Resolution
verified exact
local_arxiv, observed 2026-05-23T19:23:21.732594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-23T19:22:48.087574Z digest=sha256:4939e545fb45ea75fca01c605f09fb038f3b821bd35368e29c2589991ec76c65

Observation 4cb90199-14d9-4d0b-b297-647d44f44a60 · inbound

ICT: Image-Object Cross-Level Trusted Intervention for Mitigating Object Hallucination in Large Vision-Language Models cites this paper.

ICT: Image-Object Cross-Level Trusted Intervention for Mitigating Object Hallucination in Large Vision-Language Models CogVLM2: Visual Language Models for Image and Video Understanding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:13.200029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:13.200029Z digest=sha256:8d126862119ff4a7d0d4ec6f93086fcb381f4b83083724bb93efebe82224a447

Observation b7e92756-c100-41ac-9710-cc5b51013b8f · inbound

All Seeds Are Not Equal: Enhancing Compositional Text-to-Image Generation with Reliable Random Seeds cites this paper.

All Seeds Are Not Equal: Enhancing Compositional Text-to-Image Generation with Reliable Random Seeds CogVLM2: Visual Language Models for Image and Video Understanding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T10:57:49.249681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:57:49.249681Z digest=sha256:744bde016eff6814c9fdef1ad90513a0f0ce009ee03d0c535c9ddce8a76c16c9

Observation 542f5c65-6096-4cf0-8097-4748a8434ca1 · inbound

OpenHumanVid: A Large-Scale High-Quality Dataset for Enhancing Human-Centric Video Generation cites this paper.

OpenHumanVid: A Large-Scale High-Quality Dataset for Enhancing Human-Centric Video Generation CogVLM2: Visual Language Models for Image and Video Understanding

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T10:44:52.508007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:44:52.508007Z digest=sha256:ee2175289c06ef5fa05d2c4e58432e1e3491f0e804ac04a8fc73875558862544

Observation a21b01a0-49ff-4cb7-b178-40b54e03990d · inbound

Mimir: Improving Video Diffusion Models for Precise Text Understanding cites this paper.

Mimir: Improving Video Diffusion Models for Precise Text Understanding CogVLM2: Visual Language Models for Image and Video Understanding

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:11.697842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:11.697842Z digest=sha256:b53cf3bbb89ced28de6afc05d8da43d9a4c9383580d26297ec1a2531b58d0bf8

Observation 8bcb7c49-8e65-4f33-9d51-52465e746d80 · inbound

SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding cites this paper.

SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding CogVLM2: Visual Language Models for Image and Video Understanding

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T21:29:49.610925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:29:49.610925Z digest=sha256:b0f7c9baeddb1ee76334ce04ac2b8f8bc68d8ff1a3c269ee60dbdc92f595d8b3

Observation 18c8ea42-2691-4d54-9308-bda11e8539d9 · inbound

p-MoD: Building Mixture-of-Depths MLLMs via Progressive Ratio Decay cites this paper.

p-MoD: Building Mixture-of-Depths MLLMs via Progressive Ratio Decay CogVLM2: Visual Language Models for Image and Video Understanding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T21:31:44.375733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:31:44.375733Z digest=sha256:756ddb0ffcef1a768d6414cd5c7a0568c3c347951ead4374a6c1cc0922b29a4f

Observation 13143aad-217b-4e29-aa85-c10d28e9c651 · inbound

Uni-NaVid: A Video-based Vision-Language-Action Model for Unifying Embodied Navigation Tasks cites this paper.

Uni-NaVid: A Video-based Vision-Language-Action Model for Unifying Embodied Navigation Tasks CogVLM2: Visual Language Models for Image and Video Understanding

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-16T20:10:28.051835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T19:51:36.137985Z digest=sha256:99c57646f19304583052a26309bc723e2090d0fd41df87aa99121e9099612a06

Observation 1b5bf52b-6207-44ad-b136-1fa758a80eb3 · inbound

MIT-10M: A Large Scale Parallel Corpus of Multilingual Image Translation cites this paper.

MIT-10M: A Large Scale Parallel Corpus of Multilingual Image Translation CogVLM2: Visual Language Models for Image and Video Understanding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:36.930018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:36.930018Z digest=sha256:04b865d7902d3584d7e7d25da0346b2b64c6d38243b367981465a5d3540d6bc9

Observation 22c75973-7a83-4169-89d2-a9460907a033 · inbound

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining cites this paper.

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining CogVLM2: Visual Language Models for Image and Video Understanding

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T16:10:28.947119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:10:28.947119Z digest=sha256:656fa0d643974b6f583af8608a1da254550ec7113e2f19edd62bea601db1c0ad

Observation dd200a4d-45e3-466b-8dc2-98ea4bf95dbf · inbound

CogNav: Cognitive Process Modeling for Object Goal Navigation with LLMs cites this paper.

CogNav: Cognitive Process Modeling for Object Goal Navigation with LLMs CogVLM2: Visual Language Models for Image and Video Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T18:07:09.698224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:07:09.698224Z digest=sha256:d2be430134ec256087b888202503f31540c68b2110c24ed67ebf19a58e38934e

Observation e8fa3bc7-3dbc-485e-8264-4a12f7377600 · inbound

VCA: Video Curious Agent for Long Video Understanding cites this paper.

VCA: Video Curious Agent for Long Video Understanding CogVLM2: Visual Language Models for Image and Video Understanding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T16:51:15.146078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:51:15.146078Z digest=sha256:e89200f8fdc4e3934793aa7379d27a02f40922158cdc69a41a92b6396001f593

Observation 3853f0ae-45c3-44b3-a537-60687707fd4d · inbound

PunchBench: Benchmarking MLLMs in Multimodal Punchline Comprehension cites this paper.

PunchBench: Benchmarking MLLMs in Multimodal Punchline Comprehension CogVLM2: Visual Language Models for Image and Video Understanding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T14:32:24.061627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:32:24.061627Z digest=sha256:e086725a2fdf399cb97cdd384ba5050b024517672f5e3132d989c00eca8ef8ca

Observation 7b3b0cee-ec53-4083-ba9e-33e21ce4bb7f · inbound

HumanVBench: Probing Human-Centric Video Understanding in MLLMs with Automatically Synthesized Benchmarks cites this paper.

HumanVBench: Probing Human-Centric Video Understanding in MLLMs with Automatically Synthesized Benchmarks CogVLM2: Visual Language Models for Image and Video Understanding

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-23T07:05:29.249376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-23T07:05:08.716223Z digest=sha256:05d1917e28a73a1e0f8d8cc04bcbccec417f14cd8da2aa6baed0709b352c552d

Observation 28fe57f7-70d8-45a3-a09d-57d8f79cf6d7 · inbound

VisionReward: Fine-Grained Multi-Dimensional Human Preference Learning for Image and Video Generation cites this paper.

VisionReward: Fine-Grained Multi-Dimensional Human Preference Learning for Image and Video Generation CogVLM2: Visual Language Models for Image and Video Understanding

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T20:10:28.051835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-16T11:49:14.698249Z digest=sha256:83b326717d0fa5e39bbfa29daaa9f9f895bd54b199c2d9cdb08c5f189d84f005

Observation e17633e3-7103-4a5c-b930-5c03bcfedcc7 · inbound

EliGen: Entity-Level Controlled Image Generation with Regional Attention cites this paper.

EliGen: Entity-Level Controlled Image Generation with Regional Attention CogVLM2: Visual Language Models for Image and Video Understanding

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-10T22:38:08.640912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:38:08.640912Z digest=sha256:839d303eacecb98ad08cae1d482315f36bbf8dbaf76e9a1d7d9b453eb62ae321

Observation d343bc22-7f2c-4a94-b0f6-9c11bb61ee7d · inbound

VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction cites this paper.

VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction CogVLM2: Visual Language Models for Image and Video Understanding

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-05-17T21:08:19.716043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T21:08:19.570050Z digest=sha256:6bf41d505f19c442c9b644af802c2d93493c62b39fc45d58da4cee1f7fcc101c

Observation c33e27e4-ab57-4af2-a9f3-b12bea3ee231 · inbound

MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models cites this paper.

MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models CogVLM2: Visual Language Models for Image and Video Understanding

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-23T05:45:28.289290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-23T05:44:31.546843Z digest=sha256:3cf917c9f7e92c457ef2d3465c77effebcfab96e064e682940ec5adf828fb29a

Observation aa05e9a9-14ce-4327-bdfd-caa8d1658aa0 · inbound

Automated Generation of Challenging Multiple-Choice Questions for Vision Language Model Evaluation cites this paper.

Automated Generation of Challenging Multiple-Choice Questions for Vision Language Model Evaluation CogVLM2: Visual Language Models for Image and Video Understanding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T21:57:14.197014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:57:14.197014Z digest=sha256:784fe4859c822ef246667a552e9191b98e1c472705c61030bf451b5989576825

Observation e0be945d-b4ad-488d-954a-f4da9bd4ba9b · inbound

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions cites this paper.

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions CogVLM2: Visual Language Models for Image and Video Understanding

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T17:30:47.844735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:30:47.844735Z digest=sha256:96f82dc320faeda7ed09393c3643f56a6677c369c6087ed69832040ce073fe1d

Observation 2d921696-ed34-4e29-871d-1a83a63f7f06 · inbound

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding cites this paper.

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding CogVLM2: Visual Language Models for Image and Video Understanding

Reference 153

Resolution
verified exact
arxiv_id, observed 2026-05-16T20:10:28.051835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-11T01:19:59.603343Z digest=sha256:98c16f95c62da33c5c2c8e24222deea658120a93c3869ad8ef144cfa9c706d67

Observation 24c36bcd-dbb8-49ab-a937-3ac590d86abb · inbound

Parameter-Efficient Fine-Tuning for Foundation Models cites this paper.

Parameter-Efficient Fine-Tuning for Foundation Models CogVLM2: Visual Language Models for Image and Video Understanding

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T15:38:02.989993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:38:02.989993Z digest=sha256:c33f65bc6b41b06f6c7d53f23c28a0d7c822c6539c6932b6dcbd3673682a2639

Observation a421bbdd-b9fd-433a-bc95-a95243aa9862 · inbound

Improving Video Generation with Human Feedback cites this paper.

Improving Video Generation with Human Feedback CogVLM2: Visual Language Models for Image and Video Understanding

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-16T20:10:28.051835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-13T15:30:02.578430Z digest=sha256:c78e219665f1b3a690855dd0e9a55ba355340b3fc4a3a5232cc10a7b3c979825

Observation 5f20bb0b-ca90-4602-96a2-3f33908ac49d · inbound

Temporal Preference Optimization for Long-Form Video Understanding cites this paper.

Temporal Preference Optimization for Long-Form Video Understanding CogVLM2: Visual Language Models for Image and Video Understanding

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T15:35:30.140139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:35:30.140139Z digest=sha256:af44c017906d149c70ab329ccbb1d918442b4812c04d17b0432e83eb61b16468

Observation ac01ab06-dbdb-4897-a3e3-12b933609d0d · inbound

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers cites this paper.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers CogVLM2: Visual Language Models for Image and Video Understanding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.716137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.716137Z digest=sha256:f9ba2ee4362ae49853a5c325ca6a5fb586d31136c83c9819e87e93be559574be

Observation 41854b49-82df-4c20-af01-61df17409811 · inbound

RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs cites this paper.

RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs CogVLM2: Visual Language Models for Image and Video Understanding

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T21:39:21.697649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T21:39:21.697649Z digest=sha256:d83b3cb2d549a2e90c0bdf200409f58afde7ca19cbd0d2a93146c66a312d4b7f

Observation d5152635-448a-4b1e-884c-1d77e25ade10 · inbound

MQuant: Unleashing the Inference Potential of Multimodal Large Language Models via Full Static Quantization cites this paper.

MQuant: Unleashing the Inference Potential of Multimodal Large Language Models via Full Static Quantization CogVLM2: Visual Language Models for Image and Video Understanding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-09T19:12:56.065018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:12:56.065018Z digest=sha256:df3989eda51e38a8d3db470368cceb0eb03414affa75b3bdb61e56bf4805d897

Observation ccf3b396-b07e-4a72-b983-5689db8bcda2 · inbound

Se\~norita-2M: A High-Quality Instruction-based Dataset for General Video Editing by Video Specialists cites this paper.

Se\~norita-2M: A High-Quality Instruction-based Dataset for General Video Editing by Video Specialists CogVLM2: Visual Language Models for Image and Video Understanding

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-08T14:34:02.009863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:34:02.009863Z digest=sha256:3f5c451ed8acb0791ee305a8b6003aa6b7db8a6da01b643361c2ec910a68630f

Observation b749407b-50fe-48b0-a978-ce8c3d8f9694 · inbound

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering cites this paper.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering CogVLM2: Visual Language Models for Image and Video Understanding

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.189601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.189601Z digest=sha256:f966bcdd46ac9e024ba7f78c784c1f7bd290e31bc42280d05e69cd8f642f0626

Observation d396599b-5e37-499d-a6da-3a60182807c4 · inbound

When 'YES' Meets 'BUT': Can Large Models Comprehend Contradictory Humor Through Comparative Reasoning? cites this paper.

When 'YES' Meets 'BUT': Can Large Models Comprehend Contradictory Humor Through Comparative Reasoning? CogVLM2: Visual Language Models for Image and Video Understanding

Reference 69

Resolution
verified exact
local_arxiv, observed 2026-05-22T22:42:13.561704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-22T22:38:35.969273Z digest=sha256:5a7a3bf9d643f177a92dbb9f8d534c03a34de422f46ee1e542da845e7698f46e

Observation 41cec851-faf3-49cf-b7bf-c6f8af1032d0 · inbound

AdaMMS: Model Merging for Heterogeneous Multimodal Large Language Models with Unsupervised Coefficient Optimization cites this paper.

AdaMMS: Model Merging for Heterogeneous Multimodal Large Language Models with Unsupervised Coefficient Optimization CogVLM2: Visual Language Models for Image and Video Understanding

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-22T22:47:13.016745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-22T22:47:09.500229Z digest=sha256:446202e5f0598a98f6b63330a1b5e36b8cc65602cbf3d63ee2eee1a16c113741

Observation e153ab84-b985-404f-b405-86e22a50c706 · inbound

EVA: Evolving Semantic Adversaries for Red-Teaming GUI Agents Against Environmental Injection Attacks cites this paper.

EVA: Evolving Semantic Adversaries for Red-Teaming GUI Agents Against Environmental Injection Attacks CogVLM2: Visual Language Models for Image and Video Understanding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T15:41:22.143040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:41:22.143040Z digest=sha256:5b4f6191808d4fb6da22ff346e89c1255c3a592a9972083ef23c95b331c1b2db

Observation 6425c5f5-2628-4408-8271-946a7a686471 · inbound

TimeCausality: Evaluating the Causal Ability in Time Dimension for Vision Language Models cites this paper.

TimeCausality: Evaluating the Causal Ability in Time Dimension for Vision Language Models CogVLM2: Visual Language Models for Image and Video Understanding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T15:21:04.051014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:21:04.051014Z digest=sha256:b1a5f7d3cae4f65f02956bcb116c7abeed494d91bf0088c96b752228f301693d

Observation 572b17a5-5d92-4dd8-82a7-d80b6564bfda · inbound

MDIT-Bench: Evaluating the Dual-Implicit Toxicity in Large Multimodal Models cites this paper.

MDIT-Bench: Evaluating the Dual-Implicit Toxicity in Large Multimodal Models CogVLM2: Visual Language Models for Image and Video Understanding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:33.405736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:06:33.405736Z digest=sha256:eeb33bef6f5fb74c734c851f25ef5fd2707ea84c99d187bb3264538923e2dae0

Observation 429f532a-e99e-4441-b5b6-32b649924ea2 · inbound

EgoWalk: A Multimodal Dataset for Robot Navigation in the Wild cites this paper.

EgoWalk: A Multimodal Dataset for Robot Navigation in the Wild CogVLM2: Visual Language Models for Image and Video Understanding

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-19T12:57:17.687038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-19T12:56:30.634284Z digest=sha256:a5b2a60e5d4d9795c9bb5c5ce75b1ea53ca52b5f6941a16f6ec58bdb7bbdb6d1

Observation 537f7211-5ccc-4a13-9826-b8b7293b8a29 · inbound

Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning? cites this paper.

Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning? CogVLM2: Visual Language Models for Image and Video Understanding

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-17T05:40:56.007393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T05:40:55.944288Z digest=sha256:f818577debeca257449febc14bba0ccc3e8694f8159382805515c5a111e503de

Observation 5426ae40-e7e6-4ba6-bb60-6b115f0d2f4a · inbound

Zero-Shot 3D Visual Grounding from Vision-Language Models cites this paper.

Zero-Shot 3D Visual Grounding from Vision-Language Models CogVLM2: Visual Language Models for Image and Video Understanding

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:16.099070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:16.099070Z digest=sha256:f4b808cca6e417ce8fa0ac9b175885451257824079944cb18f311fed31c0fd04

Observation a683d7db-d212-4e2f-92ee-460c5bb7e6bc · inbound

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction cites this paper.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction CogVLM2: Visual Language Models for Image and Video Understanding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:03.626178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:03.626178Z digest=sha256:99fd831168d6bb18312f4bd775202d0f6d1ec38e51729d5fbb5b7420810a434a

Observation 91c67106-19dc-4238-9ca3-0e6c841ed35d · inbound

VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation cites this paper.

VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation CogVLM2: Visual Language Models for Image and Video Understanding

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:02.656403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:48:02.656403Z digest=sha256:2dc145b375d7fb78a6a3b63b6175e06d0755db8ded6e2693d76e8ef76fd47072

Observation f617f3a5-e179-45cd-86e8-ac1931bcd33a · inbound

MiniMax-Remover: Taming Bad Noise Helps Video Object Removal cites this paper.

MiniMax-Remover: Taming Bad Noise Helps Video Object Removal CogVLM2: Visual Language Models for Image and Video Understanding

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:21:38.829107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:21:38.829107Z digest=sha256:6faad7f83f52545677459efa1948aa01c0daae9cc328ea0f1d64c95a631debe1

Observation 8d794475-5b78-48a6-9ad6-cbca8d7d57bb · inbound

GOBench: Benchmarking Geometric Optics Generation and Understanding of MLLMs cites this paper.

GOBench: Benchmarking Geometric Optics Generation and Understanding of MLLMs CogVLM2: Visual Language Models for Image and Video Understanding

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:23.590215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:23.590215Z digest=sha256:44b0b8410ca71309ebb8bc2453835173f6a54af8438ca62eebf113a4dc7d5286

Observation 33795074-5261-4f2d-a544-74a39440397d · inbound

Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement cites this paper.

Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement CogVLM2: Visual Language Models for Image and Video Understanding

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:39:52.921977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:39:52.921977Z digest=sha256:3a9f9c5e557e5d0645be370c903f33307cc645c95af18dd1e9be8f092782bbf6

Observation 3c67805f-ff99-426e-8c56-684306e97bc9 · inbound

IllumiCraft: Unified Geometry and Illumination Diffusion for Controllable Video Generation cites this paper.

IllumiCraft: Unified Geometry and Illumination Diffusion for Controllable Video Generation CogVLM2: Visual Language Models for Image and Video Understanding

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T11:14:44.036808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:14:44.036808Z digest=sha256:bd4fe638623d87fe817f2b419eeb732a903027512f4fad53b24cdd5b3d161470

Observation 54be0b75-8e85-4d36-95de-86a8a0f15c25 · inbound

MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos cites this paper.

MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos CogVLM2: Visual Language Models for Image and Video Understanding

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:35.680171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:35.680171Z digest=sha256:98cb7555e0d66e32c795ee33b7c21291bbc18ee45cd6f327c203b2cd439b69e7

Observation 4fbc802f-525f-457b-a1cc-6b238be05542 · inbound

LayerFlow: A Unified Model for Layer-aware Video Generation cites this paper.

LayerFlow: A Unified Model for Layer-aware Video Generation CogVLM2: Visual Language Models for Image and Video Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:50.116063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:50.116063Z digest=sha256:f15e3fe7164f7bc8ff91453a913cafc6d43c1bd96869addb3a08640e5e6b2b4d

Observation 46c9b6b1-1bdf-47dc-9e13-ea4ac9dc1f34 · inbound

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs cites this paper.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs CogVLM2: Visual Language Models for Image and Video Understanding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.505880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.505880Z digest=sha256:8a44c0e58b30d8d2c3d1008d6e94721dc09de6aaf2a0e34aefe98076b6784d92

Observation f76c8697-c3b6-44c8-8c19-bafeb7f3a54f · inbound

Cross-Frame Representation Alignment for Fine-Tuning Video Diffusion Models cites this paper.

Cross-Frame Representation Alignment for Fine-Tuning Video Diffusion Models CogVLM2: Visual Language Models for Image and Video Understanding

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:55.911607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:58:55.911607Z digest=sha256:96bf08814fe550b367a42921477d6dda5c58382d354e54c51ef65600a2f94444

Observation 75504491-fbe0-40f5-9b1f-cbf7465bca9c · inbound

VideoMat: Extracting PBR Materials from Video Diffusion Models cites this paper.

VideoMat: Extracting PBR Materials from Video Diffusion Models CogVLM2: Visual Language Models for Image and Video Understanding

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T04:48:24.087480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:48:24.087480Z digest=sha256:1353fa16a8ca713e4ccfc739403c32a0c9f7e194c94d4c8493995ccda3cdf45a

Observation 19ba0531-62db-4840-8e20-4266b6af1531 · inbound

PARC: A Quantitative Framework Uncovering the Symmetries within Vision Language Models cites this paper.

PARC: A Quantitative Framework Uncovering the Symmetries within Vision Language Models CogVLM2: Visual Language Models for Image and Video Understanding

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:12:03.375223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:12:03.375223Z digest=sha256:4e5358f9e4e70e41321a0d736220cfc996af40e3c301d40d87a84be34c68bca0

Observation fdf1a535-d8e8-4c25-858e-fde0cda07149 · inbound

GenRecal: Generation after Recalibration from Large to Small Vision-Language Models cites this paper.

GenRecal: Generation after Recalibration from Large to Small Vision-Language Models CogVLM2: Visual Language Models for Image and Video Understanding

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T23:57:22.901966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:57:22.901966Z digest=sha256:b04b2bb5cc506233b00fdabd8e1b862ba1f65ad41cf4add0dfedc27761764693

Observation ae57d9f2-bfa5-4289-ad01-016bac8bedbb · inbound

AnyAni: An Interactive System with Generative AI for Animation Effect Creation and Code Understanding in Web Development cites this paper.

AnyAni: An Interactive System with Generative AI for Animation Effect Creation and Code Understanding in Web Development CogVLM2: Visual Language Models for Image and Video Understanding

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T22:18:39.318280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:18:39.318280Z digest=sha256:c57005207856f25804b0853d252f5a885274e7f58fc6e2e9b209522c34c0a8e3

Observation 8cb5907f-1712-434a-b50f-90b64a967961 · inbound

CAI: Caption-Sensitive Attention Intervention for Mitigating Object Hallucination in Large Vision-Language Models cites this paper.

CAI: Caption-Sensitive Attention Intervention for Mitigating Object Hallucination in Large Vision-Language Models CogVLM2: Visual Language Models for Image and Video Understanding

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T21:44:56.136017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:44:56.136017Z digest=sha256:433b670b59bc2e08cd1781b33989dd7a4c33d051fefc118ae1df8008b74af3e1

Observation c18bd244-ec45-4b50-8fdd-e5e641b6b1d3 · inbound

GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning cites this paper.

GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning CogVLM2: Visual Language Models for Image and Video Understanding

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-16T20:10:28.051835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-11T04:48:26.355351Z digest=sha256:b0cbdf55969d705d35954393ea46a6a8afe559ed9859ae0ef97c2a30a98fec76

Observation cf4e2d1c-22da-461c-97af-ba5d821a63f1 · inbound

LongAnimation: Long Animation Generation with Dynamic Global-Local Memory cites this paper.

LongAnimation: Long Animation Generation with Dynamic Global-Local Memory CogVLM2: Visual Language Models for Image and Video Understanding

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T20:44:00.477581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:44:00.477581Z digest=sha256:6ccea59730223a90167417c3596e208cbe3bb7f811ba029dac42656a9d5aafb8

Observation ee2df0b5-ed51-4e1a-8ae4-659f11bf2e08 · inbound

Ascending the Infinite Ladder: Benchmarking Spatial Deformation Reasoning in Vision-Language Models cites this paper.

Ascending the Infinite Ladder: Benchmarking Spatial Deformation Reasoning in Vision-Language Models CogVLM2: Visual Language Models for Image and Video Understanding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T21:22:27.120045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:22:27.120045Z digest=sha256:ece3667a3a1672018bd94cb3e4b593db4cc81ddaead89d11693232959319e277

Observation fe18d58a-759a-4aeb-a4ab-e676cca4ece0 · inbound

From Answers to Rationales: Self-Aligning Multimodal Reasoning with Answer-Oriented Chain-of-Thought cites this paper.

From Answers to Rationales: Self-Aligning Multimodal Reasoning with Answer-Oriented Chain-of-Thought CogVLM2: Visual Language Models for Image and Video Understanding

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:34.077710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:34.077710Z digest=sha256:f4a55bf9d034d9f70c4814376edf96ac6c0037eabd69a1eeda52c809fbb487f8

Observation af92a477-8330-412a-944c-438aea6291d5 · inbound

ExpStar: Towards Automatic Commentary Generation for Multi-discipline Scientific Experiments cites this paper.

ExpStar: Towards Automatic Commentary Generation for Multi-discipline Scientific Experiments CogVLM2: Visual Language Models for Image and Video Understanding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T17:55:45.294752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:55:45.294752Z digest=sha256:469f1446cd4110763b97e8ed6d2792b442d1ce42623642377d66191ff4bb89fb

Observation e1c538fe-0dbf-4454-9c19-4e70034d0425 · inbound

Foundation Model Driven Robotics: A Comprehensive Review cites this paper.

Foundation Model Driven Robotics: A Comprehensive Review CogVLM2: Visual Language Models for Image and Video Understanding

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T17:43:53.061653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:43:53.061653Z digest=sha256:485f740972142d3ab8ce4dd851a68e6f10fb9c6c4fde87804a3448ec1fcae7e6

Observation 4b0f739b-1099-42e8-ad53-78744aec49e6 · inbound

FantasyPortrait: Enhancing Multi-Character Portrait Animation with Expression-Augmented Diffusion Transformers cites this paper.

FantasyPortrait: Enhancing Multi-Character Portrait Animation with Expression-Augmented Diffusion Transformers CogVLM2: Visual Language Models for Image and Video Understanding

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T16:39:34.667066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:39:34.667066Z digest=sha256:58047783f0ae78c10e42877385705f23f9ec32543b270f1c451d8d69cb26fa96

Observation f560e7b0-eef2-4c07-84fb-f2935996c012 · inbound

Hate in Plain Sight: On the Risks of Moderating AI-Generated Hateful Illusions cites this paper.

Hate in Plain Sight: On the Risks of Moderating AI-Generated Hateful Illusions CogVLM2: Visual Language Models for Image and Video Understanding

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T11:34:26.624192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:34:26.624192Z digest=sha256:eb1c15c0738cebadf9b9dc88adac2955fd477118363fe73f84bd30a43e4c8729

Observation 88a1cdb1-3f66-4c60-9aaf-442fd10cca4d · inbound

CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text cites this paper.

CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text CogVLM2: Visual Language Models for Image and Video Understanding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T10:10:57.591421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:10:57.591421Z digest=sha256:457aa978149c627f5c08d6cc17a1f99fae951d734e5d8ea51c473354d1cfdcd6

Observation b2590721-641b-4a1e-a652-db841f948b71 · inbound

Detailed radial scale height profile of dust grains as probed by dust self-scattering in HL Tau cites this paper.

Detailed radial scale height profile of dust grains as probed by dust self-scattering in HL Tau CogVLM2: Visual Language Models for Image and Video Understanding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T05:46:59.364916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:46:59.364916Z digest=sha256:b109fce56f06f91ec2fc09ec83a15f30701965d1d2a55f9375e118cc0dfe7703

Observation eb9829c8-d950-44cf-a218-957407b31b1f · inbound

SketchAgent: Generating Structured Diagrams from Hand-Drawn Sketches cites this paper.

SketchAgent: Generating Structured Diagrams from Hand-Drawn Sketches CogVLM2: Visual Language Models for Image and Video Understanding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T05:48:15.980138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:48:15.980138Z digest=sha256:506f0b26e48bc547be8c1a58801dac8dcc73fd8b2a8856c9ffb061575e39e56b

Observation cf623911-4578-4e0b-aae1-804bbb548cce · inbound

CTA-Flux: Integrating Chinese Cultural Semantics into High-Quality English Text-to-Image Communities cites this paper.

CTA-Flux: Integrating Chinese Cultural Semantics into High-Quality English Text-to-Image Communities CogVLM2: Visual Language Models for Image and Video Understanding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T18:40:33.705789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T18:40:33.705789Z digest=sha256:fdcfbc07f1eb532621098423121d2184cab319bf276915bc1b35b31e89562c86

Observation e7d8a105-9d58-482a-a151-9c16b8c07adb · inbound

Improving Large Vision and Language Models by Learning from a Panel of Peers cites this paper.

Improving Large Vision and Language Models by Learning from a Panel of Peers CogVLM2: Visual Language Models for Image and Video Understanding

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.345868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.345868Z digest=sha256:ab15f003714acdb8b1200c9ef4cff1b3d67ea6783bc96c2bda5c555f7d36a23e

Observation 6b30dc52-6e55-4609-94cd-8a3165b86594 · inbound

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data cites this paper.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data CogVLM2: Visual Language Models for Image and Video Understanding

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T10:58:16.100830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:58:16.100830Z digest=sha256:12656c348a2e94ed1758e4b252ee9924021854f15feb3f9373ff55d7f66b38ee

Observation f19733e0-5fd9-486a-93cb-e830037cad92 · inbound

EchoVLM: Dynamic Mixture-of-Experts Vision-Language Model for Universal Ultrasound Intelligence cites this paper.

EchoVLM: Dynamic Mixture-of-Experts Vision-Language Model for Universal Ultrasound Intelligence CogVLM2: Visual Language Models for Image and Video Understanding

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-18T16:06:35.099190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-18T16:04:11.834915Z digest=sha256:3bb5765b1476cfc52ecf2ccd3c3c55c15eff623094ccccacdc797339475f017d

Observation a0652a44-4016-42ae-a5d1-601fca47aee9 · inbound

EchoVLM: Dynamic Mixture-of-Experts Vision-Language Model for Universal Ultrasound Intelligence cites this paper.

EchoVLM: Dynamic Mixture-of-Experts Vision-Language Model for Universal Ultrasound Intelligence CogVLM2: Visual Language Models for Image and Video Understanding

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T16:15:50.676012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:15:50.676012Z digest=sha256:b9fc7ace398f61092c428558855e096a4356b188e8281331e6f063a4acec9838

Observation 3215cbe9-0022-45f5-ba4b-b5c0262f5d43 · inbound

FG-CLIP 2: A Bilingual Fine-grained Vision-Language Alignment Model cites this paper.

FG-CLIP 2: A Bilingual Fine-grained Vision-Language Alignment Model CogVLM2: Visual Language Models for Image and Video Understanding

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-04T10:18:42.991905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:18:42.991905Z digest=sha256:c45663eaa503d4b5a925165142de50df0f4ab158d3c5882ac2e34abfc08bfbb3

Observation 65acc6f3-cd13-4534-b900-bac1388877c0 · inbound

Geometry-Aware Single-Image 4D Synthesis via Dense Trajectory Generation cites this paper.

Geometry-Aware Single-Image 4D Synthesis via Dense Trajectory Generation CogVLM2: Visual Language Models for Image and Video Understanding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T18:30:03.280897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:30:03.280897Z digest=sha256:444777f552f284b52acbe0cbac94333e0bb3e7ddc4cc744ec3e25c919fe26f5b

Observation cc137128-489f-4e02-a161-27f53335a197 · inbound

High-Entropy Tokens as Multimodal Failure Points in Vision-Language Models cites this paper.

High-Entropy Tokens as Multimodal Failure Points in Vision-Language Models CogVLM2: Visual Language Models for Image and Video Understanding

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-16T20:10:28.051835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T19:29:43.382392Z digest=sha256:37d91cea62912136646c26d68d2a185cb77c8beb29f812b70bb7423f475fcc89

Observation 671de08b-2c86-4874-98ed-33ccf864a62e · inbound

High-Entropy Tokens as Multimodal Failure Points in Vision-Language Models cites this paper.

High-Entropy Tokens as Multimodal Failure Points in Vision-Language Models CogVLM2: Visual Language Models for Image and Video Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T14:05:23.078713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:05:23.078713Z digest=sha256:b61e3256be4f9fce07d351b0d3631e4386cac5e6161d501448bcf892004bc38e

Observation 08af0cf1-b685-4976-8e52-8b998fb76eea · inbound

HaineiFRDM: Structure-Preserving Diffusion for Film Restoration under Fast Motion and Diverse Defects cites this paper.

HaineiFRDM: Structure-Preserving Diffusion for Film Restoration under Fast Motion and Diverse Defects CogVLM2: Visual Language Models for Image and Video Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T13:15:30.305563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:15:30.305563Z digest=sha256:8b53306ffa1fdd4cb563875984c75e070da5d39abc873784d009ce92b3db1045

Observation e4af9e13-2c93-44b5-aa23-edd890386e92 · inbound

VideoGPA: Distilling Geometry Priors for 3D-Consistent Video Generation cites this paper.

VideoGPA: Distilling Geometry Priors for 3D-Consistent Video Generation CogVLM2: Visual Language Models for Image and Video Understanding

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-16T20:10:28.051835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T09:16:57.528410Z digest=sha256:fbf0fc0642d25b2704c18dd5a10a9886adc3ab2ce90cb701584ba2f224ddf385

Observation 8ff2be2d-ac4b-4450-ba60-3fa3a91e10ae · inbound

VideoGPA: Distilling Geometry Priors for 3D-Consistent Video Generation cites this paper.

VideoGPA: Distilling Geometry Priors for 3D-Consistent Video Generation CogVLM2: Visual Language Models for Image and Video Understanding

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-03T06:13:19.065521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:13:19.065521Z digest=sha256:4c5138893a283c727be4b13f034a750940279c84f475b6a81fb8e9d1ff60dd10

Observation bdf92923-e799-48b2-b36a-69a32bcd5fdb · inbound

EgoIntent: A Pre-Outcome Micro-Step Benchmark for Understanding What, Why, and Next cites this paper.

EgoIntent: A Pre-Outcome Micro-Step Benchmark for Understanding What, Why, and Next CogVLM2: Visual Language Models for Image and Video Understanding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T05:50:27.185756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:50:27.185756Z digest=sha256:f3cc058916120823efe893186a6dae73cd24b3b36c9e0619fd0e06b7b072a711

Observation d5153d1e-8319-4c13-a2fa-345e69018981 · inbound

MIRAGE: Benchmarking and Aligning Multi-Instance Image Editing cites this paper.

MIRAGE: Benchmarking and Aligning Multi-Instance Image Editing CogVLM2: Visual Language Models for Image and Video Understanding

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T20:10:28.051835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T19:38:45.361389Z digest=sha256:5d1379827133cd8844d7a8c1da0c70f88ce8b8afa21de036304aa09a7ad32d3b

Observation 69ea4b6d-93e6-4b83-9c97-c02f0adac2ca · inbound

VersaVogue: Visual Expert Orchestration and Preference Alignment for Unified Fashion Synthesis cites this paper.

VersaVogue: Visual Expert Orchestration and Preference Alignment for Unified Fashion Synthesis CogVLM2: Visual Language Models for Image and Video Understanding

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-16T20:10:28.051835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T19:00:29.650500Z digest=sha256:67388d078cd4df67505b2955d455802f2fe25beffdfd2cc76604afbc0dcce9eb

Observation 92d96f62-5ef7-4f2a-96fa-6b2728cd23c4 · inbound

VSAS-Bench: Real-Time Evaluation of Visual Streaming Assistant Models cites this paper.

VSAS-Bench: Real-Time Evaluation of Visual Streaming Assistant Models CogVLM2: Visual Language Models for Image and Video Understanding

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-16T20:10:28.051835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T17:39:30.731688Z digest=sha256:320b4bcc070c03bb3b9109c96dfdc96453fa9415e9efa82bf147b6d253b1cb89

Observation 143042a3-3692-488a-9295-822a4320a670 · inbound

Towards Unconstrained Human-Object Interaction cites this paper.

Towards Unconstrained Human-Object Interaction CogVLM2: Visual Language Models for Image and Video Understanding

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-16T20:10:28.051835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T14:17:06.952676Z digest=sha256:b6cbeb7013927ec23abc8fed64053929d9911247d84bce40d16e3a7146b0f603

Observation 04e81a59-7c10-4f66-9e06-c8539e73423a · inbound

Culture-Aware Humorous Captioning: Multimodal Humor Generation across Cultural Contexts cites this paper.

Culture-Aware Humorous Captioning: Multimodal Humor Generation across Cultural Contexts CogVLM2: Visual Language Models for Image and Video Understanding

Reference 77

Resolution
verified exact
arxiv_id, observed 2026-05-16T20:10:28.051835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T05:33:20.557526Z digest=sha256:4c6af9d82929c7773fa4fdedbe0b51735557d9acfb9a4f7b33b6577c53cdf868

Observation 3dfaeb8b-e91b-4920-a9e8-266660ee8bef · inbound

ReCAPA: Hierarchical Predictive Correction to Mitigate Cascading Failures cites this paper.

ReCAPA: Hierarchical Predictive Correction to Mitigate Cascading Failures CogVLM2: Visual Language Models for Image and Video Understanding

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T20:10:28.051835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-09T22:21:54.577215Z digest=sha256:9d65d0fafa7d1fc8c9f9631f45013f9cb243656475eab03588562c0a0bbca06b

Observation 097dc07a-3f65-4eec-91b7-e0e9f11de37a · inbound

ReCAPA: Hierarchical Predictive Correction to Mitigate Cascading Failures cites this paper.

ReCAPA: Hierarchical Predictive Correction to Mitigate Cascading Failures CogVLM2: Visual Language Models for Image and Video Understanding

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T20:10:28.051835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-12T03:08:25.891080Z digest=sha256:83b025525c78c45251d11cf2ccbd34d7118159e9affcc7887abe6753d2404346

Observation 74f388d3-9f17-4a53-9523-351980ccc556 · inbound

KD-CVG: A Knowledge-Driven Approach for Creative Video Generation cites this paper.

KD-CVG: A Knowledge-Driven Approach for Creative Video Generation CogVLM2: Visual Language Models for Image and Video Understanding

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-16T20:10:28.051835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-09T22:51:13.632121Z digest=sha256:9dfa7e3ece464cde2d27f85f15d687b6ed028feaa34e86116a5028dd695c5da2

Observation 790d6ea8-0e57-4dcc-8f9d-3598530e9826 · inbound

CAST: Mitigating Object Hallucination in Large Vision-Language Models via Caption-Guided Visual Attention Steering cites this paper.

CAST: Mitigating Object Hallucination in Large Vision-Language Models via Caption-Guided Visual Attention Steering CogVLM2: Visual Language Models for Image and Video Understanding

Reference 53

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T20:10:28.051835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-08T18:05:38.705480Z digest=sha256:95b7dc74412887b448bdd4eacf665cb5e6766f64e195135fb53160547d2f33a3

Observation 5dd0961d-d8c1-434c-860c-4f5709476b7a · inbound

Illusion-Aware Visual Preprocessing and Anti-Illusion Prompting for Classic Illusion Understanding in Vision-Language Models cites this paper.

Illusion-Aware Visual Preprocessing and Anti-Illusion Prompting for Classic Illusion Understanding in Vision-Language Models CogVLM2: Visual Language Models for Image and Video Understanding

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-16T20:10:28.051835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-12T01:39:17.468076Z digest=sha256:188f66c9c1e08e7ff4608fa29aa84d569cec0b060132d8608c08ebd459dc820f

Observation 31174350-03f7-4a3a-9e1c-b691b3e181df · inbound

Omni-DuplexEval: Evaluating Real-time Duplex Omni-modal Interaction cites this paper.

Omni-DuplexEval: Evaluating Real-time Duplex Omni-modal Interaction CogVLM2: Visual Language Models for Image and Video Understanding

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-20T13:38:19.184600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-20T13:36:44.071188Z digest=sha256:8391bdf570bc7da7cbf361edbbd654eaf6e8fd275208185e265b21d29e105474

Observation 0e39044d-1a37-4bc6-8de4-60c62b310850 · inbound

Omni-DuplexEval: Evaluating Real-time Duplex Omni-modal Interaction cites this paper.

Omni-DuplexEval: Evaluating Real-time Duplex Omni-modal Interaction CogVLM2: Visual Language Models for Image and Video Understanding

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-07-04T01:19:20.300620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-04T01:11:42.073993Z digest=sha256:6ad7d01071cff9c7036e035e3a9b5accb97bf6f7bb2e831d155da4e6a26e1b34

Observation e07336f4-347c-40a0-9901-4679646b9a37 · inbound

From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models cites this paper.

From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models CogVLM2: Visual Language Models for Image and Video Understanding

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T05:13:21.481162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:e7de15b48eb17e69a32714699bdc7cbd81b8984a1d50e447e6467ce73732c0fb

Observation b15ad2f6-0690-4fd5-a6e4-d3236130f31c · inbound

MONET: A Massive, Open, Non-redundant and Enriched Text-to-image dataset cites this paper.

MONET: A Massive, Open, Non-redundant and Enriched Text-to-image dataset CogVLM2: Visual Language Models for Image and Video Understanding

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-21T05:23:58.461915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-21T05:21:18.369534Z digest=sha256:086b68ec87920fd92b0c4387bb8f50c286e1cec225dd2391749337bc0886e848

Observation a43ce7f1-9338-48fa-9550-81cc9503f3ef · inbound

TrackRef3D: Multi-View Consistent Track-then-Label for Open-World Referring Segmentation in 3D Gaussian Splatting cites this paper.

TrackRef3D: Multi-View Consistent Track-then-Label for Open-World Referring Segmentation in 3D Gaussian Splatting CogVLM2: Visual Language Models for Image and Video Understanding

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-06-29T18:43:50.597565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-29T18:39:28.055568Z digest=sha256:e5c7391c9aa41ba05e8ed69af870e2d148874275464d39225539d0ba9d4e7697

Observation 9d37e3c6-e5de-43f0-82cb-9982e653f131 · inbound

MotionEnhancer: Leveraging Video Diffusion for Motion-Enhanced Vision-Language Models cites this paper.

MotionEnhancer: Leveraging Video Diffusion for Motion-Enhanced Vision-Language Models CogVLM2: Visual Language Models for Image and Video Understanding

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-02T16:17:09.558511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-27T22:49:18.420491Z digest=sha256:44f35bc86d23a461a6f7b6aa1fa668114fbfa9fa625d5e35e1a2af308ca463ff

Observation eb8cf754-4144-46a4-82f6-cfbed09c3279 · inbound

PixelEyes: Decoupling Perception and Reasoning for Pinpoint Visual Evidence Seeking cites this paper.

PixelEyes: Decoupling Perception and Reasoning for Pinpoint Visual Evidence Seeking CogVLM2: Visual Language Models for Image and Video Understanding

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-02T19:57:19.117966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-02T19:52:39.269265Z digest=sha256:b9bfbbd62421b1d635d57d2232c35ef39e6ca314cafdc496ed0e4709547a496b

Observation 6de49193-a046-4c4f-b4b4-63212490ad9d · inbound

EgoPolice: A Benchmark for Egocentric Video Understanding in High-Stakes Police Body-Worn Camera Footage cites this paper.

EgoPolice: A Benchmark for Egocentric Video Understanding in High-Stakes Police Body-Worn Camera Footage CogVLM2: Visual Language Models for Image and Video Understanding

Reference 38

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T05:34:32.403838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-08T05:25:07.553526Z digest=sha256:c1b5cd5715a44ab4aa6eb38f0c55fabd42f51688bf12b8af298a2976b6e51a06

Observation 0e3f8377-2e3d-4766-97cc-b1d086c7d0c0 · inbound

WeaveEarth: Structured Evidence Construction and Reasoning for Training-Free UHR Remote Sensing Understanding cites this paper.

WeaveEarth: Structured Evidence Construction and Reasoning for Training-Free UHR Remote Sensing Understanding CogVLM2: Visual Language Models for Image and Video Understanding

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-14T14:09:30.395518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:09:30.395518Z digest=sha256:58bdf7df90eaf2ab4109a5a58b80cabefb58060e143d9080eded18a65b54cbf6

Observation 0e519978-a73f-4260-aef9-789646c7c880 · inbound

Actor as Its Own Critic: Unifying Region Understanding and Localization via CycleGRPO cites this paper.

Actor as Its Own Critic: Unifying Region Understanding and Localization via CycleGRPO CogVLM2: Visual Language Models for Image and Video Understanding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-14T04:38:05.237334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:38:05.237334Z digest=sha256:5f4c97e9c1838e4e25174c2b00403774419051ef0b6adf93e263ebdddca29702