Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-16T20:10:27.633010Z
Paper Citation Record · LEDGER
As of 12 August 2026, this Paper Citation Record lists 94 of 94 outbound references and 100 inbound Pith citation observations for arXiv:2408.16500.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-16T20:10:27.633010Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-12T14:51:13.200029Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-08T05:34:32.402425Z
94 of 94 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 9a54e822-4ab9-45b7-a39e-79850111f800 · outbound
CogVLM2: Visual Language Models for Image and Video Understanding Acharya, K
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f519b59e-a69c-4472-b0bb-d9b17202b92b · outbound
CogVLM2: Visual Language Models for Image and Video Understanding GPT-4 Technical Report
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 12e69b45-834a-484c-aabc-ec199a801209 · outbound
CogVLM2: Visual Language Models for Image and Video Understanding An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation fac6f531-5b74-4cec-bd42-bba3ee2772cb · outbound
CogVLM2: Visual Language Models for Image and Video Understanding Antol, A
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f42fe827-7dae-465d-b5df-dddf0df14e2f · outbound
CogVLM2: Visual Language Models for Image and Video Understanding Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation db0ce25f-ab37-496b-971a-64edebd88ffb · outbound
CogVLM2: Visual Language Models for Image and Video Understanding Unresolved cited work
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 399385a3-d04a-451c-a72f-118c1319ef66 · outbound
CogVLM2: Visual Language Models for Image and Video Understanding Nougat: Neural Optical Understanding for Academic Documents
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 46bc0e43-3c59-4e0a-b2be-c0b73962a863 · outbound
CogVLM2: Visual Language Models for Image and Video Understanding Byeon, B
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b7f05ec2-497f-48cd-866b-68e5f0bb5ddc · outbound
CogVLM2: Visual Language Models for Image and Video Understanding Unresolved cited work
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation be43286c-aaa6-4f05-ac68-2b86a40dd27a · outbound
CogVLM2: Visual Language Models for Image and Video Understanding Are We on the Right Way for Evaluating Large Vision-Language Models?
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation adafa0e3-6f81-4449-ad51-6dc21c67c500 · outbound
CogVLM2: Visual Language Models for Image and Video Understanding ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 62580b45-2181-4231-958f-393f3d04b395 · outbound
CogVLM2: Visual Language Models for Image and Video Understanding PaLI: A Jointly-Scaled Multilingual Language-Image Model
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 51873788-83fa-429c-b8c5-18d16cb5aa6c · outbound
CogVLM2: Visual Language Models for Image and Video Understanding InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8a0fe6c1-7d03-4073-8176-4821a424e7e4 · outbound
CogVLM2: Visual Language Models for Image and Video Understanding Unresolved cited work
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4949e172-7d11-4456-b104-8a87ba7f15eb · outbound
CogVLM2: Visual Language Models for Image and Video Understanding G-LLaVA: Solving Geometric Problem with Multi-Modal Large Language Model
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d2271787-305b-4483-ac57-d267a4f73ea6 · outbound
CogVLM2: Visual Language Models for Image and Video Understanding ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 37a4fbc5-5967-4bd6-9e07-fbfb8378c8bf · outbound
CogVLM2: Visual Language Models for Image and Video Understanding something something
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 1b6a4e1b-ab6b-4c72-866b-0d04e57b18b4 · outbound
CogVLM2: Visual Language Models for Image and Video Understanding Grauman, A
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 0a3bfad9-14b4-4706-9b86-3de7b1caf828 · outbound
CogVLM2: Visual Language Models for Image and Video Understanding Unresolved cited work
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 11bb57e1-7c3f-4c79-845c-c6eff8d4c976 · outbound
CogVLM2: Visual Language Models for Image and Video Understanding Unresolved cited work
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 48d09d39-17d0-4168-9d62-508ad1e94839 · outbound
CogVLM2: Visual Language Models for Image and Video Understanding Kafle, S
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 0bd75f35-dd08-4032-b41e-c260c900e94f · outbound
CogVLM2: Visual Language Models for Image and Video Understanding Kafle and C
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation afefd87b-f583-4aab-9f92-0a9981b0e610 · outbound
CogVLM2: Visual Language Models for Image and Video Understanding Unresolved cited work
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 5d35e18a-09b7-4db3-bfd2-bc7f5ee8c95b · outbound
CogVLM2: Visual Language Models for Image and Video Understanding GeomVerse: A Systematic Evaluation of Large Models for Geometric Reasoning
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e1baaf04-8de4-4025-a39c-71e9b8e2ab62 · outbound
CogVLM2: Visual Language Models for Image and Video Understanding Kembhavi, M
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e26de3c2-270c-4c6a-86b2-75d766afb424 · outbound
CogVLM2: Visual Language Models for Image and Video Understanding Kembhavi, M
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation da9fac4f-310f-4c21-8981-f7ab8b46e639 · outbound
CogVLM2: Visual Language Models for Image and Video Understanding Kembhavi, M
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 89f11a57-515c-414d-9240-cf1375e20207 · outbound
CogVLM2: Visual Language Models for Image and Video Understanding Unresolved cited work
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation a7a21ed4-e12f-4fc1-965f-0791d835de4e · outbound
CogVLM2: Visual Language Models for Image and Video Understanding Krishna, Y
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9885d6ee-3a90-4f91-a337-566fd07ff03e · outbound
CogVLM2: Visual Language Models for Image and Video Understanding Unresolved cited work
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4607e1dd-2765-48e0-a96f-340c3d8c9cf5 · outbound
CogVLM2: Visual Language Models for Image and Video Understanding PP-OCRv3: More Attempts for the Improvement of Ultra Lightweight OCR System
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e0d47ba8-940f-4113-95a6-4d73cdb8ecbc · outbound
CogVLM2: Visual Language Models for Image and Video Understanding Unresolved cited work
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 0652cd27-6b42-445f-8066-cb5879896a28 · outbound
CogVLM2: Visual Language Models for Image and Video Understanding VideoChat: Chat-Centric Video Understanding
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 08d38f54-f572-4214-ba13-000243b31a84 · outbound
CogVLM2: Visual Language Models for Image and Video Understanding Unresolved cited work
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 215434de-9888-454e-8479-a6ce6fd57807 · outbound
CogVLM2: Visual Language Models for Image and Video Understanding UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4752dbaa-481f-4104-8ffd-fc14d2ab159a · outbound
CogVLM2: Visual Language Models for Image and Video Understanding Unresolved cited work
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b6a8c910-c016-4550-a1a3-c16656d56357 · outbound
CogVLM2: Visual Language Models for Image and Video Understanding Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b56541ca-3f59-498b-8d37-80c9589f0766 · outbound
CogVLM2: Visual Language Models for Image and Video Understanding Unresolved cited work
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 933aeecf-ff3e-412e-a5d0-0bcdedb3c0d7 · outbound
CogVLM2: Visual Language Models for Image and Video Understanding Unresolved cited work
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9628919e-6ccf-4368-99b2-e7c76070419a · outbound
CogVLM2: Visual Language Models for Image and Video Understanding Unresolved cited work
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 006b469e-8057-48f0-993d-12705df4fd29 · outbound
CogVLM2: Visual Language Models for Image and Video Understanding Unresolved cited work
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation fba93d6d-aade-463d-ac2e-5d114150ca90 · outbound
CogVLM2: Visual Language Models for Image and Video Understanding Unresolved cited work
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ed258d1e-0ec7-43b8-9f8e-ec97490736e6 · outbound
CogVLM2: Visual Language Models for Image and Video Understanding ST-LLM: Large Language Models Are Effective Temporal Learners
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d22a74fe-19eb-4142-9be5-aade3379bff5 · outbound
CogVLM2: Visual Language Models for Image and Video Understanding MMBench: Is Your Multi-modal Model an All-around Player?
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation dda4b2f0-c538-47a5-bac8-b6bf4feddb93 · outbound
CogVLM2: Visual Language Models for Image and Video Understanding Unresolved cited work
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 3fadd1fa-3df3-4adc-a7d5-5021ba705c83 · outbound
CogVLM2: Visual Language Models for Image and Video Understanding Unresolved cited work
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8378d1cc-3b1a-4cea-825a-131a12ef1827 · outbound
CogVLM2: Visual Language Models for Image and Video Understanding Unresolved cited work
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8c5f4bca-301a-409b-b94e-fbcf5869b7da · outbound
CogVLM2: Visual Language Models for Image and Video Understanding Unresolved cited work
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 2e7da8ad-e0e9-4990-b680-b7a544d73a61 · outbound
CogVLM2: Visual Language Models for Image and Video Understanding Unresolved cited work
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation bc8c901a-6038-4a53-bf88-5c8e21cfa6ee · outbound
CogVLM2: Visual Language Models for Image and Video Understanding Unresolved cited work
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 67d4e8bc-6a3f-4ec2-976b-43b6e236dfc6 · outbound
CogVLM2: Visual Language Models for Image and Video Understanding Unresolved cited work
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 83c4514c-9187-4a43-8e84-18741ebd59e5 · outbound
CogVLM2: Visual Language Models for Image and Video Understanding Unresolved cited work
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e8f9307d-1efa-4220-8899-e3a489d2c08b · outbound
CogVLM2: Visual Language Models for Image and Video Understanding Unresolved cited work
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 49ecbadf-39ca-4f01-ae6c-126297641938 · outbound
CogVLM2: Visual Language Models for Image and Video Understanding Marino, M
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ab2f312b-a3dd-4ac5-8edb-6a4dcf3ce4a3 · outbound
CogVLM2: Visual Language Models for Image and Video Understanding Marti and H
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d620c82a-f797-4f38-99b7-ea1df47a204a · outbound
CogVLM2: Visual Language Models for Image and Video Understanding Masry, D
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c0ac1a4a-60ca-4b97-b35b-6498f7642f97 · outbound
CogVLM2: Visual Language Models for Image and Video Understanding ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 2be31272-5224-4446-b428-9ab5ec0606c1 · outbound
CogVLM2: Visual Language Models for Image and Video Understanding Masry, D
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c93e4f5f-46a4-48ae-a803-e880fa3cccdf · outbound
CogVLM2: Visual Language Models for Image and Video Understanding Mathew, V
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 5eb98273-7f29-4c13-902c-a3e918dcdebc · outbound
CogVLM2: Visual Language Models for Image and Video Understanding Mathew, D
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 60611023-95f5-44f8-a7f2-64c1d6a7e560 · outbound
CogVLM2: Visual Language Models for Image and Video Understanding Mathew, D
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4bc8b320-4139-4516-970f-837028af3720 · outbound
CogVLM2: Visual Language Models for Image and Video Understanding Mathew, D
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 105879cc-09d2-4e7c-8ba8-8b74d15ede99 · outbound
CogVLM2: Visual Language Models for Image and Video Understanding Mishra, S
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 1da3e238-20a2-4283-8d52-77202e7b22ee · outbound
CogVLM2: Visual Language Models for Image and Video Understanding Unresolved cited work
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 67e47033-07ba-4894-852d-4d263c54316e · outbound
CogVLM2: Visual Language Models for Image and Video Understanding Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ca324a57-d643-4590-a8bb-7f477c6345fe · outbound
CogVLM2: Visual Language Models for Image and Video Understanding TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4ee1ecd9-8719-4618-9e4e-fa49f8b6f344 · outbound
CogVLM2: Visual Language Models for Image and Video Understanding Schuhmann, R
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 431a3330-2eca-4b36-a522-366b88b7377b · outbound
CogVLM2: Visual Language Models for Image and Video Understanding LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 010ea674-078b-429a-8d10-5aa2a183eb75 · outbound
CogVLM2: Visual Language Models for Image and Video Understanding Schwenk, A
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6b4bff8d-8b16-4a8c-a3a9-be4adb8b7413 · outbound
CogVLM2: Visual Language Models for Image and Video Understanding Unresolved cited work
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation de1d903c-78f6-49bc-a3f4-57328560d13a · outbound
CogVLM2: Visual Language Models for Image and Video Understanding Singh, V
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d58a588d-378d-4b28-b029-c1e75ef45aa2 · outbound
CogVLM2: Visual Language Models for Image and Video Understanding Singh, V
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation a75f5f2b-8ca2-4d28-8161-68348ad4b46a · outbound
CogVLM2: Visual Language Models for Image and Video Understanding EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation aa30b319-8c08-4bd8-b180-873ab01e7054 · outbound
CogVLM2: Visual Language Models for Image and Video Understanding Unresolved cited work
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8acc3d6d-4f6e-4ccb-88be-df59bc93c15b · outbound
CogVLM2: Visual Language Models for Image and Video Understanding Unresolved cited work
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation a09a5b4c-0772-46d4-9811-b8b2837694c9 · outbound
CogVLM2: Visual Language Models for Image and Video Understanding LVBench: An Extreme Long Video Understanding Benchmark
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation defb3e9e-c668-49ce-91df-cbdba7e44a1d · outbound
CogVLM2: Visual Language Models for Image and Video Understanding CogVLM: Visual Expert for Pretrained Language Models
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8c54c02a-278b-4aa7-a14d-fae33148d94d · outbound
CogVLM2: Visual Language Models for Image and Video Understanding Unresolved cited work
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 38dd3c32-185c-47ea-9826-8967c0e5108a · outbound
CogVLM2: Visual Language Models for Image and Video Understanding Unresolved cited work
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d2d8a7ec-ac15-4a91-8c1d-9bbc1e9bd6b4 · outbound
CogVLM2: Visual Language Models for Image and Video Understanding Unresolved cited work
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 5e0aa4f6-4677-4cb0-bb0a-a5f35afd971d · outbound
CogVLM2: Visual Language Models for Image and Video Understanding Unresolved cited work
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation cf4a475c-3692-42f6-8ae1-c03796870c41 · outbound
CogVLM2: Visual Language Models for Image and Video Understanding MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 7bfaa28d-75c4-4274-8a50-bab90447ddc4 · outbound
CogVLM2: Visual Language Models for Image and Video Understanding Unresolved cited work
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e8073d0e-b871-4d18-8e33-d78f836ab8aa · outbound
CogVLM2: Visual Language Models for Image and Video Understanding A Survey on Multimodal Large Language Models
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6e246c8e-b37c-4fd6-b63b-d90deedb94c1 · outbound
CogVLM2: Visual Language Models for Image and Video Understanding MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 5f9bf804-442e-4c99-b2d8-6d45bb7631a1 · outbound
CogVLM2: Visual Language Models for Image and Video Understanding Unresolved cited work
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d2753ffe-2fff-48c6-ad30-819a8c2e8589 · outbound
CogVLM2: Visual Language Models for Image and Video Understanding Unresolved cited work
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 1e3d03fb-344b-4f66-8db2-158f0e57127a · outbound
CogVLM2: Visual Language Models for Image and Video Understanding Zhang, F
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 2027a0c2-8aa6-414c-8543-c10731b1bfbc · outbound
CogVLM2: Visual Language Models for Image and Video Understanding Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4eef3165-d815-461c-9d58-baed95544be4 · outbound
CogVLM2: Visual Language Models for Image and Video Understanding Zhang, P
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 7a856c60-8a1e-498a-be79-9ab924c70cf5 · outbound
CogVLM2: Visual Language Models for Image and Video Understanding VCR: A Task for Pixel-Level Complex Reasoning in Vision Language Models via Restoring Occluded Text
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 426d7758-4284-4e31-a0cb-05e65d850e1f · outbound
CogVLM2: Visual Language Models for Image and Video Understanding It features a blue and white circular design with a black border
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e51ea41b-eae4-428d-8390-f18a89c42103 · outbound
CogVLM2: Visual Language Models for Image and Video Understanding It consists of a stylized letter "Q" enclosed within a circular shape
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 36b71b83-e388-4da3-884a-785ad178df4d · outbound
CogVLM2: Visual Language Models for Image and Video Understanding N/A" instead).{
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ae03e68c-0b35-47c4-8a43-06f9a44a98fb · inbound
LVBench: An Extreme Long Video Understanding Benchmark CogVLM2: Visual Language Models for Image and Video Understanding
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation bfb2e576-3c5b-4b29-9add-7a0e9196be4d · inbound
CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer CogVLM2: Visual Language Models for Image and Video Understanding
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 3ed25625-ffcc-4338-b771-77e891f93c84 · inbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models CogVLM2: Visual Language Models for Image and Video Understanding
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation eb35f914-a198-4dbf-9193-7bd2b7f97aef · inbound
PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling CogVLM2: Visual Language Models for Image and Video Understanding
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 80bab832-fcb6-443e-9f08-8096876b0a1b · inbound
S$^4$ST: A Strong, Self-transferable, faSt, and Simple Scale Transformation for Transferable Targeted Attack CogVLM2: Visual Language Models for Image and Video Understanding
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4cb90199-14d9-4d0b-b297-647d44f44a60 · inbound
ICT: Image-Object Cross-Level Trusted Intervention for Mitigating Object Hallucination in Large Vision-Language Models CogVLM2: Visual Language Models for Image and Video Understanding
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7e92756-c100-41ac-9710-cc5b51013b8f · inbound
All Seeds Are Not Equal: Enhancing Compositional Text-to-Image Generation with Reliable Random Seeds CogVLM2: Visual Language Models for Image and Video Understanding
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 542f5c65-6096-4cf0-8097-4748a8434ca1 · inbound
OpenHumanVid: A Large-Scale High-Quality Dataset for Enhancing Human-Centric Video Generation CogVLM2: Visual Language Models for Image and Video Understanding
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a21b01a0-49ff-4cb7-b178-40b54e03990d · inbound
Mimir: Improving Video Diffusion Models for Precise Text Understanding CogVLM2: Visual Language Models for Image and Video Understanding
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8bcb7c49-8e65-4f33-9d51-52465e746d80 · inbound
SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding CogVLM2: Visual Language Models for Image and Video Understanding
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18c8ea42-2691-4d54-9308-bda11e8539d9 · inbound
p-MoD: Building Mixture-of-Depths MLLMs via Progressive Ratio Decay CogVLM2: Visual Language Models for Image and Video Understanding
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13143aad-217b-4e29-aa85-c10d28e9c651 · inbound
Uni-NaVid: A Video-based Vision-Language-Action Model for Unifying Embodied Navigation Tasks CogVLM2: Visual Language Models for Image and Video Understanding
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 1b5bf52b-6207-44ad-b136-1fa758a80eb3 · inbound
MIT-10M: A Large Scale Parallel Corpus of Multilingual Image Translation CogVLM2: Visual Language Models for Image and Video Understanding
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22c75973-7a83-4169-89d2-a9460907a033 · inbound
Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining CogVLM2: Visual Language Models for Image and Video Understanding
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd200a4d-45e3-466b-8dc2-98ea4bf95dbf · inbound
CogNav: Cognitive Process Modeling for Object Goal Navigation with LLMs CogVLM2: Visual Language Models for Image and Video Understanding
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8fa3bc7-3dbc-485e-8264-4a12f7377600 · inbound
VCA: Video Curious Agent for Long Video Understanding CogVLM2: Visual Language Models for Image and Video Understanding
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3853f0ae-45c3-44b3-a537-60687707fd4d · inbound
PunchBench: Benchmarking MLLMs in Multimodal Punchline Comprehension CogVLM2: Visual Language Models for Image and Video Understanding
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b3b0cee-ec53-4083-ba9e-33e21ce4bb7f · inbound
HumanVBench: Probing Human-Centric Video Understanding in MLLMs with Automatically Synthesized Benchmarks CogVLM2: Visual Language Models for Image and Video Understanding
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 28fe57f7-70d8-45a3-a09d-57d8f79cf6d7 · inbound
VisionReward: Fine-Grained Multi-Dimensional Human Preference Learning for Image and Video Generation CogVLM2: Visual Language Models for Image and Video Understanding
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e17633e3-7103-4a5c-b930-5c03bcfedcc7 · inbound
EliGen: Entity-Level Controlled Image Generation with Regional Attention CogVLM2: Visual Language Models for Image and Video Understanding
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d343bc22-7f2c-4a94-b0f6-9c11bb61ee7d · inbound
VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction CogVLM2: Visual Language Models for Image and Video Understanding
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c33e27e4-ab57-4af2-a9f3-b12bea3ee231 · inbound
MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models CogVLM2: Visual Language Models for Image and Video Understanding
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation aa05e9a9-14ce-4327-bdfd-caa8d1658aa0 · inbound
Automated Generation of Challenging Multiple-Choice Questions for Vision Language Model Evaluation CogVLM2: Visual Language Models for Image and Video Understanding
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0be945d-b4ad-488d-954a-f4da9bd4ba9b · inbound
ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions CogVLM2: Visual Language Models for Image and Video Understanding
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d921696-ed34-4e29-871d-1a83a63f7f06 · inbound
VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding CogVLM2: Visual Language Models for Image and Video Understanding
Reference 153
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 24c36bcd-dbb8-49ab-a937-3ac590d86abb · inbound
Parameter-Efficient Fine-Tuning for Foundation Models CogVLM2: Visual Language Models for Image and Video Understanding
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a421bbdd-b9fd-433a-bc95-a95243aa9862 · inbound
Improving Video Generation with Human Feedback CogVLM2: Visual Language Models for Image and Video Understanding
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 5f20bb0b-ca90-4602-96a2-3f33908ac49d · inbound
Temporal Preference Optimization for Long-Form Video Understanding CogVLM2: Visual Language Models for Image and Video Understanding
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac01ab06-dbdb-4897-a3e3-12b933609d0d · inbound
FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers CogVLM2: Visual Language Models for Image and Video Understanding
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41854b49-82df-4c20-af01-61df17409811 · inbound
RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs CogVLM2: Visual Language Models for Image and Video Understanding
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5152635-448a-4b1e-884c-1d77e25ade10 · inbound
MQuant: Unleashing the Inference Potential of Multimodal Large Language Models via Full Static Quantization CogVLM2: Visual Language Models for Image and Video Understanding
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ccf3b396-b07e-4a72-b983-5689db8bcda2 · inbound
Se\~norita-2M: A High-Quality Instruction-based Dataset for General Video Editing by Video Specialists CogVLM2: Visual Language Models for Image and Video Understanding
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b749407b-50fe-48b0-a978-ce8c3d8f9694 · inbound
EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering CogVLM2: Visual Language Models for Image and Video Understanding
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d396599b-5e37-499d-a6da-3a60182807c4 · inbound
When 'YES' Meets 'BUT': Can Large Models Comprehend Contradictory Humor Through Comparative Reasoning? CogVLM2: Visual Language Models for Image and Video Understanding
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 41cec851-faf3-49cf-b7bf-c6f8af1032d0 · inbound
AdaMMS: Model Merging for Heterogeneous Multimodal Large Language Models with Unsupervised Coefficient Optimization CogVLM2: Visual Language Models for Image and Video Understanding
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e153ab84-b985-404f-b405-86e22a50c706 · inbound
EVA: Evolving Semantic Adversaries for Red-Teaming GUI Agents Against Environmental Injection Attacks CogVLM2: Visual Language Models for Image and Video Understanding
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6425c5f5-2628-4408-8271-946a7a686471 · inbound
TimeCausality: Evaluating the Causal Ability in Time Dimension for Vision Language Models CogVLM2: Visual Language Models for Image and Video Understanding
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 572b17a5-5d92-4dd8-82a7-d80b6564bfda · inbound
MDIT-Bench: Evaluating the Dual-Implicit Toxicity in Large Multimodal Models CogVLM2: Visual Language Models for Image and Video Understanding
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 429f532a-e99e-4441-b5b6-32b649924ea2 · inbound
EgoWalk: A Multimodal Dataset for Robot Navigation in the Wild CogVLM2: Visual Language Models for Image and Video Understanding
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 537f7211-5ccc-4a13-9826-b8b7293b8a29 · inbound
Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning? CogVLM2: Visual Language Models for Image and Video Understanding
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 5426ae40-e7e6-4ba6-bb60-6b115f0d2f4a · inbound
Zero-Shot 3D Visual Grounding from Vision-Language Models CogVLM2: Visual Language Models for Image and Video Understanding
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a683d7db-d212-4e2f-92ee-460c5bb7e6bc · inbound
RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction CogVLM2: Visual Language Models for Image and Video Understanding
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91c67106-19dc-4238-9ca3-0e6c841ed35d · inbound
VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation CogVLM2: Visual Language Models for Image and Video Understanding
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f617f3a5-e179-45cd-86e8-ac1931bcd33a · inbound
MiniMax-Remover: Taming Bad Noise Helps Video Object Removal CogVLM2: Visual Language Models for Image and Video Understanding
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d794475-5b78-48a6-9ad6-cbca8d7d57bb · inbound
GOBench: Benchmarking Geometric Optics Generation and Understanding of MLLMs CogVLM2: Visual Language Models for Image and Video Understanding
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33795074-5261-4f2d-a544-74a39440397d · inbound
Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement CogVLM2: Visual Language Models for Image and Video Understanding
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c67805f-ff99-426e-8c56-684306e97bc9 · inbound
IllumiCraft: Unified Geometry and Illumination Diffusion for Controllable Video Generation CogVLM2: Visual Language Models for Image and Video Understanding
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54be0b75-8e85-4d36-95de-86a8a0f15c25 · inbound
MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos CogVLM2: Visual Language Models for Image and Video Understanding
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4fbc802f-525f-457b-a1cc-6b238be05542 · inbound
LayerFlow: A Unified Model for Layer-aware Video Generation CogVLM2: Visual Language Models for Image and Video Understanding
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46c9b6b1-1bdf-47dc-9e13-ea4ac9dc1f34 · inbound
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs CogVLM2: Visual Language Models for Image and Video Understanding
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f76c8697-c3b6-44c8-8c19-bafeb7f3a54f · inbound
Cross-Frame Representation Alignment for Fine-Tuning Video Diffusion Models CogVLM2: Visual Language Models for Image and Video Understanding
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75504491-fbe0-40f5-9b1f-cbf7465bca9c · inbound
VideoMat: Extracting PBR Materials from Video Diffusion Models CogVLM2: Visual Language Models for Image and Video Understanding
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19ba0531-62db-4840-8e20-4266b6af1531 · inbound
PARC: A Quantitative Framework Uncovering the Symmetries within Vision Language Models CogVLM2: Visual Language Models for Image and Video Understanding
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fdf1a535-d8e8-4c25-858e-fde0cda07149 · inbound
GenRecal: Generation after Recalibration from Large to Small Vision-Language Models CogVLM2: Visual Language Models for Image and Video Understanding
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae57d9f2-bfa5-4289-ad01-016bac8bedbb · inbound
AnyAni: An Interactive System with Generative AI for Animation Effect Creation and Code Understanding in Web Development CogVLM2: Visual Language Models for Image and Video Understanding
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8cb5907f-1712-434a-b50f-90b64a967961 · inbound
CAI: Caption-Sensitive Attention Intervention for Mitigating Object Hallucination in Large Vision-Language Models CogVLM2: Visual Language Models for Image and Video Understanding
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c18bd244-ec45-4b50-8fdd-e5e641b6b1d3 · inbound
GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning CogVLM2: Visual Language Models for Image and Video Understanding
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation cf4e2d1c-22da-461c-97af-ba5d821a63f1 · inbound
LongAnimation: Long Animation Generation with Dynamic Global-Local Memory CogVLM2: Visual Language Models for Image and Video Understanding
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee2df0b5-ed51-4e1a-8ae4-659f11bf2e08 · inbound
Ascending the Infinite Ladder: Benchmarking Spatial Deformation Reasoning in Vision-Language Models CogVLM2: Visual Language Models for Image and Video Understanding
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe18d58a-759a-4aeb-a4ab-e676cca4ece0 · inbound
From Answers to Rationales: Self-Aligning Multimodal Reasoning with Answer-Oriented Chain-of-Thought CogVLM2: Visual Language Models for Image and Video Understanding
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af92a477-8330-412a-944c-438aea6291d5 · inbound
ExpStar: Towards Automatic Commentary Generation for Multi-discipline Scientific Experiments CogVLM2: Visual Language Models for Image and Video Understanding
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1c538fe-0dbf-4454-9c19-4e70034d0425 · inbound
Foundation Model Driven Robotics: A Comprehensive Review CogVLM2: Visual Language Models for Image and Video Understanding
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b0f739b-1099-42e8-ad53-78744aec49e6 · inbound
FantasyPortrait: Enhancing Multi-Character Portrait Animation with Expression-Augmented Diffusion Transformers CogVLM2: Visual Language Models for Image and Video Understanding
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f560e7b0-eef2-4c07-84fb-f2935996c012 · inbound
Hate in Plain Sight: On the Risks of Moderating AI-Generated Hateful Illusions CogVLM2: Visual Language Models for Image and Video Understanding
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88a1cdb1-3f66-4c60-9aaf-442fd10cca4d · inbound
CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text CogVLM2: Visual Language Models for Image and Video Understanding
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2590721-641b-4a1e-a652-db841f948b71 · inbound
Detailed radial scale height profile of dust grains as probed by dust self-scattering in HL Tau CogVLM2: Visual Language Models for Image and Video Understanding
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb9829c8-d950-44cf-a218-957407b31b1f · inbound
SketchAgent: Generating Structured Diagrams from Hand-Drawn Sketches CogVLM2: Visual Language Models for Image and Video Understanding
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf623911-4578-4e0b-aae1-804bbb548cce · inbound
CTA-Flux: Integrating Chinese Cultural Semantics into High-Quality English Text-to-Image Communities CogVLM2: Visual Language Models for Image and Video Understanding
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7d8a105-9d58-482a-a151-9c16b8c07adb · inbound
Improving Large Vision and Language Models by Learning from a Panel of Peers CogVLM2: Visual Language Models for Image and Video Understanding
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b30dc52-6e55-4609-94cd-8a3165b86594 · inbound
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data CogVLM2: Visual Language Models for Image and Video Understanding
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f19733e0-5fd9-486a-93cb-e830037cad92 · inbound
EchoVLM: Dynamic Mixture-of-Experts Vision-Language Model for Universal Ultrasound Intelligence CogVLM2: Visual Language Models for Image and Video Understanding
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation a0652a44-4016-42ae-a5d1-601fca47aee9 · inbound
EchoVLM: Dynamic Mixture-of-Experts Vision-Language Model for Universal Ultrasound Intelligence CogVLM2: Visual Language Models for Image and Video Understanding
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3215cbe9-0022-45f5-ba4b-b5c0262f5d43 · inbound
FG-CLIP 2: A Bilingual Fine-grained Vision-Language Alignment Model CogVLM2: Visual Language Models for Image and Video Understanding
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65acc6f3-cd13-4534-b900-bac1388877c0 · inbound
Geometry-Aware Single-Image 4D Synthesis via Dense Trajectory Generation CogVLM2: Visual Language Models for Image and Video Understanding
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc137128-489f-4e02-a161-27f53335a197 · inbound
High-Entropy Tokens as Multimodal Failure Points in Vision-Language Models CogVLM2: Visual Language Models for Image and Video Understanding
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 671de08b-2c86-4874-98ed-33ccf864a62e · inbound
High-Entropy Tokens as Multimodal Failure Points in Vision-Language Models CogVLM2: Visual Language Models for Image and Video Understanding
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08af0cf1-b685-4976-8e52-8b998fb76eea · inbound
HaineiFRDM: Structure-Preserving Diffusion for Film Restoration under Fast Motion and Diverse Defects CogVLM2: Visual Language Models for Image and Video Understanding
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4af9e13-2c93-44b5-aa23-edd890386e92 · inbound
VideoGPA: Distilling Geometry Priors for 3D-Consistent Video Generation CogVLM2: Visual Language Models for Image and Video Understanding
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8ff2be2d-ac4b-4450-ba60-3fa3a91e10ae · inbound
VideoGPA: Distilling Geometry Priors for 3D-Consistent Video Generation CogVLM2: Visual Language Models for Image and Video Understanding
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bdf92923-e799-48b2-b36a-69a32bcd5fdb · inbound
EgoIntent: A Pre-Outcome Micro-Step Benchmark for Understanding What, Why, and Next CogVLM2: Visual Language Models for Image and Video Understanding
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5153d1e-8319-4c13-a2fa-345e69018981 · inbound
MIRAGE: Benchmarking and Aligning Multi-Instance Image Editing CogVLM2: Visual Language Models for Image and Video Understanding
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 69ea4b6d-93e6-4b83-9c97-c02f0adac2ca · inbound
VersaVogue: Visual Expert Orchestration and Preference Alignment for Unified Fashion Synthesis CogVLM2: Visual Language Models for Image and Video Understanding
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 92d96f62-5ef7-4f2a-96fa-6b2728cd23c4 · inbound
VSAS-Bench: Real-Time Evaluation of Visual Streaming Assistant Models CogVLM2: Visual Language Models for Image and Video Understanding
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 143042a3-3692-488a-9295-822a4320a670 · inbound
Towards Unconstrained Human-Object Interaction CogVLM2: Visual Language Models for Image and Video Understanding
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 04e81a59-7c10-4f66-9e06-c8539e73423a · inbound
Culture-Aware Humorous Captioning: Multimodal Humor Generation across Cultural Contexts CogVLM2: Visual Language Models for Image and Video Understanding
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 3dfaeb8b-e91b-4920-a9e8-266660ee8bef · inbound
ReCAPA: Hierarchical Predictive Correction to Mitigate Cascading Failures CogVLM2: Visual Language Models for Image and Video Understanding
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 097dc07a-3f65-4eec-91b7-e0e9f11de37a · inbound
ReCAPA: Hierarchical Predictive Correction to Mitigate Cascading Failures CogVLM2: Visual Language Models for Image and Video Understanding
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 74f388d3-9f17-4a53-9523-351980ccc556 · inbound
KD-CVG: A Knowledge-Driven Approach for Creative Video Generation CogVLM2: Visual Language Models for Image and Video Understanding
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 790d6ea8-0e57-4dcc-8f9d-3598530e9826 · inbound
CAST: Mitigating Object Hallucination in Large Vision-Language Models via Caption-Guided Visual Attention Steering CogVLM2: Visual Language Models for Image and Video Understanding
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 5dd0961d-d8c1-434c-860c-4f5709476b7a · inbound
Illusion-Aware Visual Preprocessing and Anti-Illusion Prompting for Classic Illusion Understanding in Vision-Language Models CogVLM2: Visual Language Models for Image and Video Understanding
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 31174350-03f7-4a3a-9e1c-b691b3e181df · inbound
Omni-DuplexEval: Evaluating Real-time Duplex Omni-modal Interaction CogVLM2: Visual Language Models for Image and Video Understanding
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 0e39044d-1a37-4bc6-8de4-60c62b310850 · inbound
Omni-DuplexEval: Evaluating Real-time Duplex Omni-modal Interaction CogVLM2: Visual Language Models for Image and Video Understanding
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e07336f4-347c-40a0-9901-4679646b9a37 · inbound
From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models CogVLM2: Visual Language Models for Image and Video Understanding
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b15ad2f6-0690-4fd5-a6e4-d3236130f31c · inbound
MONET: A Massive, Open, Non-redundant and Enriched Text-to-image dataset CogVLM2: Visual Language Models for Image and Video Understanding
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation a43ce7f1-9338-48fa-9550-81cc9503f3ef · inbound
TrackRef3D: Multi-View Consistent Track-then-Label for Open-World Referring Segmentation in 3D Gaussian Splatting CogVLM2: Visual Language Models for Image and Video Understanding
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9d37e3c6-e5de-43f0-82cb-9982e653f131 · inbound
MotionEnhancer: Leveraging Video Diffusion for Motion-Enhanced Vision-Language Models CogVLM2: Visual Language Models for Image and Video Understanding
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation eb8cf754-4144-46a4-82f6-cfbed09c3279 · inbound
PixelEyes: Decoupling Perception and Reasoning for Pinpoint Visual Evidence Seeking CogVLM2: Visual Language Models for Image and Video Understanding
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6de49193-a046-4c4f-b4b4-63212490ad9d · inbound
EgoPolice: A Benchmark for Egocentric Video Understanding in High-Stakes Police Body-Worn Camera Footage CogVLM2: Visual Language Models for Image and Video Understanding
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 0e3f8377-2e3d-4766-97cc-b1d086c7d0c0 · inbound
WeaveEarth: Structured Evidence Construction and Reasoning for Training-Free UHR Remote Sensing Understanding CogVLM2: Visual Language Models for Image and Video Understanding
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e519978-a73f-4260-aef9-789646c7c880 · inbound
Actor as Its Own Critic: Unifying Region Understanding and Localization via CycleGRPO CogVLM2: Visual Language Models for Image and Video Understanding
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.