Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T16:32:43.351429Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 100 of 121 outbound references and 1 inbound Pith citation observation for arXiv:2508.18179.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T16:32:43.351429Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-03T00:55:26.251613Z
A source-named dated measurement, never combined with another source.
Source: cited_works
100 of 121 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 525eff74-1564-420d-9e24-250910d23a05 · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Pixtral 12B
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7c13701-c6ce-4f33-a3c6-e0747a3d2222 · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Flamingo: a visual language model for few-shot learning
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9392d820-17be-4b84-ab4f-116c7d4fdf35 · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models The claude 3 model family: Opus, sonnet, haiku, 2024
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81f11666-1191-4d31-a78d-ccd0d09c9ebd · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Claude 3.7 sonnet system card, 2025 a
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6b6bc5d-0c83-4f22-8acd-3f74deaa4eed · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models System card: Claude opus 4 & claude sonnet 4, 2025 b
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22522544-7446-4faf-b137-7ab1c8a56688 · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Vqa: Visual question answering
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ebe7bf2-9001-4afe-ad5d-3ca79f1dec28 · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19c3fbbf-be30-42ee-a0f4-6b1c0e30c1c3 · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Qwen Technical Report
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 908f5923-476e-4e50-9f6c-1f3d7eb22f97 · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Qwen2.5-VL Technical Report
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d27bcae4-8b22-46ed-b3e6-93138cd72d66 · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6368c1c4-988f-430e-8956-038e765585a1 · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Uniter: Universal image-text representation learning
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4299bd4a-fbe2-4ae1-afd2-876fe3333ec2 · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a261484-5d95-4ead-8c18-0509e1fba46a · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e0a6a27-b13f-48d8-bd7f-4c4f6711677a · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ddd3b3f-e5ba-497b-b01b-9f998cfffb9f · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Chess.com, 2025
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cfcead19-d472-4bd3-8bb4-880af03df0db · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e533b79-c3dc-4cdf-875c-326ccaec9818 · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Opencompass: A universal evaluation platform for foundation models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6a4f47f-e237-4390-bf52-d4f3f140aec1 · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Holistic Analysis of Hallucination in GPT-4V(ision): Bias and Interference Challenges
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 909f5dc2-0451-4b43-85d4-928a6b10d722 · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models music21: A toolkit for computer-aided musicology and symbolic music analysis
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 623cc85f-a0fb-480c-a91e-9d459d15d2e9 · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 042497b8-5e67-4690-a45f-bedec4b34b1f · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Gemini: A Family of Highly Capable Multimodal Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 557dbf17-5f15-44e3-b677-7e18b7fa79c2 · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b548fcf9-34da-403c-a0d9-bc519078c568 · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Gemma: Open Models Based on Gemini Research and Technology
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d40e41b-4df4-4cc2-bd5e-5be2ec409b9b · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Gemma 2: Improving Open Language Models at a Practical Size
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc974796-0076-4805-9ca2-2817667dbec0 · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Introducing gemini 2.0: our new ai model for the agentic era, 2025 a
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a836bc7-e02c-4b25-9585-f169472e5801 · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Gemma 3 technical report, 2025 b
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d96fd3f-bb25-4d3b-8d93-ca52ab6b8a85 · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Portable game notation specification and implementation guide, 1994
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3be8c70-b887-4b2a-94e1-ee46191f300f · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Pmr: Prototypical modal rebalance for multimodal learning
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1bba78f3-bf2a-4d2f-8492-81413d2aaa6c · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Layoutgpt: Compositional visual planning and generation with large language models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50fd2ba7-bd79-455c-b4f2-a8ccf058a8a9 · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models python-chess: A chess library for python, 2025
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed43d763-ea8c-4457-b4ef-7bfcbe69bd21 · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Blink: Multimodal large language models can see but not perceive
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 663e06fa-ef94-4a8c-84ef-ed5e5653fea9 · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5c2725d-325d-40b0-883e-9efcdd44eae3 · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1096c998-ce92-4cad-b094-3c43de91fa1a · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models The Llama 3 Herd of Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b936b9a4-6caf-4baa-a381-a1514f502bd4 · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models What can large language models do in chemistry? a comprehensive benchmark on eight tasks
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d11e22d-3154-4b94-bbf3-74e16e85c72c · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Exploring network structure, dynamics, and function using networkx
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb84384c-ec52-4cbc-a89d-46289c2dd6be · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models ChartLlama: A Multimodal LLM for Chart Understanding and Generation
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d5b0801-966a-45f1-9e27-1ecb18d0c745 · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea38c2f4-6165-4719-8c34-37e37b7bb1a4 · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Deciphering Cross-Modal Alignment in Large Vision-Language Models with Modality Integration Rate
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd51e676-1cdc-497e-864d-136105f4a8f3 · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Modality competition: What makes joint training of multi-modal network fail in deep learning?(provably)
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8cd357c3-b2bf-436c-a51e-d3bd53657c32 · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models MANTIS: Interleaved Multi-Image Instruction Tuning
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81bf7ea1-e689-4002-a7fd-4b3695743f52 · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Spin: Sparsifying and integrating internal neurons in large language models for text classification
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7077a98-198b-45e1-b241-c10348bd59ba · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Pubchem 2025 update
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f53ec5da-aa9a-4c56-b10b-808a6a046c4b · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Large language models are zero-shot reasoners
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6bf091bc-e6df-4375-b0ba-726f35f368dc · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Learning multiple layers of features from tiny images
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5927e38e-4511-4f55-a206-f1a6aa216002 · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Gonzalez, Hao Zhang, and Ion Stoica
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67eeec96-0dd9-40a7-a7a1-aeed20b5168c · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Snap: A general-purpose network analysis and graph-mining library
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65efaed9-bf46-43f4-a761-9ef8ee62548d · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models LLaVA-OneVision: Easy Visual Task Transfer
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b7f39c6-b46e-49f8-b175-1e9ee5e70d0b · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Seed-bench: Benchmarking multimodal large language models
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8c289cd-b6e4-4188-bf6e-c8e9db1b9195 · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3316da2c-5ec2-4606-93c8-29592ec1201f · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e27a4424-be51-4d4c-85e0-bdeb6965c1b0 · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5131e450-4984-4ce0-85b1-f3b7bdfb6127 · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Oscar: Object-semantics aligned pre-training for vision-language tasks
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 504b61d1-6c62-438d-870e-bb3cf0d968fe · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Lichess evaluation database, 2025 a
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05561b7a-905a-4928-9308-9fd88d1928a2 · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Lichess: Free online chess, 2025 b
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 830962d1-0d78-42cf-a6fd-e832264415a5 · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Lichess puzzle database, 2025 c
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1da1321c-6954-4054-bf68-99d4958d54bd · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Microsoft coco: Common objects in context
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41cd2bdc-7836-45af-b3ac-248c134213fa · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd0db3e4-c09f-463f-82dc-9ac001915bfa · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Improved baselines with visual instruction tuning, 2023 b
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f01e33cf-2675-40c4-b4c8-a5a2f57928d9 · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Visual instruction tuning, 2023 c
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01222fc1-4d5b-4862-9035-864045560d18 · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Llava-next: Improved reasoning, ocr, and world knowledge, January 2024 a
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c38a33c-f9d0-4df9-bac2-6b1044cf1a3e · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Mmbench: Is your multi-modal model an all-around player? In European conference on computer vision, pp.\ 216--233
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d0d69f5-1320-4526-b7e3-7e7e7976f46a · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5ce209a-9bee-4fc2-b001-34eefa640d60 · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 305e7673-8562-4b24-9930-1ccd5454ec3b · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Ok-vqa: A visual question answering benchmark requiring external knowledge
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d87dc009-474d-4dce-8a68-80272dfd1454 · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Gpt-4v(ision) system card, 2023
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 3380c342-0974-4bcc-80b2-1ce4e181596a · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models GPT-4o System Card
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c60b406e-824d-4d6c-9ace-be98ff1e8954 · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Gpt-4o mini: Advancing cost-efficient intelligence, 2024 b
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation bd1b13ce-edee-4e6a-a115-d94fc2f79a8a · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Openai o1 system card, December 2024 c
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e9fcfc15-f111-494e-bd35-a5d55318156d · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Gpt-4.5 system card, 2025 a
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 5d283bba-1925-4d51-9066-9845a0fb8af7 · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Gpt-5 system card, 2025 b
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation dff49c99-74e0-4a3d-9c22-846f2ce032c1 · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Openai o3-mini system card, February 2025 c
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 59195a71-c2e5-4545-9cc0-44a80e816065 · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Openai o3 and o4-mini system card, 2025 d
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation aecaf710-6ce3-4431-9d76-183a18a5bf07 · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Cross-modal Attention Congruence Regularization for Vision-Language Relation Alignment
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 1102cc23-3b0f-46de-8988-d3a531a0f46d · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs?
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e58ada7-be1e-4694-8a90-9b7dc3640228 · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Balanced multimodal learning via on-the-fly gradient modulation
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c86ab290-9ad4-4a06-9a8a-5cfe6ac0e7dd · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Rdkit: Open-source cheminformatics, 2025
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 030b90b6-00d8-4128-8338-84030711a36f · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models The music encoding initiative (mei)
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 717576f2-79b3-4890-8ff8-b1ad6d093217 · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Music abc notation with music theory dataset, 2025
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 1ead4494-65b0-47ce-81c4-edce3d41d116 · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models LXMERT: Learning Cross-Modality Encoder Representations from Transformers
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7bfb62e6-aa72-44b2-ab6d-5f3d298cc645 · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models NOTA: Multimodal Music Notation Understanding for Visual Large Language Model
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a33936c0-6748-4474-a156-d578b814e4cb · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Cambrian-1: A fully open, vision-centric exploration of multimodal llms
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 4f4a7637-c902-492c-8510-17f51cdf7608 · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models LLaMA: Open and Efficient Foundation Language Models
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5e8a843-d464-4c35-a635-ef36e6506200 · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9db96dbe-cb7c-44e1-816d-70403c041b18 · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Towards generalist biomedical ai
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bcd0eec3-4e6b-426a-995d-ffb50898b36d · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Visualizing data using t-sne
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6325104-99ae-4d35-ac8a-7ef7ed205511 · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models LIMCA: LLM for Automating Analog In-Memory Computing Architecture Design Exploration
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7364adca-5cf4-40f5-bb5d-e5596374413b · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Abc notation standard, 2004
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 9997b4c6-3a97-4216-b471-9464fb84e6e2 · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Is a picture worth a thousand words? delving into spatial reasoning for vision language models
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 62b0c2fd-bb29-43ab-998c-45fd3d7599f8 · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Multilingual E5 Text Embeddings: A Technical Report
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation efab14d9-2a73-482c-a6b2-3b0a7ed6279c · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f7a46f2-a96e-4f35-8c92-1046bfe948a6 · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Smiles, a chemical language and information system
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 388eac7b-f4f0-4743-b579-35c03b1ed15b · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Tunesformer: Forming irish tunes with control codes by bar patching
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 4a82fe10-49e3-4c20-8230-99660a8f39ed · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models SciMMIR: Benchmarking Scientific Multi-modal Information Retrieval
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fcd16cd4-e99e-429c-be21-0192a690c04b · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Qwen2.5-Omni Technical Report
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7631358-7355-417e-b85b-f59ba3d580eb · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Lvlm-ehub: A comprehensive evaluation benchmark for large vision-language models
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 922727b0-2cc7-4087-acb8-f82ec6c17722 · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Qwen2 Technical Report
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f4aed58-150c-4fef-8a03-127c4dbe76ab · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Qwen2.5 Technical Report
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a76350f-abfb-43f5-87c0-66dfb21b3e17 · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Report Cards: Qualitative Evaluation of Language Models Using Natural Language Summaries
Reference 99
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c13d70d0-f14c-4ed2-a0d3-a8795c32a985 · outbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 351a9023-0a61-4de3-8da0-72858a7b8fe0 · inbound
TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.