Pith. sign in

Paper Citation Record · LEDGER

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models

As of 18 August 2026, this Paper Citation Record lists 100 of 121 outbound references and 1 inbound Pith citation observation for arXiv:2508.18179.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.18179 v1

Coverage vector

measured 100 of 121 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T16:32:43.351429Z

measured 101 of 101 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T00:55:26.251613Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 121 outbound references displayed

  • verified exact2
  • verified fuzzy18
  • unresolved80
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 525eff74-1564-420d-9e24-250910d23a05 · outbound

This paper cites Pixtral 12B.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Pixtral 12B

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:42.044547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:32:42.044547Z digest=sha256:adad9e4ece9a144e24f3ed6c150de40187d8f93af34cec9f5ae6b68384603de2

Observation f7c13701-c6ce-4f33-a3c6-e0747a3d2222 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Flamingo: a visual language model for few-shot learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:42.161491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:32:42.161491Z digest=sha256:a3d04fc16b2d01155e3fb00e2bcd6a313ca36c774f74d831158edeedbc5dae03

Observation 9392d820-17be-4b84-ab4f-116c7d4fdf35 · outbound

This paper cites The claude 3 model family: Opus, sonnet, haiku, 2024.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models The claude 3 model family: Opus, sonnet, haiku, 2024

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:42.248888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:32:42.248888Z digest=sha256:1eada138ba83b55e26ea62c843056a64292413bfff87f523a12eae26a5f0647d

Observation 81f11666-1191-4d31-a78d-ccd0d09c9ebd · outbound

This paper cites Claude 3.7 sonnet system card, 2025 a.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Claude 3.7 sonnet system card, 2025 a

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:42.328402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:32:42.328402Z digest=sha256:9b80778f94a4fbe748fa4f4b633d74c85f45c69874b04e73680e5831fa7ffbf1

Observation d6b6bc5d-0c83-4f22-8acd-3f74deaa4eed · outbound

This paper cites System card: Claude opus 4 & claude sonnet 4, 2025 b.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models System card: Claude opus 4 & claude sonnet 4, 2025 b

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:42.392037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:32:42.392037Z digest=sha256:fd5fc8ede376d814d41a7e953525dc51e2df6b7ebcf14401f64c04d17a192cc9

Observation 22522544-7446-4faf-b137-7ab1c8a56688 · outbound

This paper cites Vqa: Visual question answering.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Vqa: Visual question answering

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:42.477810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:32:42.477810Z digest=sha256:f80447cf6b4b4e95ed33e3145f4422d4ee4241b936b335f9ee1b3bcc62b8a9b4

Observation 2ebe7bf2-9001-4afe-ad5d-3ca79f1dec28 · outbound

This paper cites OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:42.555428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:32:42.555428Z digest=sha256:6377c83c94baba287da54b4edfb9e60c79387ba6a300b10f2c9e6d77e913b342

Observation 19c3fbbf-be30-42ee-a0f4-6b1c0e30c1c3 · outbound

This paper cites Qwen Technical Report.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Qwen Technical Report

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:42.617018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:32:42.617018Z digest=sha256:a9719f575204f067f96d62f57a98c69ca1606b90ac6503143ec7e5d527172ea2

Observation 908f5923-476e-4e50-9f6c-1f3d7eb22f97 · outbound

This paper cites Qwen2.5-VL Technical Report.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Qwen2.5-VL Technical Report

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:42.703832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:32:42.703832Z digest=sha256:96aff94c4c79b00fcd32fbb047cc11572e1daed524f0369dbfd523cb2abff745

Observation d27bcae4-8b22-46ed-b3e6-93138cd72d66 · outbound

This paper cites MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:42.732195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:32:42.732195Z digest=sha256:7524342d2f14cfc3cdb54b6f46f265d06e9c2b7a00ad2f3e2e3e54bcc042d572

Observation 6368c1c4-988f-430e-8956-038e765585a1 · outbound

This paper cites Uniter: Universal image-text representation learning.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Uniter: Universal image-text representation learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:42.738792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:32:42.738792Z digest=sha256:6c1dab44da49e236f4f825e4471fb122b5823a4eb50683c9666638cf86114cb1

Observation 4299bd4a-fbe2-4ae1-afd2-876fe3333ec2 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:42.744491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:32:42.744491Z digest=sha256:f9dab81708b0e5e980136803be910330d3e0cb12e767eb5f2035a0abfdcc99b9

Observation 3a261484-5d95-4ead-8c18-0509e1fba46a · outbound

This paper cites How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:42.750413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:32:42.750413Z digest=sha256:ade02e03548b0006cf6246ff3f5a938d4c3e17d793dd0c52f1465ef19ecff09e

Observation 6e0a6a27-b13f-48d8-bd7f-4c4f6711677a · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:42.757959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:32:42.757959Z digest=sha256:0062aeb6929874f19753b2c7076457dfe9890ea203bccecf2d034408c298fd3c

Observation 1ddd3b3f-e5ba-497b-b01b-9f998cfffb9f · outbound

This paper cites Chess.com, 2025.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Chess.com, 2025

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:42.764237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:32:42.764237Z digest=sha256:1329b76f64fe5bc9ebeea618657530b2e30c082963617acddefd367679b7ca06

Observation cfcead19-d472-4bd3-8bb4-880af03df0db · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:42.770937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:32:42.770937Z digest=sha256:9cd7a7ef0fcdb0740688bc1a224083819c3bbab02a311162529b8f6713cf9712

Observation 9e533b79-c3dc-4cdf-875c-326ccaec9818 · outbound

This paper cites Opencompass: A universal evaluation platform for foundation models.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Opencompass: A universal evaluation platform for foundation models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:42.777528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:32:42.777528Z digest=sha256:4a3457e00f6877cca8e4523e20e3ae1373b1cb1c035ad5cc3cacf154056aa975

Observation f6a4f47f-e237-4390-bf52-d4f3f140aec1 · outbound

This paper cites Holistic Analysis of Hallucination in GPT-4V(ision): Bias and Interference Challenges.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Holistic Analysis of Hallucination in GPT-4V(ision): Bias and Interference Challenges

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:42.784068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:32:42.784068Z digest=sha256:0564b2b636e74ddacfd4c043b3fa3c7d326edc2fe306de090982aba88903e2d1

Observation 909f5dc2-0451-4b43-85d4-928a6b10d722 · outbound

This paper cites music21: A toolkit for computer-aided musicology and symbolic music analysis.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models music21: A toolkit for computer-aided musicology and symbolic music analysis

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:42.790967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:32:42.790967Z digest=sha256:38318d227ac700374d4ce205a583862ee5cd58ba59ff222d0ee51dae75590ac8

Observation 623cc85f-a0fb-480c-a91e-9d459d15d2e9 · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:42.800336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:32:42.800336Z digest=sha256:8aedaca3c35acb4c189c6a68a27c3cb2786ccf129bd4934985a66707a7cbb06e

Observation 042497b8-5e67-4690-a45f-bedec4b34b1f · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Gemini: A Family of Highly Capable Multimodal Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:42.809612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:32:42.809612Z digest=sha256:57c3e85683279ac06f58a70ec344ae99706dee91bac15a6a32f4ba726bb0ebe5

Observation 557dbf17-5f15-44e3-b677-7e18b7fa79c2 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:42.817100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:32:42.817100Z digest=sha256:d4739aae77691f3e7ecb229f06aedf17fa9baa517a0233edc24212286bcdd8dc

Observation b548fcf9-34da-403c-a0d9-bc519078c568 · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Gemma: Open Models Based on Gemini Research and Technology

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:42.823442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:32:42.823442Z digest=sha256:e5fae508469766d3ca0424b0480eebe381bb476adbaa93bbe67d8f498e569098

Observation 2d40e41b-4df4-4cc2-bd5e-5be2ec409b9b · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Gemma 2: Improving Open Language Models at a Practical Size

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:42.831534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:32:42.831534Z digest=sha256:0273ddec2ceff728ca11b7a7cb2f174e63eea8022684a8e26810be3015a70c07

Observation dc974796-0076-4805-9ca2-2817667dbec0 · outbound

This paper cites Introducing gemini 2.0: our new ai model for the agentic era, 2025 a.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Introducing gemini 2.0: our new ai model for the agentic era, 2025 a

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:42.848467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:32:42.848467Z digest=sha256:dcdc4ff883f1a98ae09fee595cf88efd691f191044ad7ed6a05b4a6cf5683207

Observation 5a836bc7-e02c-4b25-9585-f169472e5801 · outbound

This paper cites Gemma 3 technical report, 2025 b.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Gemma 3 technical report, 2025 b

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:42.854575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:32:42.854575Z digest=sha256:5adb808edadba0017d2497f5f2029f7ef749b107252e8fb6d7ef22a27b2f1cdb

Observation 6d96fd3f-bb25-4d3b-8d93-ca52ab6b8a85 · outbound

This paper cites Portable game notation specification and implementation guide, 1994.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Portable game notation specification and implementation guide, 1994

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:42.864402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:32:42.864402Z digest=sha256:bd15b73c44022fe9792b3c8f895ea80b76521939f3d0a90979461980eea11dd0

Observation b3be8c70-b887-4b2a-94e1-ee46191f300f · outbound

This paper cites Pmr: Prototypical modal rebalance for multimodal learning.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Pmr: Prototypical modal rebalance for multimodal learning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:42.870783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:32:42.870783Z digest=sha256:fc87005e9411176c54d65dcbef0ca4e75e362f5da7e2fece09982bc811357870

Observation 1bba78f3-bf2a-4d2f-8492-81413d2aaa6c · outbound

This paper cites Layoutgpt: Compositional visual planning and generation with large language models.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Layoutgpt: Compositional visual planning and generation with large language models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:42.877320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:32:42.877320Z digest=sha256:208750d9f5e936ba34af26b3290e7430eb0284647a122277db6c399bf41bf417

Observation 50fd2ba7-bd79-455c-b4f2-a8ccf058a8a9 · outbound

This paper cites python-chess: A chess library for python, 2025.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models python-chess: A chess library for python, 2025

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:42.882769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:32:42.882769Z digest=sha256:84cf61f5ab13409ccd8babf429f75d4e2903cc49a3eb8ddfd9ed0b208c6df5f3

Observation ed43d763-ea8c-4457-b4ef-7bfcbe69bd21 · outbound

This paper cites Blink: Multimodal large language models can see but not perceive.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Blink: Multimodal large language models can see but not perceive

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:42.891709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:32:42.891709Z digest=sha256:b8f46d3b6bdfbe008908d1eb32250f8ee746440688d1984dbbad44cac175a243

Observation 663e06fa-ef94-4a8c-84ef-ed5e5653fea9 · outbound

This paper cites LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:42.899699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:32:42.899699Z digest=sha256:f2c7f0d34b4d77edbcb163cc78dfd293f97d25a316c2cf7abff557448dcf3aa9

Observation c5c2725d-325d-40b0-883e-9efcdd44eae3 · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answering.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Making the v in vqa matter: Elevating the role of image understanding in visual question answering

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:42.904966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:32:42.904966Z digest=sha256:b53571289447e2e08bad608f823caccaa208cae9b9b430b7c23172367e453789

Observation 1096c998-ce92-4cad-b094-3c43de91fa1a · outbound

This paper cites The Llama 3 Herd of Models.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models The Llama 3 Herd of Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:42.912222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:32:42.912222Z digest=sha256:4434e227b8c30d82c77090e52c6a37505a6de67aec06263f3f31aa4b18f610ff

Observation b936b9a4-6caf-4baa-a381-a1514f502bd4 · outbound

This paper cites What can large language models do in chemistry? a comprehensive benchmark on eight tasks.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models What can large language models do in chemistry? a comprehensive benchmark on eight tasks

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:42.917902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:32:42.917902Z digest=sha256:638ae3e4f77d49432a36f2425e37a16377637266e09764a588324cfd3f298ac2

Observation 8d11e22d-3154-4b94-bbf3-74e16e85c72c · outbound

This paper cites Exploring network structure, dynamics, and function using networkx.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Exploring network structure, dynamics, and function using networkx

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:42.922914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:32:42.922914Z digest=sha256:44a15d347b453b8e1fb36a748893aca061be7de5985031b801916a018ef8f096

Observation bb84384c-ec52-4cbc-a89d-46289c2dd6be · outbound

This paper cites ChartLlama: A Multimodal LLM for Chart Understanding and Generation.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models ChartLlama: A Multimodal LLM for Chart Understanding and Generation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:42.928350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:32:42.928350Z digest=sha256:083e1b79cd2652e9373ce2b10b0f2e1ebd207cafc5566481bc3539a21684666f

Observation 9d5b0801-966a-45f1-9e27-1ecb18d0c745 · outbound

This paper cites Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:42.934041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:32:42.934041Z digest=sha256:74a914e3ff587c2e87018eecb25984a963e44bdfd0814216b4f184b5225d3d6f

Observation ea38c2f4-6165-4719-8c34-37e37b7bb1a4 · outbound

This paper cites Deciphering Cross-Modal Alignment in Large Vision-Language Models with Modality Integration Rate.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Deciphering Cross-Modal Alignment in Large Vision-Language Models with Modality Integration Rate

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:42.939167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:32:42.939167Z digest=sha256:58dcd152ca36d6f18a8fa247c4c736c98556c1ba1c7d33b5f084d4b4d03523ce

Observation fd51e676-1cdc-497e-864d-136105f4a8f3 · outbound

This paper cites Modality competition: What makes joint training of multi-modal network fail in deep learning?(provably).

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Modality competition: What makes joint training of multi-modal network fail in deep learning?(provably)

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:42.945106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:32:42.945106Z digest=sha256:3fd9da8031ddc22cbff74a2f570cf256100144d87f60a31499ed65a37a0b5b2c

Observation 8cd357c3-b2bf-436c-a51e-d3bd53657c32 · outbound

This paper cites MANTIS: Interleaved Multi-Image Instruction Tuning.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models MANTIS: Interleaved Multi-Image Instruction Tuning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:42.950450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:32:42.950450Z digest=sha256:db9aed2d002575f6ba089710ccad5e71388facaf75c703d8b362029681924f38

Observation 81bf7ea1-e689-4002-a7fd-4b3695743f52 · outbound

This paper cites Spin: Sparsifying and integrating internal neurons in large language models for text classification.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Spin: Sparsifying and integrating internal neurons in large language models for text classification

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:42.956607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:32:42.956607Z digest=sha256:bf58cdb1dd90d47945e495c0ec4a95264524d43c2a9265a4d6d117fd3589b707

Observation c7077a98-198b-45e1-b241-c10348bd59ba · outbound

This paper cites Pubchem 2025 update.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Pubchem 2025 update

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:42.963041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:32:42.963041Z digest=sha256:9c53fb4457bcefb5afcc390ac3862c53db5abb96ab8f57fe3290d29688a92026

Observation f53ec5da-aa9a-4c56-b10b-808a6a046c4b · outbound

This paper cites Large language models are zero-shot reasoners.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Large language models are zero-shot reasoners

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:42.984383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:32:42.984383Z digest=sha256:7acfcd0f565e4505b9cee5022bdc8e0d7770fb0eb4486f101745003c6e8be953

Observation 6bf091bc-e6df-4375-b0ba-726f35f368dc · outbound

This paper cites Learning multiple layers of features from tiny images.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Learning multiple layers of features from tiny images

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:43.006862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:32:43.006862Z digest=sha256:ab4aa24472aa787d695b2e772f4ce157c1e5fa09ba809fe2d9cc92b6d40c21d0

Observation 5927e38e-4511-4f55-a206-f1a6aa216002 · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Gonzalez, Hao Zhang, and Ion Stoica

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:43.013709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:32:43.013709Z digest=sha256:9f3cc70a0cad7baeabd838994392f63455af3a529737979d8ebc9c2477b8d6db

Observation 67eeec96-0dd9-40a7-a7a1-aeed20b5168c · outbound

This paper cites Snap: A general-purpose network analysis and graph-mining library.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Snap: A general-purpose network analysis and graph-mining library

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:43.020978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:32:43.020978Z digest=sha256:97f24193cbadd695a4a7e5927e5e56cc67839a3403906b36f891e1e20ff75898

Observation 65efaed9-bf46-43f4-a761-9ef8ee62548d · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models LLaVA-OneVision: Easy Visual Task Transfer

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:43.026292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:32:43.026292Z digest=sha256:740e9d2987cce391819e217ef81ada90f86b3fa795392ec16d9363946a55bfcd

Observation 3b7f39c6-b46e-49f8-b175-1e9ee5e70d0b · outbound

This paper cites Seed-bench: Benchmarking multimodal large language models.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Seed-bench: Benchmarking multimodal large language models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:43.033345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:32:43.033345Z digest=sha256:3fd8d5a387e3b631ff94e87e21a9c4844132b3aacfc7f35c5aa43232fce17a96

Observation c8c289cd-b6e4-4188-bf6e-c8e9db1b9195 · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:43.040398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:32:43.040398Z digest=sha256:fed924621a3aee3e7814577b4cfef0eb5adaca517f3a33f33a9a445f0df4fa9a

Observation 3316da2c-5ec2-4606-93c8-29592ec1201f · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:43.045873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:32:43.045873Z digest=sha256:fbaa666d22aa9bd6c96ec2477787d097bf1d0cb05ba11fc893bcd3cef0c3061b

Observation e27a4424-be51-4d4c-85e0-bdeb6965c1b0 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:43.050715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:32:43.050715Z digest=sha256:4b0aaa56438fdc0aa66ae4f7af7f7fad4bf944e8c585e042c8042c439af408a8

Observation 5131e450-4984-4ce0-85b1-f3b7bdfb6127 · outbound

This paper cites Oscar: Object-semantics aligned pre-training for vision-language tasks.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Oscar: Object-semantics aligned pre-training for vision-language tasks

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:43.055662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:32:43.055662Z digest=sha256:b6828bcf2bb44e06250828cf633e254835b9f5c9486105a8239e5391ad9901a0

Observation 504b61d1-6c62-438d-870e-bb3cf0d968fe · outbound

This paper cites Lichess evaluation database, 2025 a.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Lichess evaluation database, 2025 a

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:43.060050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:32:43.060050Z digest=sha256:48e418d729a4bceb66371e20e9b890f82a797ecee8cca2fa0c01a792c49514e0

Observation 05561b7a-905a-4928-9308-9fd88d1928a2 · outbound

This paper cites Lichess: Free online chess, 2025 b.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Lichess: Free online chess, 2025 b

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:43.065022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:32:43.065022Z digest=sha256:b0d00f47adea77df9331b8d384e70abda680dcbf362490c78cc8150a5534df1b

Observation 830962d1-0d78-42cf-a6fd-e832264415a5 · outbound

This paper cites Lichess puzzle database, 2025 c.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Lichess puzzle database, 2025 c

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:43.070040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:32:43.070040Z digest=sha256:b9735ee593c37974c509c6c70c90eacd11b0532633d842f667d01d9a0cf44e53

Observation 1da1321c-6954-4054-bf68-99d4958d54bd · outbound

This paper cites Microsoft coco: Common objects in context.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Microsoft coco: Common objects in context

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:43.074656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:32:43.074656Z digest=sha256:31942edec7e005733282dadea1158981c003b6e50eb257b7abc77bbe35d3b5ca

Observation 41cd2bdc-7836-45af-b3ac-248c134213fa · outbound

This paper cites HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:43.079490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:32:43.079490Z digest=sha256:ceb690294c6ec0723a4d2701d8169797742dec7c6ec112a56ce44a94124b662a

Observation dd0db3e4-c09f-463f-82dc-9ac001915bfa · outbound

This paper cites Improved baselines with visual instruction tuning, 2023 b.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Improved baselines with visual instruction tuning, 2023 b

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:43.086560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:32:43.086560Z digest=sha256:348966c57a6be7240fc4b1cb7359d033abecb48aa4fee08dfbc41979ba0db49c

Observation f01e33cf-2675-40c4-b4c8-a5a2f57928d9 · outbound

This paper cites Visual instruction tuning, 2023 c.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Visual instruction tuning, 2023 c

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:43.093718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:32:43.093718Z digest=sha256:b4775d4379f1e12e77010d234e25e1e13e844f202cafa60d9c361b58b7a01fa4

Observation 01222fc1-4d5b-4862-9035-864045560d18 · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge, January 2024 a.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Llava-next: Improved reasoning, ocr, and world knowledge, January 2024 a

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:43.099558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:32:43.099558Z digest=sha256:05ef41bf3f80b064755fa47ea7ae5092f62d78759657021c7e78e130ee6e2180

Observation 0c38a33c-f9d0-4df9-bac2-6b1044cf1a3e · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? In European conference on computer vision, pp.\ 216--233.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Mmbench: Is your multi-modal model an all-around player? In European conference on computer vision, pp.\ 216--233

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:43.106880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:32:43.106880Z digest=sha256:d818c45c5cb59816ee75c768e6d8f7dead82972116d5de6fd1b9b2a9295bbb9c

Observation 7d0d69f5-1320-4526-b7e3-7e7e7976f46a · outbound

This paper cites Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:43.116167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:32:43.116167Z digest=sha256:5d8e1afa3eafee0b5fdd94a4614815049715b68cd86b81bfa68d254eec8dcdd6

Observation e5ce209a-9bee-4fc2-b001-34eefa640d60 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:43.123282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:32:43.123282Z digest=sha256:2caef1ee5b04eed05efb84621277535687a3be57bc95f0fa6864651ab249d37a

Observation 305e7673-8562-4b24-9930-1ccd5454ec3b · outbound

This paper cites Ok-vqa: A visual question answering benchmark requiring external knowledge.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Ok-vqa: A visual question answering benchmark requiring external knowledge

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:43.129520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:32:43.129520Z digest=sha256:857bd44c789a872c3cea140f2695f31e0e1bcf0d1fd54c3b74b6ebf9c15e0ef6

Observation d87dc009-474d-4dce-8a68-80272dfd1454 · outbound

This paper cites Gpt-4v(ision) system card, 2023.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Gpt-4v(ision) system card, 2023

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:32:45.434566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-05T16:32:43.134751Z digest=sha256:18643208a5640f02ae86f7e1019a6bdf4744286ae32ad5ed768f04064948636a

Observation 3380c342-0974-4bcc-80b2-1ce4e181596a · outbound

This paper cites GPT-4o System Card.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models GPT-4o System Card

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:43.140039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:32:43.140039Z digest=sha256:4c6d9a0ab7008ddc8c5ca9e3db8d7975bb7098ceb802b3fe12f2b2affabdfadf

Observation c60b406e-824d-4d6c-9ace-be98ff1e8954 · outbound

This paper cites Gpt-4o mini: Advancing cost-efficient intelligence, 2024 b.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Gpt-4o mini: Advancing cost-efficient intelligence, 2024 b

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:32:45.411773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-05T16:32:43.146683Z digest=sha256:54d8cdc1fd38c66d36d7711863ef84e92db63d1d55cba9902b1ac0532996bb0b

Observation bd1b13ce-edee-4e6a-a115-d94fc2f79a8a · outbound

This paper cites Openai o1 system card, December 2024 c.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Openai o1 system card, December 2024 c

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:32:45.388851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-05T16:32:43.152920Z digest=sha256:163f2f27ef2ed5cebd792a23b4a34da44732dce6a3f0dc323fec35ebbd870f08

Observation e9fcfc15-f111-494e-bd35-a5d55318156d · outbound

This paper cites Gpt-4.5 system card, 2025 a.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Gpt-4.5 system card, 2025 a

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:32:45.366029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-05T16:32:43.161131Z digest=sha256:ca12b38fa0695c3075bea657264374a9febb02d5120f26dee5c9023213ca6d5c

Observation 5d283bba-1925-4d51-9066-9845a0fb8af7 · outbound

This paper cites Gpt-5 system card, 2025 b.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Gpt-5 system card, 2025 b

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:32:45.338683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-05T16:32:43.167373Z digest=sha256:00e5b8f4ed82058709587c152fa7ff18b4233f6cc929cf3792c7a4f4d5bbe07d

Observation dff49c99-74e0-4a3d-9c22-846f2ce032c1 · outbound

This paper cites Openai o3-mini system card, February 2025 c.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Openai o3-mini system card, February 2025 c

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:32:45.322244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-05T16:32:43.174490Z digest=sha256:48efbad2510f4b436936a480371b9ab6d53f999907b4cec84d6c58d31a64e9d6

Observation 59195a71-c2e5-4545-9cc0-44a80e816065 · outbound

This paper cites Openai o3 and o4-mini system card, 2025 d.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Openai o3 and o4-mini system card, 2025 d

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:32:45.301339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-05T16:32:43.180287Z digest=sha256:381a144488613311505058c93f6f2e27abd5ab1ad154bd2f718450aecb01ef8c

Observation aecaf710-6ce3-4431-9d76-183a18a5bf07 · outbound

This paper cites Cross-modal Attention Congruence Regularization for Vision-Language Relation Alignment.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Cross-modal Attention Congruence Regularization for Vision-Language Relation Alignment

Reference 74

Resolution
verified exact
local_arxiv, observed 2026-08-05T16:32:44.173425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-05T16:32:43.185194Z digest=sha256:179d3479b856cde9ec71be2a7114b88b863daa894c3f04fc66169a8f5c31be52

Observation 1102cc23-3b0f-46de-8988-d3a531a0f46d · outbound

This paper cites Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs?.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs?

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:43.191067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:32:43.191067Z digest=sha256:4c7c840751d1b5c0e8e320af04226129b5d528417eda5802d8aa4d171def03f5

Observation 1e58ada7-be1e-4694-8a90-9b7dc3640228 · outbound

This paper cites Balanced multimodal learning via on-the-fly gradient modulation.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Balanced multimodal learning via on-the-fly gradient modulation

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:32:45.274551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-05T16:32:43.197495Z digest=sha256:504f6bf2c8b5f243d7860fe1d75dd149b8231b33bf3905c70e0df6fb6964a40b

Observation c86ab290-9ad4-4a06-9a8a-5cfe6ac0e7dd · outbound

This paper cites Rdkit: Open-source cheminformatics, 2025.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Rdkit: Open-source cheminformatics, 2025

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:32:45.253283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-05T16:32:43.204721Z digest=sha256:2927958fb39875ab64c296cb3fb4ff4c420bb96325e82db57fb39d543c6f9f01

Observation 030b90b6-00d8-4128-8338-84030711a36f · outbound

This paper cites The music encoding initiative (mei).

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models The music encoding initiative (mei)

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:32:45.227880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-05T16:32:43.213015Z digest=sha256:a2a8429965d77f0a0277a953374b1e022d2fd1c59c50ac96520eac05fc140a65

Observation 717576f2-79b3-4890-8ff8-b1ad6d093217 · outbound

This paper cites Music abc notation with music theory dataset, 2025.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Music abc notation with music theory dataset, 2025

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:32:45.207573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-05T16:32:43.217878Z digest=sha256:67c57645e92dcadda2cd13edd928ff802156c186f6cf74cbdd016218999a89eb

Observation 1ead4494-65b0-47ce-81c4-edce3d41d116 · outbound

This paper cites LXMERT: Learning Cross-Modality Encoder Representations from Transformers.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models LXMERT: Learning Cross-Modality Encoder Representations from Transformers

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:43.223845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:32:43.223845Z digest=sha256:4eb04615ea4418a73363387f1c9f3a1e2d9bae7432563cbefc3e405412b9cce4

Observation 7bfb62e6-aa72-44b2-ab6d-5f3d298cc645 · outbound

This paper cites NOTA: Multimodal Music Notation Understanding for Visual Large Language Model.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models NOTA: Multimodal Music Notation Understanding for Visual Large Language Model

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:43.228974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:32:43.228974Z digest=sha256:8c1719d17d4c9d3253ba4710d63ba69faa17b16f450e3cdde3cb3d2af26a3114

Observation a33936c0-6748-4474-a156-d578b814e4cb · outbound

This paper cites Cambrian-1: A fully open, vision-centric exploration of multimodal llms.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Cambrian-1: A fully open, vision-centric exploration of multimodal llms

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:32:45.183943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-05T16:32:43.240271Z digest=sha256:1c68eb4253517bbe15a628bffe22c98e646205cc7f7440d409e7ba39583bf866

Observation 4f4a7637-c902-492c-8510-17f51cdf7608 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models LLaMA: Open and Efficient Foundation Language Models

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:43.245876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:32:43.245876Z digest=sha256:5164a35d32c414781b260d09375c492c1e59aa3786661583ccd3463d065ae1df

Observation d5e8a843-d464-4c35-a635-ef36e6506200 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:43.252491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:32:43.252491Z digest=sha256:6f293d1d0c0549b4f4f2d7d2c9c0f6327f5c5594704af83a11be01a809b06b38

Observation 9db96dbe-cb7c-44e1-816d-70403c041b18 · outbound

This paper cites Towards generalist biomedical ai.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Towards generalist biomedical ai

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:43.262505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:32:43.262505Z digest=sha256:2f06aa9d4069bb3b8258c7977a46484c3187781a60da0fe048f5131363964cd2

Observation bcd0eec3-4e6b-426a-995d-ffb50898b36d · outbound

This paper cites Visualizing data using t-sne.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Visualizing data using t-sne

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:43.268755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:32:43.268755Z digest=sha256:da94607bf68dbb689002287cf99a9c47e0c558e7bd25bd4438724481d2b4726b

Observation c6325104-99ae-4d35-ac8a-7ef7ed205511 · outbound

This paper cites LIMCA: LLM for Automating Analog In-Memory Computing Architecture Design Exploration.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models LIMCA: LLM for Automating Analog In-Memory Computing Architecture Design Exploration

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:43.274487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:32:43.274487Z digest=sha256:b1efb24d485f561e7b8b963bca65eb019e4be0d9df31fdd76da371b920ce4970

Observation 7364adca-5cf4-40f5-bb5d-e5596374413b · outbound

This paper cites Abc notation standard, 2004.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Abc notation standard, 2004

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:32:45.120670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-05T16:32:43.280710Z digest=sha256:4eb7b4c72ca6ba812ed932a53e3fb39d8fc9d1366ab281547176aed21d6c3435

Observation 9997b4c6-3a97-4216-b471-9464fb84e6e2 · outbound

This paper cites Is a picture worth a thousand words? delving into spatial reasoning for vision language models.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Is a picture worth a thousand words? delving into spatial reasoning for vision language models

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:32:45.096888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-05T16:32:43.285992Z digest=sha256:bbe19e3b4f05eb605c8fcd04652b407f1659bc6da60fdaab33d95eb6eda44eba

Observation 62b0c2fd-bb29-43ab-998c-45fd3d7599f8 · outbound

This paper cites Multilingual E5 Text Embeddings: A Technical Report.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Multilingual E5 Text Embeddings: A Technical Report

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:43.290623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:32:43.290623Z digest=sha256:0526cf046389d02fd0b8bd37c07b74e7a7e8de76dcdc078e5561f42c36d812ac

Observation efab14d9-2a73-482c-a6b2-3b0a7ed6279c · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:43.296670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:32:43.296670Z digest=sha256:e4c5c2e57d621b462738779bc6292587a339a8821a0e4487af78c04739191c85

Observation 2f7a46f2-a96e-4f35-8c92-1046bfe948a6 · outbound

This paper cites Smiles, a chemical language and information system.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Smiles, a chemical language and information system

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:32:45.076071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-05T16:32:43.303240Z digest=sha256:f55d9a8dd011492a3f7a95eac9e5e7545e205cdc351de063d1f64dcfacdf3c09

Observation 388eac7b-f4f0-4743-b579-35c03b1ed15b · outbound

This paper cites Tunesformer: Forming irish tunes with control codes by bar patching.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Tunesformer: Forming irish tunes with control codes by bar patching

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:32:45.053819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-05T16:32:43.308302Z digest=sha256:c54e4a0503aa887d2b9ccbe20cda0a5754baa99801ebf23a22837e2d1dd494ca

Observation 4a82fe10-49e3-4c20-8230-99660a8f39ed · outbound

This paper cites SciMMIR: Benchmarking Scientific Multi-modal Information Retrieval.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models SciMMIR: Benchmarking Scientific Multi-modal Information Retrieval

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:43.314453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:32:43.314453Z digest=sha256:9b4728cf6f8307213863552d2e8b48c9d3a4ada919a3b3aa2522d8ea85711932

Observation fcd16cd4-e99e-429c-be21-0192a690c04b · outbound

This paper cites Qwen2.5-Omni Technical Report.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Qwen2.5-Omni Technical Report

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:43.321030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:32:43.321030Z digest=sha256:3175aba4476a700f8878fc61981bedf68f8d60753714ceba999085153ee02fe7

Observation f7631358-7355-417e-b85b-f59ba3d580eb · outbound

This paper cites Lvlm-ehub: A comprehensive evaluation benchmark for large vision-language models.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Lvlm-ehub: A comprehensive evaluation benchmark for large vision-language models

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:32:45.035191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-05T16:32:43.328301Z digest=sha256:f99ddd9b3e6b3fc3fbde6f357f2661e40047fb4c65397418b3cb104454224fb5

Observation 922727b0-2cc7-4087-acb8-f82ec6c17722 · outbound

This paper cites Qwen2 Technical Report.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Qwen2 Technical Report

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:43.334674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:32:43.334674Z digest=sha256:0af04128d51bd997e4d548b58dd9b5f74538522872efb99ba67312eda77f5402

Observation 8f4aed58-150c-4fef-8a03-127c4dbe76ab · outbound

This paper cites Qwen2.5 Technical Report.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Qwen2.5 Technical Report

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:43.341074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:32:43.341074Z digest=sha256:9de1c5d8cdb4735afb19f0c9ec0c340276131484a0588ea9bafe60582f6cb655

Observation 5a76350f-abfb-43f5-87c0-66dfb21b3e17 · outbound

This paper cites Report Cards: Qualitative Evaluation of Language Models Using Natural Language Summaries.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Report Cards: Qualitative Evaluation of Language Models Using Natural Language Summaries

Reference 99

Resolution
verified exact
local_arxiv, observed 2026-08-05T16:32:43.767477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-05T16:32:43.345892Z digest=sha256:5f02dd5cde17b911ee7a668c07f9b5b63bbe0e7bb4c8c078c84e01ac33f0bb2f

Observation c13d70d0-f14c-4ed2-a0d3-a8795c32a985 · outbound

This paper cites mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:32:45.012584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-05T16:32:43.351429Z digest=sha256:184f2f4fc719e4e686e634c4b925cf0eeca3cc570d1950d4c67d68c12a10c7af

Pith citing papers

Observation 351a9023-0a61-4de3-8da0-72858a7b8fe0 · inbound

TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs cites this paper.

TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T00:55:26.251613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:55:26.251613Z digest=sha256:3cd6731db324b92a9882f1698d064184e2ec34acc8bae64e1ab5f02b60077a9f