Pith. sign in

Paper Citation Record · LEDGER

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs?

As of 18 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 2 inbound Pith citation observations for arXiv:2501.02669.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.02669 v2

Coverage vector

measured 32 of 32 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T22:13:01.238326Z

measured 34 of 34 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T16:32:43.191067Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T08:53:16.341543Z

Reference resolution

32 of 32 outbound references displayed

  • verified exact0
  • verified fuzzy14
  • unresolved16
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d9b929af-00a5-4a12-a07f-ffbbe87e0c48 · outbound

This paper cites shape type), the model first needs to correctly enumerate the attribute values (e.g.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? shape type), the model first needs to correctly enumerate the attribute values (e.g

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:13:03.945341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T22:13:00.590633Z digest=sha256:9f68b3fda9a20475ff53491a32f4fab413bd657c40caef27ac9c47e26f26724c

Observation 4738033a-e60a-4a35-956d-6ed14f3af148 · outbound

This paper cites an unresolved cited work.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:13:04.316329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T22:13:00.491671Z digest=sha256:59cb234127a263182acd073b98d7992716b46483ad545de755d3ebddb3beb95f

Observation 091f7cb5-6e71-487f-9f52-4ced86728735 · outbound

This paper cites backtracking.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? backtracking

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:13:04.185468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T22:13:00.514749Z digest=sha256:714fa52fc479eadf974bd8c5d6dfc584c8b1da1252c1381328d93701952208fd

Observation 8e50b250-444e-4e1c-bf56-e893c2fd84f8 · outbound

This paper cites • To reason about the query: the model needs to correctly enumerate the attribute values for each image in the query similarly.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? • To reason about the query: the model needs to correctly enumerate the attribute values for each image in the query similarly

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:13:03.583444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T22:13:00.715938Z digest=sha256:bfa18823cc3764f0c40a1c1d265744ec0e2b69a9c6666c06f0d2305a53d61e00

Observation 6c348bf4-4c2b-4212-a0e3-77dd9ac44850 · outbound

This paper cites Singh, A., Natarajan, V ., Shah, M., Jiang, Y ., Chen, X., Batra, D., Parikh, D., and Rohrbach, M.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? Singh, A., Natarajan, V ., Shah, M., Jiang, Y ., Chen, X., Batra, D., Parikh, D., and Rohrbach, M

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:13:04.534759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T22:13:00.344750Z digest=sha256:bb7b64d50bb736e3eff25f4dc9ad08fd77f7ec24437b59ff72759082285ec1f1

Observation c049058d-af3d-4975-a6dd-142ba2e4f579 · outbound

This paper cites Sun, Z., Yu, L., Shen, Y ., Liu, W., Yang, Y ., Welleck, S., and Gan, C.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? Sun, Z., Yu, L., Shen, Y ., Liu, W., Yang, Y ., Welleck, S., and Gan, C

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:13:04.485464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T22:13:00.394167Z digest=sha256:86d95c26440a57e5327fe7fed2db6bde359d1167029de95745bd003bbe5c55d8

Observation b1f6d2a1-d1f7-456c-823d-73939131dfb8 · outbound

This paper cites Are Large-Language Models Graph Algorithmic Reasoners?.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? Are Large-Language Models Graph Algorithmic Reasoners?

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T22:13:00.399401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:13:00.399401Z digest=sha256:676290ff64bb99c6f84771708973c0c599084835d5f3eeadaaeeb2c5d8238f51

Observation e3404e2f-cd44-4751-b728-0c2a890cb8ae · outbound

This paper cites MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T22:13:00.424752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:13:00.424752Z digest=sha256:8b103a39a518a404e7658eed1193e8dad331a85bc21a3f65d14be0019f76be9b

Observation 978db19d-91f2-44da-93e8-ee725fb2899f · outbound

This paper cites single-hop.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? single-hop

Reference 10

Resolution
malformed identifier
raw_fallback, observed 2026-08-10T22:13:04.444754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T22:13:00.454746Z digest=sha256:30fc1e481dd8664bc58e356b38eb1277c37f96a12a244642b3ce569be7048587

Observation 08ac75a3-4241-45ca-a8be-d40be35b0bad · outbound

This paper cites an unresolved cited work.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:13:03.794769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T22:13:00.626799Z digest=sha256:7268a5d64ea5b44cfa3d9df8bbb05f296b5c5ecc8c9a3424fc4e002d1db6c2fd

Observation 06843895-1faa-4f65-bca3-c48f112ca72d · outbound

This paper cites an unresolved cited work.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:13:03.679614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T22:13:00.674277Z digest=sha256:df06a7dbf16612b51353fca8c0a51ffc40549d6d547fe37eed6422d8b1e9856f

Observation 1eb820e9-c6d1-4b40-9924-28cc28c105b4 · outbound

This paper cites (line type , XOR), the model needs to identify the correct values of the attribute domain d for each option image and the correct relationr.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? (line type , XOR), the model needs to identify the correct values of the attribute domain d for each option image and the correct relationr

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:13:03.474745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T22:13:00.754749Z digest=sha256:a8f628b2cc8c57877528e01f5c1bf2f6191fe99adc011012259249d9a73e1b8d

Observation 8b419063-6839-4715-8156-031645c3f018 · outbound

This paper cites 46 Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? Table 14.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? 46 Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? Table 14

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:13:03.364749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T22:13:00.784808Z digest=sha256:fe3721bf3a013da441e79a0cf332d4fe04ebfcf683dbe5410092011c0a8613c8

Observation 7ae0808f-cb8f-4690-a4f4-9fc0568f3f2c · outbound

This paper cites an unresolved cited work.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:13:03.263094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T22:13:00.834746Z digest=sha256:9840137137f951576a1ec724d001c5ad1aa6a635306f839e63e568ac31da111e

Observation ff329e0e-2a4a-4a88-bc39-a405a7262fdb · outbound

This paper cites Answer: 95 November December Figure 32.ASIMPLEexample fromTable Readout.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? Answer: 95 November December Figure 32.ASIMPLEexample fromTable Readout

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:13:03.182008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T22:13:00.874880Z digest=sha256:0fa09d943084647ae1e61f8470521b20059b262443b2466902f3b7a409cb074b

Observation 6309a005-e4a2-45e2-aa3e-7b52bb1c3aaa · outbound

This paper cites an unresolved cited work.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:13:03.094820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T22:13:00.920852Z digest=sha256:78adfddaec6bf922b3a0201c6e1197c2c658a38928d829f9562009e4f8c53348

Observation 68bb3852-c4eb-48c8-ad6d-d2dfe5c2ff34 · outbound

This paper cites Answer: 233 Figure 33.AHARDexample fromTable Readout.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? Answer: 233 Figure 33.AHARDexample fromTable Readout

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:13:03.025587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T22:13:00.961279Z digest=sha256:d954b307e3d07d854d12a77344916fc502c65ba44a03f9d93837f34572275de7

Observation 353eb138-7832-49d5-9b38-081d30de6fb4 · outbound

This paper cites The grid is filled up with objects, which you will be asked to recognize and collect, and obstacles, which you should avoid.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? The grid is filled up with objects, which you will be asked to recognize and collect, and obstacles, which you should avoid

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:13:02.904748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T22:13:00.994759Z digest=sha256:85c7370b42b9fe6ec7118a6d8688fee847e8653a0e74c3edcbf8ee9a6f3e5317

Observation bb8675ab-8c9c-4b48-91ce-fa88e39e2ca1 · outbound

This paper cites an unresolved cited work.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:13:02.788039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T22:13:01.044070Z digest=sha256:7d9de209397b06059d2524b0b6b6e815bc1ed53be361e6ec30a561de516d2577

Observation f517d880-22c6-4f55-8aad-13066805e3a0 · outbound

This paper cites The grid is filled up with objects, which you will be asked to recognize and collect, and obstacles, which you should avoid.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? The grid is filled up with objects, which you will be asked to recognize and collect, and obstacles, which you should avoid

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:13:02.663678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T22:13:01.059180Z digest=sha256:5abd66c504da3514fed5294f794fe776e2592cda87191c6666c6e4ecc5df84e2

Observation 93957649-5c7b-4775-8d6c-c13322deba03 · outbound

This paper cites an unresolved cited work.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:13:02.569937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T22:13:01.063216Z digest=sha256:e0fa662917e24eccfdd89b6d51b4e607978ca8a43867ed1bdfbe82beb35f4464

Observation 244b66e0-b917-4c17-9f2c-6a9a29e89968 · outbound

This paper cites 51 Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? The image shows a a puzzle in a 3 by 3 grid followed by 4 options.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? 51 Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? The image shows a a puzzle in a 3 by 3 grid followed by 4 options

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:13:02.435490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T22:13:01.086083Z digest=sha256:ecbf18b73eb8c80321d79ba53d2efdbe7cabec3625c1ec24feda08dea7514b28

Observation b3021570-a7b2-4bce-8e26-1ffd20f0dd7f · outbound

This paper cites … position: Image 1: (1, 0), (0, 2) Image 2: (0, 2), (1, 1) Image 3: (0, 2) This suggests the AND relation.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? … position: Image 1: (1, 0), (0, 2) Image 2: (0, 2), (1, 1) Image 3: (0, 2) This suggests the AND relation

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:13:02.378205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T22:13:01.125099Z digest=sha256:45d81de511fa163770a8e5eb555347c9036b7f4987a32a4f847155ba97672b13

Observation ff24848f-4171-4f07-85ec-70295e6d4e8f · outbound

This paper cites an unresolved cited work.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:13:02.234271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T22:13:01.146836Z digest=sha256:dbf17a2eff47e2221f94af193a49a191cd0cf14d494fd68401bcebb9b422a3e9

Observation c6d5e29a-2eca-46f5-a933-a7fba6130d87 · outbound

This paper cites color: Image 1: 189, 135 Image 2: 189 Image 3: 189 This suggests the AND relation.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? color: Image 1: 189, 135 Image 2: 189 Image 3: 189 This suggests the AND relation

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:13:02.124845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T22:13:01.192600Z digest=sha256:b6af17e9535d51c07a8d259a894b95d052380c2e8dbd7bcea6a6206d2b74080b

Observation 4fd0a3a0-7f46-41c1-b50e-e97941bb5a47 · outbound

This paper cites an unresolved cited work.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:13:02.015722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T22:13:01.238326Z digest=sha256:f65094c0f8414763298db8c281d94935f5a6ab85d4dc5f405473ee396f948207

Observation bde4c299-e96c-446b-bf43-ee2ffed36c6d · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-10T22:13:00.234747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:13:00.234747Z digest=sha256:7afc2ed237a1e0f7b3cf98cc4c3e57b6e8071026c6bb7b0c0e08fe84151c3736

Observation a8a3da1c-c5fe-4d93-8c2b-3099b57c98ba · outbound

This paper cites Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision

Reference 576

Resolution
unresolved
no resolver link, observed 2026-08-10T22:13:00.216677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:13:00.216677Z digest=sha256:fb948bbaf17ee7cb1008f322a7ff8498b47ea25f11eafaf5418a82b242b2f92c

Observation 5bc9d72c-a524-41a7-89f5-4345f31dbb77 · outbound

This paper cites Convert,.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? Convert,

Reference 2015

Resolution
malformed identifier
raw_fallback, observed 2026-08-10T22:13:04.074754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T22:13:00.544835Z digest=sha256:1afdb6de3802bd81e2cd7dceb5117cb4bd17a7ef8605ad56f7c373d947d418eb

Observation 9a97e119-ad7e-4962-bab7-09c74feee0eb · outbound

This paper cites doi: 10.18653/v1/2023.acl-short.43.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? doi: 10.18653/v1/2023.acl-short.43

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-10T22:13:00.299514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:13:00.299514Z digest=sha256:fe0715b33aab4a5d5164513ba8a52c85596a7ee83f02c686380ab70daa4f6eca

Observation 42a01c31-c3bc-444b-8b05-59224c67ef9c · outbound

This paper cites On Pre-training of Multimodal Language Models Customized for Chart Understanding.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? On Pre-training of Multimodal Language Models Customized for Chart Understanding

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-10T22:13:00.264753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:13:00.264753Z digest=sha256:c066c11e7b25136643d79f5e33e38080c33dd37ab196582bc57bf548b119cb3a

Observation db99015a-11da-4001-9936-ae37af6886cb · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-10T22:13:00.271469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:13:00.271469Z digest=sha256:9cdae9707a4988d3c801d140c041784345298424f6688b1201773485b8d63063

Pith citing papers

Observation 1102cc23-3b0f-46de-8988-d3a531a0f46d · inbound

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models cites this paper.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs?

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:43.191067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:32:43.191067Z digest=sha256:ce5b06d627ae029d7196582297e99f77ccbc4f1c38ceb45ee094317ded302b40

Observation 5f397b56-197c-44e0-9ad9-cc6c9852c6d3 · inbound

GDSD: Reinforcement Learning as Guided Denoiser Self-Distillation for Diffusion Language Models cites this paper.

GDSD: Reinforcement Learning as Guided Denoiser Self-Distillation for Diffusion Language Models Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs?

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:53:16.343122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-29T08:44:53.969301Z digest=sha256:490c50d7049f5e2312d72d9493af96f62eda528e5347525eeb586b39f5c5a506