Pith. sign in

Paper Citation Record · LEDGER

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs?

As of 18 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 2 inbound Pith citation observations for arXiv:2501.02669.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.02669 v2

Coverage vector

measured 32 of 32 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T22:13:01.238326Z

measured 34 of 34 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T16:32:43.191067Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T08:53:16.341543Z

Reference resolution

32 of 32 outbound references displayed

  • verified exact0
  • verified fuzzy14
  • unresolved16
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d9b929af-00a5-4a12-a07f-ffbbe87e0c48 · outbound

This paper cites shape type), the model first needs to correctly enumerate the attribute values (e.g.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? shape type), the model first needs to correctly enumerate the attribute values (e.g

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:13:03.945341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T22:13:00.590633Z digest=sha256:51e43ab2ebc2474e5c8e6a37cd9e594c51cccd9e7eff5fb4fd012265027fd084

Observation 4738033a-e60a-4a35-956d-6ed14f3af148 · outbound

This paper cites an unresolved cited work.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:13:04.316329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T22:13:00.491671Z digest=sha256:bd4c1a71a0b71ab7ba7d9c8e102c2853fe2a1455f130997b14a85dd5549cced4

Observation 091f7cb5-6e71-487f-9f52-4ced86728735 · outbound

This paper cites backtracking.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? backtracking

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:13:04.185468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T22:13:00.514749Z digest=sha256:5934e9dca63d60bf1d62bb46ae2c14c444736339f839d1a7f5d7bc490d4feb16

Observation 8e50b250-444e-4e1c-bf56-e893c2fd84f8 · outbound

This paper cites • To reason about the query: the model needs to correctly enumerate the attribute values for each image in the query similarly.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? • To reason about the query: the model needs to correctly enumerate the attribute values for each image in the query similarly

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:13:03.583444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T22:13:00.715938Z digest=sha256:b0451d1f550f6b56e96895d57630a44305c7c0f98e691cad8d4ccab14e2a8f07

Observation 6c348bf4-4c2b-4212-a0e3-77dd9ac44850 · outbound

This paper cites Singh, A., Natarajan, V ., Shah, M., Jiang, Y ., Chen, X., Batra, D., Parikh, D., and Rohrbach, M.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? Singh, A., Natarajan, V ., Shah, M., Jiang, Y ., Chen, X., Batra, D., Parikh, D., and Rohrbach, M

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:13:04.534759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T22:13:00.344750Z digest=sha256:ca78d57057613465a8219ca17c0be932c3b91f2c10a2700ba7ff7de5d108a5a5

Observation c049058d-af3d-4975-a6dd-142ba2e4f579 · outbound

This paper cites Sun, Z., Yu, L., Shen, Y ., Liu, W., Yang, Y ., Welleck, S., and Gan, C.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? Sun, Z., Yu, L., Shen, Y ., Liu, W., Yang, Y ., Welleck, S., and Gan, C

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:13:04.485464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T22:13:00.394167Z digest=sha256:92fbd0ce8c814164b1830f86a8bd32d575ea6f9ad964e17c51482cd3c75119fc

Observation b1f6d2a1-d1f7-456c-823d-73939131dfb8 · outbound

This paper cites Are Large-Language Models Graph Algorithmic Reasoners?.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? Are Large-Language Models Graph Algorithmic Reasoners?

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T22:13:00.399401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:13:00.399401Z digest=sha256:b89a7a578805636fe9f46b85815c782568a100deddba854bc7f20a89e0413757

Observation e3404e2f-cd44-4751-b728-0c2a890cb8ae · outbound

This paper cites MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T22:13:00.424752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:13:00.424752Z digest=sha256:967be0b8daabe5357b5a85e069902da24ba43f1dee1ef9e63246606ce501b53d

Observation 978db19d-91f2-44da-93e8-ee725fb2899f · outbound

This paper cites single-hop.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? single-hop

Reference 10

Resolution
malformed identifier
raw_fallback, observed 2026-08-10T22:13:04.444754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T22:13:00.454746Z digest=sha256:2e8e1ddf75f8940a76a1a9829e35c37dc5a67d457dda3f62469cd3cc33d91264

Observation 08ac75a3-4241-45ca-a8be-d40be35b0bad · outbound

This paper cites an unresolved cited work.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:13:03.794769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T22:13:00.626799Z digest=sha256:9bbb8519c293ba1f0fa4931958e94627a2e36b9d8301029b459641781579fd24

Observation 06843895-1faa-4f65-bca3-c48f112ca72d · outbound

This paper cites an unresolved cited work.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:13:03.679614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T22:13:00.674277Z digest=sha256:bc95c466caf14ddde04c5a070d6bf9c1480bdefb4617431e6a7be0662ba9fea1

Observation 1eb820e9-c6d1-4b40-9924-28cc28c105b4 · outbound

This paper cites (line type , XOR), the model needs to identify the correct values of the attribute domain d for each option image and the correct relationr.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? (line type , XOR), the model needs to identify the correct values of the attribute domain d for each option image and the correct relationr

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:13:03.474745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T22:13:00.754749Z digest=sha256:6633cae080d001dd37e277928dcc049c342f5bec37dcf9f7ae2b14c6e17b1e90

Observation 8b419063-6839-4715-8156-031645c3f018 · outbound

This paper cites 46 Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? Table 14.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? 46 Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? Table 14

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:13:03.364749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T22:13:00.784808Z digest=sha256:00c5ae766f91418dfde22aec21d3b59d8acb465fa05b303c4d3e074fdc0171fd

Observation 7ae0808f-cb8f-4690-a4f4-9fc0568f3f2c · outbound

This paper cites an unresolved cited work.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:13:03.263094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T22:13:00.834746Z digest=sha256:cfe2ab3da2d629779a4f12d249251a29b79d9214350a548ef04570cee8d8b91e

Observation ff329e0e-2a4a-4a88-bc39-a405a7262fdb · outbound

This paper cites Answer: 95 November December Figure 32.ASIMPLEexample fromTable Readout.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? Answer: 95 November December Figure 32.ASIMPLEexample fromTable Readout

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:13:03.182008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T22:13:00.874880Z digest=sha256:f3445a0398040770a352ad8826efaf03f71b3baed9a631a6e70bf559c958abb0

Observation 6309a005-e4a2-45e2-aa3e-7b52bb1c3aaa · outbound

This paper cites an unresolved cited work.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:13:03.094820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T22:13:00.920852Z digest=sha256:b7a26ccfc2bdc25480958df4ccf03d73f8ddfbe73f43e031c137b2ca3d09af71

Observation 68bb3852-c4eb-48c8-ad6d-d2dfe5c2ff34 · outbound

This paper cites Answer: 233 Figure 33.AHARDexample fromTable Readout.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? Answer: 233 Figure 33.AHARDexample fromTable Readout

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:13:03.025587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T22:13:00.961279Z digest=sha256:6806c603e29781c0df9c70457ad57a9fcc55b36cf5702822c809ff387712790d

Observation 353eb138-7832-49d5-9b38-081d30de6fb4 · outbound

This paper cites The grid is filled up with objects, which you will be asked to recognize and collect, and obstacles, which you should avoid.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? The grid is filled up with objects, which you will be asked to recognize and collect, and obstacles, which you should avoid

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:13:02.904748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T22:13:00.994759Z digest=sha256:fbf288430a582f9946228c0fbfda338a7e8ba79cf3108a65d7a73f8281468778

Observation bb8675ab-8c9c-4b48-91ce-fa88e39e2ca1 · outbound

This paper cites an unresolved cited work.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:13:02.788039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T22:13:01.044070Z digest=sha256:0ca79efe59e36e6f4eea71a6ceecc9c9e68ffa6f29a2f210bfad5d179e301938

Observation f517d880-22c6-4f55-8aad-13066805e3a0 · outbound

This paper cites The grid is filled up with objects, which you will be asked to recognize and collect, and obstacles, which you should avoid.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? The grid is filled up with objects, which you will be asked to recognize and collect, and obstacles, which you should avoid

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:13:02.663678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T22:13:01.059180Z digest=sha256:57472052a07690c9c0ead1f47f5467bbf2782e42e4a5d76df322cda2419c2c62

Observation 93957649-5c7b-4775-8d6c-c13322deba03 · outbound

This paper cites an unresolved cited work.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:13:02.569937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T22:13:01.063216Z digest=sha256:7406f92f4aaa61f70af34e271e0b390df5f3f0c2bce1c452dca7512e803cc23e

Observation 244b66e0-b917-4c17-9f2c-6a9a29e89968 · outbound

This paper cites 51 Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? The image shows a a puzzle in a 3 by 3 grid followed by 4 options.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? 51 Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? The image shows a a puzzle in a 3 by 3 grid followed by 4 options

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:13:02.435490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T22:13:01.086083Z digest=sha256:90f71078e205caef217c4bf8154a1f435c94c1fbccd7e76e05b7c3a31d52150b

Observation b3021570-a7b2-4bce-8e26-1ffd20f0dd7f · outbound

This paper cites … position: Image 1: (1, 0), (0, 2) Image 2: (0, 2), (1, 1) Image 3: (0, 2) This suggests the AND relation.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? … position: Image 1: (1, 0), (0, 2) Image 2: (0, 2), (1, 1) Image 3: (0, 2) This suggests the AND relation

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:13:02.378205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T22:13:01.125099Z digest=sha256:55a5556ab9f5851f804b540910ed93994ae587cd9251c1684f50f7d033e64460

Observation ff24848f-4171-4f07-85ec-70295e6d4e8f · outbound

This paper cites an unresolved cited work.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:13:02.234271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T22:13:01.146836Z digest=sha256:04b48d91d0186fd79d3591fdf340b053f60fa7f26a5bf20ed51984a72088802b

Observation c6d5e29a-2eca-46f5-a933-a7fba6130d87 · outbound

This paper cites color: Image 1: 189, 135 Image 2: 189 Image 3: 189 This suggests the AND relation.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? color: Image 1: 189, 135 Image 2: 189 Image 3: 189 This suggests the AND relation

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:13:02.124845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T22:13:01.192600Z digest=sha256:40b53ff0e69c3cc2ae7d12fb9d7195f41741df2e20f172f8f766df5dbda1bd88

Observation 4fd0a3a0-7f46-41c1-b50e-e97941bb5a47 · outbound

This paper cites an unresolved cited work.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:13:02.015722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T22:13:01.238326Z digest=sha256:1cbcb32173dcd0a63ac0dc1bc8f666afe7b872ddd6c48581a237e231efdb9b73

Observation bde4c299-e96c-446b-bf43-ee2ffed36c6d · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-10T22:13:00.234747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:13:00.234747Z digest=sha256:38f4ff998290c6f179c760275f7bda09f3627601403d425dd9da2ef1d99cf225

Observation a8a3da1c-c5fe-4d93-8c2b-3099b57c98ba · outbound

This paper cites Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision

Reference 576

Resolution
unresolved
no resolver link, observed 2026-08-10T22:13:00.216677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:13:00.216677Z digest=sha256:a54bf63fe94bfbdd01bb5e69bda35d3e801e3076b6dca61b0bf1c075885cc649

Observation 5bc9d72c-a524-41a7-89f5-4345f31dbb77 · outbound

This paper cites Convert,.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? Convert,

Reference 2015

Resolution
malformed identifier
raw_fallback, observed 2026-08-10T22:13:04.074754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T22:13:00.544835Z digest=sha256:172d4912f1ac5bdb39ec9221bd8a06a5376e5372cc611e300b3754900ed5b90e

Observation 9a97e119-ad7e-4962-bab7-09c74feee0eb · outbound

This paper cites doi: 10.18653/v1/2023.acl-short.43.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? doi: 10.18653/v1/2023.acl-short.43

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-10T22:13:00.299514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:13:00.299514Z digest=sha256:79ec88fa8805afb943ed2b06c5f1d1e1aa1e68b35d8c91e28d7c711696cacdd5

Observation 42a01c31-c3bc-444b-8b05-59224c67ef9c · outbound

This paper cites On Pre-training of Multimodal Language Models Customized for Chart Understanding.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? On Pre-training of Multimodal Language Models Customized for Chart Understanding

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-10T22:13:00.264753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:13:00.264753Z digest=sha256:b7cf241c0a9e5674ba89c1d40918505cd2df2d79907b41b6904a713f8c432345

Observation db99015a-11da-4001-9936-ae37af6886cb · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs? MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-10T22:13:00.271469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:13:00.271469Z digest=sha256:b9c40daf4f74f232328bf868c7ea2fe972114ed4f2e868056b7cb9bf6f9c4783

Pith citing papers

Observation 1102cc23-3b0f-46de-8988-d3a531a0f46d · inbound

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models cites this paper.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs?

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:43.191067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:32:43.191067Z digest=sha256:4c7c840751d1b5c0e8e320af04226129b5d528417eda5802d8aa4d171def03f5

Observation 5f397b56-197c-44e0-9ad9-cc6c9852c6d3 · inbound

GDSD: Reinforcement Learning as Guided Denoiser Self-Distillation for Diffusion Language Models cites this paper.

GDSD: Reinforcement Learning as Guided Denoiser Self-Distillation for Diffusion Language Models Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs?

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:53:16.343122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-29T08:44:53.969301Z digest=sha256:63fda4fbe1de7aea8618baa4878e9622d94974912098d83b1cf33315844347b4