Pith. sign in

Paper Citation Record · LEDGER

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward

As of 18 August 2026, this Paper Citation Record lists 85 of 85 outbound references and 0 inbound Pith citation observations for arXiv:2504.16727.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.16727 v3

Coverage vector

measured 85 of 85 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:01:40.085249Z

measured 85 of 85 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

85 of 85 outbound references displayed

  • verified exact1
  • verified fuzzy1
  • unresolved83
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 97ec323a-ec9f-430b-8b31-3c15c92376e1 · outbound

This paper cites online" 'onlinestring :=.

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward online" 'onlinestring :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T11:01:39.631336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:01:39.631336Z digest=sha256:e11bdff9ac9d1c5927c5c1cb5864f3506316f86f0e9bc6b582f950298cd45627

Observation cf68b0c9-acf6-4121-abd1-8a57c79aa8e7 · outbound

This paper cites write newline.

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T11:01:39.638039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:01:39.638039Z digest=sha256:1ec579139ac576463d8d6254cbf9f1177f6b148cd71993ff916bc94bebebe737

Observation 8b7442c2-e228-4152-a724-7dbc47cff0d0 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T11:01:39.643026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:01:39.643026Z digest=sha256:7eab717ba49c89570194788a1a8d629e428099efdb28e47d96993598952aca49

Observation da685c7d-6460-4f2b-be84-b125ec9432a2 · outbound

This paper cites Understanding intermediate layers using linear classifier probes.

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward Understanding intermediate layers using linear classifier probes

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T11:01:39.648013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:01:39.648013Z digest=sha256:f1692e70b2cdda4ec7bf36df282a4f82fe1863f4f4c39b294edc7b9dc7902227

Observation b424f949-bb93-4945-be3e-2b4e92f4ed84 · outbound

This paper cites Flamingo: a Visual Language Model for Few-Shot Learning.

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward Flamingo: a Visual Language Model for Few-Shot Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T11:01:39.653003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:01:39.653003Z digest=sha256:6271e9651a8b22e22353498633de7a69af4b1aa8545f4d854bb514d3ac0a90a1

Observation f2fc880b-7d27-4861-9d9e-e4b37fe1c67a · outbound

This paper cites See It from My Perspective: How Language Affects Cultural Bias in Image Understanding.

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward See It from My Perspective: How Language Affects Cultural Bias in Image Understanding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T11:01:39.658149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:01:39.658149Z digest=sha256:119009a16cae3c1854c94a19ca26e968eccc1fad3677debb8941e48d915fdda9

Observation e4248c0b-1f24-43c5-abc2-4a51ae17d25f · outbound

This paper cites an unresolved cited work.

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T11:01:39.664947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:01:39.664947Z digest=sha256:0a5864daf68c770ede9d7adeb8966a8e9f47766996a0295e9c13be190292f3cf

Observation 1980a252-4adb-461b-b17b-3fa046bdadb5 · outbound

This paper cites Qwen Technical Report.

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward Qwen Technical Report

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T11:01:39.670195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:01:39.670195Z digest=sha256:5a41bd0eb4cb64a533ad19a66c9a19e8bf21f4f9d8aeb17a23ae7bb8eb5f0289

Observation 3ec1fcf5-3aee-41f4-acaa-12adbed600eb · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T11:01:39.675722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:01:39.675722Z digest=sha256:413f3b282d1b9465eb6db02b84c3805734292367161ae51f6ea32b0e133e00ee

Observation a9b024f8-a1d7-4ebd-8dab-97c118d63501 · outbound

This paper cites an unresolved cited work.

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T11:01:39.682091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:01:39.682091Z digest=sha256:a9a247bb7507cc6aafbeaea00061970edfc69135674015a114503888a3e03ebb

Observation 1c1b7bb4-ff73-4353-a000-fafd903f8a3f · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T11:01:39.687461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:01:39.687461Z digest=sha256:622d706e684a787fb560a4da1409afef73ab0a92daab35eb3a7e7f4730bf9bd7

Observation ba242164-361b-49a3-96e9-e71874d61941 · outbound

This paper cites GMAI-MMBench: A Comprehensive Multimodal Evaluation Benchmark Towards General Medical AI.

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward GMAI-MMBench: A Comprehensive Multimodal Evaluation Benchmark Towards General Medical AI

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T11:01:39.697849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:01:39.697849Z digest=sha256:76e7a3edfdacd74f98f711b96bc0cca0e0fe693d74c53c5dd695c109da618bd8

Observation 60d1f8f6-8ade-406f-a342-b55155c565cb · outbound

This paper cites an unresolved cited work.

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T11:01:39.703064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:01:39.703064Z digest=sha256:b170efa202ae5aefbc1a59354a6c9b34f31791639c56ed717b440c73ae665537

Observation fca509fc-b2f6-4c0a-86e8-06eb45f042cc · outbound

This paper cites Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling.

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T11:01:39.708780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:01:39.708780Z digest=sha256:d78026d8388837e12b50d39b602fee020c21c52472c58b963420c4a24a1304c3

Observation b1542e83-8b05-4610-995a-f0db96387a9a · outbound

This paper cites RH20T-P: A Primitive-Level Robotic Dataset Towards Composable Generalization Agents.

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward RH20T-P: A Primitive-Level Robotic Dataset Towards Composable Generalization Agents

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T11:01:39.714202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:01:39.714202Z digest=sha256:5392ecfb178a0da3fda08ce837ab0597f98bee229a23cf652d92f82a2a91c9bc

Observation 8c6ec801-4426-42e7-a595-52e89c717985 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T11:01:39.719645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:01:39.719645Z digest=sha256:fbe1dd00016ff33e59f25284fd41de853744244a2d55400d9bf4e69df8fa6520

Observation 693b231c-ab09-4cf2-baad-70fe1ce10b44 · outbound

This paper cites an unresolved cited work.

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T11:01:39.725020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:01:39.725020Z digest=sha256:9e3fe13c5668a2d7dbe7df4631dd2f7a080319dd7c516317950839d85a8a3ae2

Observation 44fe01aa-94a9-4ed2-a7ea-24bb92d2d80a · outbound

This paper cites an unresolved cited work.

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T11:01:39.730778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:01:39.730778Z digest=sha256:79303082f467184043a649fa946afd14eedddac286838f0a3228313f65bce680

Observation 85226e3d-5393-4bd9-bd66-d86989026f72 · outbound

This paper cites Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models.

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T11:01:39.735440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:01:39.735440Z digest=sha256:f977ca5f1012feb8d98296ff0ef1086abfb38cf69b971fad2870ccf73cccefc0

Observation e3cff407-7cbf-46af-b677-2ac089e6a2c3 · outbound

This paper cites an unresolved cited work.

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T11:01:39.740442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:01:39.740442Z digest=sha256:fc136f9cc7a9a9d83ee7bf523528ca50022bfc33287d1f133eee5d859efbf573

Observation aafbde9b-c0ec-454c-9533-bc725b410cf4 · outbound

This paper cites The Llama 3 Herd of Models.

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward The Llama 3 Herd of Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T11:01:39.746252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:01:39.746252Z digest=sha256:22e99bc81f10f8204b51cf11b38ef039cad069ea475b4c03fb5db25d21a95d55

Observation dd2fd0d3-5c13-4eab-b940-e3e09d4d100e · outbound

This paper cites an unresolved cited work.

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-16T11:01:41.911010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T11:01:39.751662Z digest=sha256:038b614f9901fa7c19e12f7a52402aed1e12a313489777b9b06abe0828310817

Observation 8949cbf4-ae19-4478-8f7f-4085d9fa726b · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T11:01:39.759537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:01:39.759537Z digest=sha256:81ca916aa93f162a93fb3ef891cff6530a0a61901dbbee0a857543c12172c734

Observation c96bee23-c5ae-414a-8fc6-9959e594e21c · outbound

This paper cites BLINK: Multimodal Large Language Models Can See but Not Perceive.

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T11:01:39.764632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:01:39.764632Z digest=sha256:27aafa3e8496ec341deb2199bc992de5748b54448fa7900dcf3c0e7acc571346

Observation f01f9e68-b6da-4d45-a0ca-c6e246a1d7a6 · outbound

This paper cites H2OVL-Mississippi Vision Language Models Technical Report.

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward H2OVL-Mississippi Vision Language Models Technical Report

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T11:01:39.771355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:01:39.771355Z digest=sha256:baa76217c70cece16b17c3f42e863fa34defb526cfc4869aaab04cef52985672

Observation 25042b6a-658e-47d1-93fa-a9a7d86658ae · outbound

This paper cites ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools.

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T11:01:39.782666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:01:39.782666Z digest=sha256:e330339eaac1da7f33e4869c90b38c9eeb979f48071602917770b97d34e4f9af

Observation c6f89848-9af7-411b-9f49-3511e5cf0191 · outbound

This paper cites an unresolved cited work.

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T11:01:39.787783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:01:39.787783Z digest=sha256:3425d3c6ed039daa8292c1b5a90922fe4b230d0691fbf1a7752232729e0566ba

Observation eca983f9-67c8-4c1a-9d0f-f9dc3b9d33a7 · outbound

This paper cites Self-Correction is More than Refinement: A Learning Framework for Visual and Language Reasoning Tasks.

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward Self-Correction is More than Refinement: A Learning Framework for Visual and Language Reasoning Tasks

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T11:01:39.792742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:01:39.792742Z digest=sha256:18348cbb2163a2e069d78b77234cc24d3e0568885c8c8a8fd9d5dbd5b3387adc

Observation 54b81bdb-9f80-45fe-87d4-7f5dacbedb96 · outbound

This paper cites MMBoundary: Advancing MLLM Knowledge Boundary Awareness through Reasoning Step Confidence Calibration.

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward MMBoundary: Advancing MLLM Knowledge Boundary Awareness through Reasoning Step Confidence Calibration

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T11:01:39.798306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:01:39.798306Z digest=sha256:7d47d07ea43e8dda3a2efc8c3d4179654ebc221b9ebd1625328d7cf1c39ccd32

Observation 36f9864a-35c8-46e2-a99e-84055ae4c87f · outbound

This paper cites Uncovering Bias in Large Vision-Language Models with Counterfactuals.

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward Uncovering Bias in Large Vision-Language Models with Counterfactuals

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T11:01:39.803579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:01:39.803579Z digest=sha256:9f81212bf189ecf040385195e83468bed4a2bfd602d6641e8171ee0fe80de46a

Observation 5e34acb8-8c12-4868-8def-13f8dd03e60a · outbound

This paper cites OmniMedVQA: A New Large-Scale Comprehensive Evaluation Benchmark for Medical LVLM.

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward OmniMedVQA: A New Large-Scale Comprehensive Evaluation Benchmark for Medical LVLM

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T11:01:39.808655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:01:39.808655Z digest=sha256:82d35167b0a29b7dc76fddd6de0f01c1b534523b3429d68f95149a3628f976f4

Observation b5946d6e-295f-4244-82bc-1b2c16b7be5e · outbound

This paper cites an unresolved cited work.

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T11:01:39.813407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:01:39.813407Z digest=sha256:13a7ea2602e8f7b5ecaa87b6d5633df7b4fe4802c20b46cfd4f8be002965024a

Observation 633f170d-095b-48f3-8475-dc4ec9a8ca24 · outbound

This paper cites an unresolved cited work.

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T11:01:39.817880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:01:39.817880Z digest=sha256:41393e8406588615287885bf978313d7451115bdac321191c36792aa531ecd66

Observation ab161e2b-d8f7-4740-9e9b-6779058a7b92 · outbound

This paper cites Mistral 7B.

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward Mistral 7B

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T11:01:39.822888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:01:39.822888Z digest=sha256:9431cb58fbf000b521d98e1622927437705ef7ce5bbce2a5b54e7cba077e2494

Observation 0eb79cc7-c5c0-4aa1-a4e8-52cf1a68a009 · outbound

This paper cites an unresolved cited work.

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T11:01:39.827635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:01:39.827635Z digest=sha256:50ba769a33723cb89757011da32b1c7878db8d19fd9bee029ea96dfa7d65ee78

Observation 5648ec8d-639b-4df0-bcf0-3c04f82cfe5a · outbound

This paper cites Benchmarking Cognitive Biases in Large Language Models as Evaluators.

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T11:01:39.832116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:01:39.832116Z digest=sha256:1fc94a81e2d636acda74b04bf2dfd328205d5500468fd19884138380ff825af1

Observation 41f79523-1c8a-409c-bfaf-be2694fd13c2 · outbound

This paper cites VHELM: A Holistic Evaluation of Vision Language Models.

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward VHELM: A Holistic Evaluation of Vision Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-16T11:01:39.837833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:01:39.837833Z digest=sha256:122caa49c11280d778256f4f5c96856069d374fb864aa01a7a0a5c3bfa461f00

Observation 04cfaf20-cb91-46c6-84d0-78f252d6d8d5 · outbound

This paper cites The Power of Scale for Parameter-Efficient Prompt Tuning.

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward The Power of Scale for Parameter-Efficient Prompt Tuning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-16T11:01:39.842600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:01:39.842600Z digest=sha256:c84ac690f4a405e5248e6f96a8821a471ef41cb955bb74a38b0ef897b6a0c920

Observation 042b4cef-96fe-4d96-8fa7-c15e39bbe193 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward LLaVA-OneVision: Easy Visual Task Transfer

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T11:01:39.847533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:01:39.847533Z digest=sha256:07a3b0250aad87f57d2a6783500fd1c5b85f3dde07f0ce03ad70db9c46919f09

Observation 3c9cb6a9-59a9-47e1-a056-66c14a63f050 · outbound

This paper cites SEED-Bench-2: Benchmarking Multimodal Large Language Models.

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward SEED-Bench-2: Benchmarking Multimodal Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T11:01:39.852160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:01:39.852160Z digest=sha256:ee52c47eff8a280510ec3009966cdb0e5558d1ecbec1bd0472ca430b383aa0c7

Observation f82ada41-a7d9-4c4c-8d63-23da69b5f3d6 · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T11:01:39.857078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:01:39.857078Z digest=sha256:0ec1d6c508bc793b889c2e53a63db903113bdcd9fb97a2294529cbb53cd91611

Observation 5c99ab68-9c42-462c-9338-5a7ac2288f8c · outbound

This paper cites M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning.

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-16T11:01:39.862054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:01:39.862054Z digest=sha256:c562096103834c65c3c5bd6af776c973fc0c959037da8ee7945c287f002adae9

Observation 3863bc25-fa56-4808-aeda-0a59ec104726 · outbound

This paper cites HEMM: Holistic Evaluation of Multimodal Foundation Models.

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward HEMM: Holistic Evaluation of Multimodal Foundation Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T11:01:39.867457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:01:39.867457Z digest=sha256:c9566b6997399825a2a01ad053614a80e03d68afc971be0059e14003fdc87157

Observation 2b90c306-6cf0-4e7d-927c-d5ed2af54666 · outbound

This paper cites an unresolved cited work.

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-16T11:01:41.888068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T11:01:39.872051Z digest=sha256:72f15b90f3bd493b257e081b57b1279ed946a4ab83e3948c441c2cda4773f179

Observation 0bce2420-03b9-462e-a859-65f946ed7e04 · outbound

This paper cites an unresolved cited work.

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-16T11:01:41.866022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T11:01:39.876928Z digest=sha256:27d066cc793a55ff063e9c6d07e7f971d3f32daef60e06b7005c0a12072af2a5

Observation c9e1fed2-aa99-4f41-9449-1fbbe98bfa3a · outbound

This paper cites an unresolved cited work.

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward Unresolved cited work

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-16T11:01:39.882037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:01:39.882037Z digest=sha256:54fedc032101027073a5a633576489621d97196102210b7c9e4a44b9eeacb28a

Observation fac4da8b-d4c2-4e60-b8ad-ea8aacbae4db · outbound

This paper cites Visual Instruction Tuning.

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward Visual Instruction Tuning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-16T11:01:39.887255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:01:39.887255Z digest=sha256:2ae52d0f33b5157155fab4424dfc52a435a3a20007e6779025452bf9ad8cc8c2

Observation a6876fe4-bec2-4872-8f29-1fabff24bb08 · outbound

This paper cites an unresolved cited work.

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward Unresolved cited work

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-16T11:01:39.892089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:01:39.892089Z digest=sha256:f6d3719600441d4fb231cdf4a7648f808390992dede08485220f343949de5cec

Observation 898a838e-6404-4bf2-9d2d-4d6d06c0073c · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward MMBench: Is Your Multi-modal Model an All-around Player?

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-16T11:01:39.897134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:01:39.897134Z digest=sha256:7335a3bca655d2b03324d74b4e468a95254713334d61f776ffb004bf9864b437

Observation 7e00243f-d649-4568-8cf5-44f699ec3f4f · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-16T11:01:39.902066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:01:39.902066Z digest=sha256:b891b67b79d199d245649654d80c25817f0b2adc73fa650644f6c4cfd198e787

Observation feb1094c-effd-45c9-85bf-56c9a2e52b6a · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-16T11:01:39.907492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:01:39.907492Z digest=sha256:7c85b9c577c2b0a1c20ba4d88d0536ff28317b51aa8e17b3afcb1d0e2c8061d0

Observation ae93ff32-ac72-4f78-b76d-a369f3ecf321 · outbound

This paper cites RePaint: Inpainting using Denoising Diffusion Probabilistic Models.

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward RePaint: Inpainting using Denoising Diffusion Probabilistic Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-16T11:01:39.912475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:01:39.912475Z digest=sha256:ebcf7aa055fa1aad9fe3cfd1592de04c2114f2842de085c0f43f0fae6d27cd89

Observation 55f14170-9609-4de9-a0ee-8b47686a30f6 · outbound

This paper cites Mono-InternVL: Pushing the Boundaries of Monolithic Multimodal Large Language Models with Endogenous Visual Pre-training.

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward Mono-InternVL: Pushing the Boundaries of Monolithic Multimodal Large Language Models with Endogenous Visual Pre-training

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-16T11:01:39.917497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:01:39.917497Z digest=sha256:e358eb5790442bb74b8d8d5d529310aae12a6fc73e59179d4631f46b16dc61a8

Observation 33272fe0-72d2-4a0d-bcd6-7fd10d8f0f57 · outbound

This paper cites JailBreakV: A Benchmark for Assessing the Robustness of MultiModal Large Language Models against Jailbreak Attacks.

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward JailBreakV: A Benchmark for Assessing the Robustness of MultiModal Large Language Models against Jailbreak Attacks

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-16T11:01:39.922578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:01:39.922578Z digest=sha256:c6364d8f71ae4ac707210dd4f153209cc4a5c459bb24e14ec96ae63c0e703ca0

Observation e6638828-2023-4b5d-bffe-659279e774c9 · outbound

This paper cites Understanding the Effective Receptive Field in Deep Convolutional Neural Networks.

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward Understanding the Effective Receptive Field in Deep Convolutional Neural Networks

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-16T11:01:39.927595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:01:39.927595Z digest=sha256:455dab5e71902452d62f1feb512a171b33b2df3b2b7a8513f80286478c5eb9e6

Observation feec921a-8090-4aa3-9e7b-82fb30a57fe8 · outbound

This paper cites GPT-4 Technical Report.

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward GPT-4 Technical Report

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-16T11:01:39.932697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:01:39.932697Z digest=sha256:06836571288681c32e69cc0fa7186139f1c81b6255d3c8de212c4606adf0a194

Observation 0c773764-d7c0-4a67-a71b-260a501209dc · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward DINOv2: Learning Robust Visual Features without Supervision

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-16T11:01:39.937755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:01:39.937755Z digest=sha256:b5a70e9f8b16cf689af9de03440cd3d9957baa2a625d2407cb505939d57e0d23

Observation 6e2d15b7-193b-404c-bca8-c3895309f7ea · outbound

This paper cites Fung, Weizhu Chen, Minhao Cheng, and Furu Wei.

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward Fung, Weizhu Chen, Minhao Cheng, and Furu Wei

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-16T11:01:39.942314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:01:39.942314Z digest=sha256:a74010fe91d120780fd34896e9c93229e9cc7c900659ac7fb898bca0e4fd16a0

Observation 59fac2be-1fed-4eb4-93ce-d1457a97f957 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward Learning Transferable Visual Models From Natural Language Supervision

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-16T11:01:39.946890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:01:39.946890Z digest=sha256:2cf7fed849b34d67fd1267ac65fc62a19023d5cda5bdf4f47dc8e243f916e79e

Observation 8887fab4-65cf-4412-9199-5ecaedfbfb23 · outbound

This paper cites Do Vision Transformers See Like Convolutional Neural Networks?.

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward Do Vision Transformers See Like Convolutional Neural Networks?

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-16T11:01:39.951802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:01:39.951802Z digest=sha256:95c740b68131c9f1ec70e71e4bc7d5907344cd15a526c9b6fc5a8626b3042569

Observation 4fc1710d-c783-46fa-aa99-7f0434b99d22 · outbound

This paper cites an unresolved cited work.

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward Unresolved cited work

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-16T11:01:39.956494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:01:39.956494Z digest=sha256:865742ba61a2fb909ff4637caf2f0326f602edd494a3ac6c65a2e9c9912b78fd

Observation 8d42653b-fb23-4fa9-9598-e93da7c83d6f · outbound

This paper cites an unresolved cited work.

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward Unresolved cited work

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-16T11:01:39.961208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:01:39.961208Z digest=sha256:56de785f0d3f28a5fb71cd690b092d0daa0b606f0713f7a53f2206d8310e1003

Observation 6e7b3624-e1ef-44ea-b6c7-b0ad977327e2 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward LLaMA: Open and Efficient Foundation Language Models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-16T11:01:39.966661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:01:39.966661Z digest=sha256:fca16a6f9cc4135bd94e272ee3e3b390533ffb4b4f75da4309b86fd99acb0fda

Observation f4adc990-3dc9-41f2-94b0-f50969ae1594 · outbound

This paper cites Is ChatGPT a Good NLG Evaluator? A Preliminary Study.

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward Is ChatGPT a Good NLG Evaluator? A Preliminary Study

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-16T11:01:39.971339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:01:39.971339Z digest=sha256:0058ff9c56ef988d037b50f37a06cf67be4186d5eb99dceb0b29b8e5e4a49c5b

Observation 8b9c1474-9b5a-4f77-9bfe-1f1c04b2e7e5 · outbound

This paper cites To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning.

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-16T11:01:39.976358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:01:39.976358Z digest=sha256:f874ec4f8413c1ac85ba07e5b98d1fcfc0737a16776c3b6f92a396908093e8e7

Observation de59de8d-922f-4c0d-bb8b-d3596fbe2d23 · outbound

This paper cites Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception.

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-16T11:01:39.981135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:01:39.981135Z digest=sha256:f27445d8f69bc141f867f04343aa7f32bebdaeedbd2f911c36ef68aa6ee95f54

Observation 7ccbd96a-56da-4428-b16e-9c6134bf9dc0 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-16T11:01:39.985980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:01:39.985980Z digest=sha256:629bebb22ccd29aaa85c9b1e0e82b7947446f6a3cf0e721f8222f2454385ddf9

Observation 0d8d37d1-8d89-41ba-8052-6ca97de4ea45 · outbound

This paper cites VLBiasBench: A Comprehensive Benchmark for Evaluating Bias in Large Vision-Language Model.

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward VLBiasBench: A Comprehensive Benchmark for Evaluating Bias in Large Vision-Language Model

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-16T11:01:39.990724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:01:39.990724Z digest=sha256:147a4672f330217de0d0316001bb4e92a3abd65cac5d4e11d419af288d6a4ef8

Observation 3eee00be-a2b6-4228-9583-7cfae4873657 · outbound

This paper cites EmbodiedScan: A Holistic Multi-Modal 3D Perception Suite Towards Embodied AI.

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward EmbodiedScan: A Holistic Multi-Modal 3D Perception Suite Towards Embodied AI

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-16T11:01:39.995700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:01:39.995700Z digest=sha256:8f8334501f959e4bc29fe3e9f3066f8e401ea79a94b8797cf1beff3e5271cf0a

Observation 972bbe9d-0cd6-4950-ae6b-8ab2272aabcf · outbound

This paper cites CALM: Unleashing the Cross-Lingual Self-Aligning Ability of Language Model Question Answering.

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward CALM: Unleashing the Cross-Lingual Self-Aligning Ability of Language Model Question Answering

Reference 70

Resolution
verified exact
local_arxiv, observed 2026-08-16T11:01:40.448586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T11:01:40.000629Z digest=sha256:7da475c07d00676085f291ca177643e0e407220d7034f6ef83519828f0323de6

Observation d085772f-4ad8-4177-9522-0bbf42e505ff · outbound

This paper cites V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs.

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-16T11:01:40.010747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:01:40.010747Z digest=sha256:3aaab967a53fd6ccf07a16022721e0651269a20159b29dd6e9c29d7f36ab9f86

Observation 54803c1c-4f9f-43f5-8e39-a41e4de698cc · outbound

This paper cites an unresolved cited work.

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward Unresolved cited work

Reference 72

Resolution
unresolved
raw_fallback, observed 2026-08-16T11:01:41.824079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T11:01:40.016299Z digest=sha256:84856c3df7af01ef2e17efa00953b9ea662be6726ea8cf7a9dc0bd73d6f586d7

Observation e2415dfe-1ed6-4450-ad08-508b124350ed · outbound

This paper cites Sadler, Dinesh Manocha, and Amrit Bedi.

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward Sadler, Dinesh Manocha, and Amrit Bedi

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:01:41.805654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T11:01:40.020920Z digest=sha256:eafc686f03561fd5a6141306c720ee9255bdcf689353b6a7eecb9722a10fb1ef

Observation d3382be4-a904-4460-8a3b-1e7de1b34b7d · outbound

This paper cites CARES: A Comprehensive Benchmark of Trustworthiness in Medical Vision Language Models.

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward CARES: A Comprehensive Benchmark of Trustworthiness in Medical Vision Language Models

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-16T11:01:40.025936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:01:40.025936Z digest=sha256:00a58087210bd277382d6947135ed580498a2669c87746c0b595c125e9e0c6b2

Observation 0922a4bb-0077-4694-b316-048a59fc417c · outbound

This paper cites GPT4Tools: Teaching Large Language Model to Use Tools via Self-instruction.

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward GPT4Tools: Teaching Large Language Model to Use Tools via Self-instruction

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-16T11:01:40.030970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:01:40.030970Z digest=sha256:9cb55ad3224c7e6c2c3de160e4b8b839894e5d65f6770cd473f5b1d08b410693

Observation 6aa326c0-2ad9-43b2-aedd-a54905e6fd8e · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-16T11:01:40.035975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:01:40.035975Z digest=sha256:cd0615ba3a5bf94de6625ad160ed8d1f743f23414caef3c180b93353e5b82b3e

Observation 4c6d869c-3eb4-4f67-b1e1-6f2f1bfcabc1 · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-16T11:01:40.040858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:01:40.040858Z digest=sha256:818df081a68654094cd3513a3375d99f9ed2b805e068cf34399d169f748202b1

Observation a75edee2-559d-49d5-8e73-38109433dc06 · outbound

This paper cites AppAgent: Multimodal Agents as Smartphone Users.

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward AppAgent: Multimodal Agents as Smartphone Users

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-16T11:01:40.046199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:01:40.046199Z digest=sha256:a301f4e796b308ed11d7ac90fe33535b65fae1eded8d2cb0582d5690904e77bf

Observation f46a637a-2a90-47e8-a35d-77a7e578fd7f · outbound

This paper cites B-AVIBench: Towards Evaluating the Robustness of Large Vision-Language Model on Black-box Adversarial Visual-Instructions.

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward B-AVIBench: Towards Evaluating the Robustness of Large Vision-Language Model on Black-box Adversarial Visual-Instructions

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-16T11:01:40.051890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:01:40.051890Z digest=sha256:ea24c8a9d5cf233f24d95a4c4aba833103762fa026ccd6c1da75bfabffab33b3

Observation d7057694-7f5b-47c6-ba10-83085e0531f5 · outbound

This paper cites VLM2-Bench: A Closer Look at How Well VLMs Implicitly Link Explicit Matching Visual Cues.

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward VLM2-Bench: A Closer Look at How Well VLMs Implicitly Link Explicit Matching Visual Cues

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-16T11:01:40.057148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:01:40.057148Z digest=sha256:87e7144bf679458ad7586f19af366b254840f45eb95e8d7c75d05cb85adf3ce2

Observation 88c25f89-f030-4fce-b34f-1cf6a821bafd · outbound

This paper cites MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?.

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-16T11:01:40.063205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:01:40.063205Z digest=sha256:41e2f2c1f0dc9f6f78fe843934e41e7cb5567b1f707ab6606a97b9402e291d51

Observation dabc7c38-4091-404a-81af-d52e9aeafb39 · outbound

This paper cites SVIT: Scaling up Visual Instruction Tuning.

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward SVIT: Scaling up Visual Instruction Tuning

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-16T11:01:40.068908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:01:40.068908Z digest=sha256:1d79c0f08509a5daa12354626be79494775ded7e3b8a326352bc0fd2c635348d

Observation 80d7c525-3310-4514-b9ec-67852d498a50 · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-16T11:01:40.074153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:01:40.074153Z digest=sha256:385dbe7f756f62b73a98e63958f0b14a95f019edb286117705b418b0658fd764

Observation dc96f8b9-6047-48d9-ab5a-b7f50bfb9628 · outbound

This paper cites Calibrated Self-Rewarding Vision Language Models.

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward Calibrated Self-Rewarding Vision Language Models

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-16T11:01:40.079295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:01:40.079295Z digest=sha256:f15b92a5042187d146f22879f2418f77ed5325d5665fdb196bf8ff252704fb2b

Observation 168bf0b1-a005-45c2-8981-8aab8554305f · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-16T11:01:40.085249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:01:40.085249Z digest=sha256:d83256b0b90e577a922be60d06c393c7a49aae9f415ec4ef979d9dae49df5b98

Pith citing papers

No inbound Pith citation observations are available.