Pith. sign in

Paper Citation Record · LEDGER

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens

As of 13 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 0 inbound Pith citation observations for arXiv:2411.14725.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.14725 v2

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T15:03:24.836356Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

61 of 61 outbound references displayed

  • verified exact1
  • verified fuzzy19
  • unresolved41
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 29f63658-53bf-4aba-9f23-ce43457dc585 · outbound

This paper cites GPT-4 Technical Report.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:23.770624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:23.770624Z digest=sha256:a1802441285cc9a1da028ee1faf9b4897621cb36f609f7fbfbf6a9b1749d9919

Observation 17e6553f-9f94-4f44-aac6-1cd611c55cae · outbound

This paper cites UniBench: Visual Reasoning Requires Rethinking Vision-Language Beyond Scaling.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens UniBench: Visual Reasoning Requires Rethinking Vision-Language Beyond Scaling

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:23.786991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:23.786991Z digest=sha256:b8a6d133a720fc9e4bde47d3463f9d5eb7c272ea96ee058ed7d17b53b9fb04fc

Observation 4b7f9521-41e0-4bed-bb9c-4d07204918cc · outbound

This paper cites Claude 3.5 sonnet.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Claude 3.5 sonnet

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:03:27.104020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:03:23.818926Z digest=sha256:ca98f58b3db7372bc7142e3f312770ad090a30e4e4984b742eb2e0c6ebbfce8a

Observation 4dd63f1c-c7d0-4a21-b8ec-672b916c5bad · outbound

This paper cites Non-Determinism of "Deterministic" LLM Settings.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Non-Determinism of "Deterministic" LLM Settings

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:23.837472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:23.837472Z digest=sha256:8c12de50720aeb0cbdf57396e8274faa9d81d4883c232a2af64721d61cb719d6

Observation 010a7473-fc9e-4e04-957e-d292964a54b4 · outbound

This paper cites Qwen Technical Report.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Qwen Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:23.871962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:23.871962Z digest=sha256:88a52a227af5b66a45c7e48610a9670ed7e502e2c799d6ed5b1e6b3e10aeea3a

Observation f6e0c276-55a1-40d6-a0cf-39bbb76421d6 · outbound

This paper cites Qwen2.5-VL Technical Report.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Qwen2.5-VL Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:23.904276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:23.904276Z digest=sha256:5e8c744a1d61ab887fbb82f055db72940ffde83bafcd420ffed3f853d8411b18

Observation 592d8c6c-0a72-40e8-b66a-d1f9e90797b0 · outbound

This paper cites Eureka: Evaluating and Understanding Large Foundation Models.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Eureka: Evaluating and Understanding Large Foundation Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:23.957999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:23.957999Z digest=sha256:07e90a5dab3a5628374c8f8f15d6d896a7a7cdd04d2382c1fa55ec935d20e744

Observation d75277d9-30a9-4b51-958b-d3ff16e000d1 · outbound

This paper cites Language Models are Few-Shot Learners.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Language Models are Few-Shot Learners

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:23.965262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:23.965262Z digest=sha256:da7641b41a11b3698c17c5b4d226ec70b557363fc65a736c28dcd43c39aa13e1

Observation 9e811f0f-c327-4f70-9912-7d228f702d79 · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:23.972836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:23.972836Z digest=sha256:1ef7d38cc2b90e66bd7ebb9238323c6e5c1df88d84e210d1cb48bdaece056966

Observation 21a2f1b4-1b4d-4656-9d21-9cc70730fd1e · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:23.978719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:23.978719Z digest=sha256:de6049c6d5c46411e22e2d13476b599866072cb9cc28f4fe511d5c9ef9a2d0d5

Observation f054d048-c168-4083-8c5b-19c9e477b4dd · outbound

This paper cites NVLM: Open Frontier-Class Multimodal LLMs.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens NVLM: Open Frontier-Class Multimodal LLMs

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:23.984309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:23.984309Z digest=sha256:54eb7cf979506ff92403c80c24710ae05f324985f7b8b4d46a931bcec5ac0671

Observation bd75654c-1b80-4804-bf6f-7f17fc094c34 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:23.989904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:23.989904Z digest=sha256:08301a5508ed44e4b49600ead3a0924f12fe3293b8bcceaf9c6c5b1c0f6af547

Observation 142f118c-2802-4fee-89e7-c344e8f7350e · outbound

This paper cites GPT-4o System Card.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens GPT-4o System Card

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:23.995882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:23.995882Z digest=sha256:a3fe39ce1fc5ade517630e4adcd3b7723e7ca48c4112b6b48c261cf857f8e1c9

Observation 8958e039-5a48-4600-8a83-0d73af34ad82 · outbound

This paper cites Editing Models with Task Arithmetic.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Editing Models with Task Arithmetic

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:24.002580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:24.002580Z digest=sha256:1e2e9e0f3e23d3035cec4ef003e222c0139c66cf1905a0436e2ca58f88c5d603

Observation 8020696b-2d88-4235-9430-16c73e0b7b33 · outbound

This paper cites MANTIS: Interleaved Multi-Image Instruction Tuning.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens MANTIS: Interleaved Multi-Image Instruction Tuning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:24.010418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:24.010418Z digest=sha256:de7c1f18e02beb9afe3e7f54aab9bf4752a2aedef21a46b94dd2f06180027f28

Observation 53f61553-04d1-475e-b527-02875aa567fa · outbound

This paper cites Kembhavi, M.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Kembhavi, M

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:03:27.081220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:03:24.020558Z digest=sha256:09a7eaf63d210807e0fad5bbb938ef25dedd73527f096455ecf9cb89a51a34dc

Observation e10dfaac-e4cb-4126-8a3a-aac64e9c0850 · outbound

This paper cites an unresolved cited work.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:03:26.981473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:03:24.025849Z digest=sha256:2f1278ff54dc23453cf86280b003b7802b46c44ac17acad3d7a7913bdd91a024

Observation 8d2e9f02-2eda-4165-9a63-59487947c66c · outbound

This paper cites an unresolved cited work.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:03:26.956819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:03:24.031716Z digest=sha256:f25e96cb252f9aec18ebfb2b7e07ef3cd87a4a7590de9db58cda80c8546a59cb

Observation e8c890f3-d616-4444-b28c-1e4c30500148 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens LLaVA-OneVision: Easy Visual Task Transfer

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:24.037367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:24.037367Z digest=sha256:51209468348b75bd4cf612bb8840d2b80a935d40c19363dc8b27231ea5407603

Observation 021cca4d-b97f-457a-94cc-63aa3825ecdc · outbound

This paper cites an unresolved cited work.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:24.043058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:24.043058Z digest=sha256:25704951f94ec68e3cfd519d40482d1d16091a4d3ad219302e7fe2f29367fdb3

Observation 6b522d49-fecd-4b49-a8aa-e01599866372 · outbound

This paper cites an unresolved cited work.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:03:26.926888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:03:24.061742Z digest=sha256:d49e9a02d4fc21f33f03f92e8137f8b618bf55ab1c6cd37a5e8c480a60fb6094

Observation 5e19d423-3489-49aa-bfb3-53f14d09e3f6 · outbound

This paper cites an unresolved cited work.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:24.077979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:24.077979Z digest=sha256:a8d4f34f05abb2a53955619cf73c9c1bf8562d51270589a4877507cd8188f646

Observation 2db77923-8e89-4ab5-b815-4b1b233f7607 · outbound

This paper cites an unresolved cited work.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:24.105463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:24.105463Z digest=sha256:ab09c7e998fc80e4b2dafe537a397f7ee6448327e38f947a99d31f1370e32f82

Observation 263e4985-f8a3-42ad-92fc-1f79f61dd4ae · outbound

This paper cites an unresolved cited work.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:03:26.886416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:03:24.129970Z digest=sha256:e3c16ed81cf14d7e51f8b81c90e6c92b5db26a36d4262a37beee673d5b7af5b6

Observation 0edbedc9-203e-4399-ba68-21537ae5f149 · outbound

This paper cites POINTS: Improving Your Vision-language Model with Affordable Strategies.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens POINTS: Improving Your Vision-language Model with Affordable Strategies

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:24.193907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:24.193907Z digest=sha256:0db887eacb1617452cb653e3f268ac2321cf0041d78028e8fd04845aad29d06b

Observation 347de614-63f3-4d72-a2fb-a01538441053 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:24.261105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:24.261105Z digest=sha256:0a7e9e0f30ea710350c3291bbb18f2308f8fe68f995e50c04dff37854db5b335

Observation 2599b4e7-255d-4427-bf93-cc868e3785a3 · outbound

This paper cites Ovis: Structural Embedding Alignment for Multimodal Large Language Model.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:24.301248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:24.301248Z digest=sha256:7dfe612eee0ba1a9f22f76c2ceb1ec81519c2517b3130be1618a64bfd11463d4

Observation 0cec2fa3-9d64-4e00-bbf2-9a0d6e90c994 · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:24.319323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:24.319323Z digest=sha256:7eb61f2f01566b0433d0401a69a1c7b99ec96bcf5cbf7076d8602b55affd3b99

Observation b7f4f425-487d-4f19-9e97-541f6085ab88 · outbound

This paper cites Mathew, D.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Mathew, D

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:03:26.868949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:03:24.324376Z digest=sha256:60297f3d35acb8c63ed86437d1ff2ac1f0faac903351248ec882bb0de957d5a1

Observation bc9917e4-889c-470f-abda-c12dc38b9e3c · outbound

This paper cites Unearthing Skill-Level Insights for Understanding Trade-Offs of Foundation Models.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Unearthing Skill-Level Insights for Understanding Trade-Offs of Foundation Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:24.329453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:24.329453Z digest=sha256:674819ad846c909ec7e582fcf7d105bc62d730a2d12ab03c9549785b8e584eff

Observation 68d6272c-3291-41b4-a9df-6fe3b227fcde · outbound

This paper cites Gptv system card, 2024.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Gptv system card, 2024

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:03:26.847137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:03:24.340481Z digest=sha256:3864071b79b6e97bb288d0e55f877e8c2fe196b5cb67f036539ec6535b9f0ea7

Observation 0765cc2b-5eb3-474d-a2d0-b26d576bdb56 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens DINOv2: Learning Robust Visual Features without Supervision

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:24.345745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:24.345745Z digest=sha256:8666dad234ff80d31362c88697bca29a2c6313654d3209cb6dbe36a0672115b4

Observation 1bc1e516-d884-49a6-a17c-5ee536a18e6c · outbound

This paper cites an unresolved cited work.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:24.351317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:24.351317Z digest=sha256:2595014ad7c051d4ec9995a7486635b3f2d29344709473698c30ddbcb8c41cfa

Observation e991f5d8-d7e9-4a40-9872-834a9dc98c8d · outbound

This paper cites Radford, J.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Radford, J

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:24.356784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:24.356784Z digest=sha256:00b94e0a82b80ff7f62f5bb0edd93c274b235270f754a8dc8099376c24b74ac7

Observation ceee076c-99d7-4175-8225-be5f4adbde4a · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:24.362312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:24.362312Z digest=sha256:9bcb2592dd5318f530f08f864f1bb5cded8070e70bb03eace26f82e8b20aa1f9

Observation caf8f6ce-7441-477c-8a97-eac1e07e505c · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens LLaMA: Open and Efficient Foundation Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:24.367607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:24.367607Z digest=sha256:c789d13bbef64418bec01651b76cad89939a2523526fc895c6bd6ba51e8e7b5a

Observation 00d683b2-4ffb-4c06-b471-c4b5326f1ec2 · outbound

This paper cites MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:24.373327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:24.373327Z digest=sha256:0ce30cc1794d6ec4025ca9c7d2e3057d3459d59e1039ddcd4ba5056a345fc5a6

Observation a72697f6-6066-407f-bd7f-d8b07573fb79 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:24.378871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:24.378871Z digest=sha256:d7ab6d5eb696a51d87417f45661ec716e7715368e57a6e7973573673fb6d90a1

Observation 224732bf-b199-4101-8915-c6a23b8bba40 · outbound

This paper cites Bounding Box Stability against Feature Dropout Reflects Detector Generalization across Environments.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Bounding Box Stability against Feature Dropout Reflects Detector Generalization across Environments

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-08-12T15:03:25.023717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:03:24.384210Z digest=sha256:8f434264549b5e568b94a587e91a9dabfaba490f814f3918f924ed214ec43ace

Observation 1c421364-0ecb-4800-aa2c-120399cb8363 · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens MMBench: Is Your Multi-modal Model an All-around Player?

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:24.389774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:24.389774Z digest=sha256:68d8ee0f0814f0b7e44e74f6ed10b38ca996aebefe9cf7c8b06813b53b07ad9d

Observation 4a1c64b2-93a3-4586-846b-58f0c01c7b66 · outbound

This paper cites an unresolved cited work.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:03:26.679958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:03:24.395596Z digest=sha256:34f98f0810ea9953fd1198d9f5aebec73f019bd6fd195fa72da6f58f50dcf50c

Observation 61c2b739-a403-4914-a3cd-718564c4f772 · outbound

This paper cites an unresolved cited work.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:03:26.618098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:03:24.415999Z digest=sha256:f22a3d82a57ca3f8bff5262d83ab1dbf67d43f58d9090924e0ebfafdd02cca37

Observation 624ae08f-463a-4c14-b7e6-c1c29655e177 · outbound

This paper cites LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:24.449765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:24.449765Z digest=sha256:4ae6b2bff05f1546322351b2e3d8ec3aa77d6faf8a8c8fbafdf830ce8f72415f

Observation 07c16031-c15b-45eb-8a98-14b19ba41a23 · outbound

This paper cites Unveiling the Tapestry of Consistency in Large Vision-Language Models.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Unveiling the Tapestry of Consistency in Large Vision-Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:24.493884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:24.493884Z digest=sha256:24a8c7ec8a3b456324ca98d9d319a56f815c39dc12c5645269e872ea288b64da

Observation d1ca39bc-2e83-4c11-a2d1-f441b52e5a69 · outbound

This paper cites Benchmarking Multi-Image Understanding in Vision and Language Models: Perception, Knowledge, Reasoning, and Multi-Hop Reasoning.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Benchmarking Multi-Image Understanding in Vision and Language Models: Perception, Knowledge, Reasoning, and Multi-Hop Reasoning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:24.558647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:24.558647Z digest=sha256:f8f7478336a123601cc1fc2c60493d29c7cb11bdaee85071e14062021498e910

Observation 45ff90b9-7d09-4f1d-8fd9-6c536de9e14b · outbound

This paper cites Guidelines: • The answer NA means that the abstract and introduction do not include the claims made in the paper.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Guidelines: • The answer NA means that the abstract and introduction do not include the claims made in the paper

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:03:26.598812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:03:24.636112Z digest=sha256:ed224fc27f7611d2429c18b730e04ffb7a5b288b00b7507f8418ed724d8d5318

Observation f7916f18-88b2-4c1a-bc82-43740aec36d6 · outbound

This paper cites Limitations.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Limitations

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:03:26.526008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:03:24.658335Z digest=sha256:a7b9b7a90d195a59663aa49c0458dd361d91f8dad9a70ccef4d8c8a511e2e3ce

Observation e12fbe19-6997-4f7d-909e-b17297f5c9cf · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include theoretical results.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Guidelines: • The answer NA means that the paper does not include theoretical results

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:03:26.402925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:03:24.684847Z digest=sha256:85e0807ec5b8dfe95816692a1885ef5d0be651182487d0b12e5f2892ccb5ad2f

Observation 20bebfd8-650d-49d4-ac7d-ae17e22fbe7f · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include experiments.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Guidelines: • The answer NA means that the paper does not include experiments

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:03:26.384883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:03:24.703238Z digest=sha256:21655bd20643d83e4a5782b300c5e3c7e5d6cbfbf91942bf476cce58dd4d60d3

Observation 64bf1c8a-3c2f-4701-b9c2-65e916ba274a · outbound

This paper cites Guidelines: • The answer NA means that paper does not include experiments requiring code.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Guidelines: • The answer NA means that paper does not include experiments requiring code

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:03:26.367041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:03:24.722927Z digest=sha256:6a0630d05fc4a1d114e7d5996f01e5d8ed17a269a5d60a8dbb226b6ae68e85af

Observation f41e0017-36f1-4589-a6d3-7c23cf26867a · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include experiments.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Guidelines: • The answer NA means that the paper does not include experiments

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:03:26.347802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:03:24.742266Z digest=sha256:db3cfeafbeaa98af3c83c83c12b51f462c3b8453b11e44819d37d5725c912f1d

Observation d627cf61-1a2a-44b5-b3d9-84ef68722d59 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include experiments.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Guidelines: • The answer NA means that the paper does not include experiments

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:03:26.201497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:03:24.774916Z digest=sha256:ea15a73a18a8cd5b520e24bbbbb86fc848efe4543fc780e506581b8a8ea5a972

Observation 908586c6-847a-490b-974e-25670ae78301 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include experiments.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Guidelines: • The answer NA means that the paper does not include experiments

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:03:26.130394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:03:24.781282Z digest=sha256:fcbe4c4a5b8000b6673169721df8a90e6a86d5829b350686b28563657e15f494

Observation d8408868-447e-4c1e-9f17-52ab8c050c08 · outbound

This paper cites Guidelines: • The answer NA means that the authors have not reviewed the NeurIPS Code of Ethics.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Guidelines: • The answer NA means that the authors have not reviewed the NeurIPS Code of Ethics

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:03:26.104548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:03:24.788121Z digest=sha256:632371b0dfbf81db8822be15a014754b2db169e5c4a3dc281197730ce0514586

Observation 359ef756-7ca3-479b-b8e2-0173901c1f08 · outbound

This paper cites Guidelines: • The answer NA means that there is no societal impact of the work performed.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Guidelines: • The answer NA means that there is no societal impact of the work performed

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:03:25.979567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:03:24.793680Z digest=sha256:eb15b3980706d6aa3a20860fd9f4e392a5152978ea24d0c47b5bc4ee3d189a5c

Observation 9b67f61e-c732-45a4-8ca9-4e1c04f7cc21 · outbound

This paper cites Guidelines: • The answer NA means that the paper poses no such risks.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Guidelines: • The answer NA means that the paper poses no such risks

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:03:25.879257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:03:24.800239Z digest=sha256:012800982893686732813aa8473a3314564e5b1d885a32996cfad159a6fa3ced

Observation 203b5458-eb86-4803-895f-f06ba7791d03 · outbound

This paper cites We will elaborate on it on the Appendix.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens We will elaborate on it on the Appendix

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:03:25.856187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:03:24.806116Z digest=sha256:f5700d192593dddb59aef04348a12d161992c3787ed32f73583cea3d33756f7f

Observation a65abba1-5764-4268-b798-b0d52d976d09 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not release new assets.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Guidelines: • The answer NA means that the paper does not release new assets

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:03:25.835417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:03:24.812627Z digest=sha256:20f51e04829767d9749e7e4177597ec25a04c7a35d36d5b8ddfb7c2702b09845

Observation f4685861-48c2-48e5-87dd-eb3e2c22db39 · outbound

This paper cites an unresolved cited work.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:03:25.780005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:03:24.821604Z digest=sha256:3c50bd11a3e8c219373fa29398a4c58453be738762cb4c57814037860ba41dbd

Observation 634334d2-a7d1-4e3d-9c9c-61e4764e8d8b · outbound

This paper cites Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:03:25.665905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:03:24.829140Z digest=sha256:88b1ccbb0763ce32175573d4e9d05de5496b3e331190dbc467c11f5a299894f4

Observation b6332e21-180d-40d5-93e2-f7353a5f43ba · outbound

This paper cites Answer: [NA] Justification: We use LLM for paper writing.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Answer: [NA] Justification: We use LLM for paper writing

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:03:25.647440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:03:24.836356Z digest=sha256:a1f3cc63647387eeba6af824c6e07318e265314ed3c182d71c4b6054be5c7aee

Pith citing papers

No inbound Pith citation observations are available.