Pith. sign in

Paper Citation Record · LEDGER

LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

As of 11 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 76 inbound Pith citation observations for arXiv:2407.12772.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.12772 v2

Coverage vector

measured 24 of 24 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-17T05:19:22.423762Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 76 of 76 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T21:34:35.657096Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

24 of 24 outbound references displayed

  • verified exact0
  • verified fuzzy11
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch7

External citation measurements

1
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 41a065a0-33ba-4505-96d7-c12348c473d8 · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T05:19:22.442626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T05:19:22.423762Z digest=sha256:d16380f80cefa21065855c412532a717544160eacb42a4e54629131d8c71dc81

Observation ac237b95-9dcc-430d-8a2e-a64a6f60f2d4 · outbound

This paper cites InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD.

LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T05:19:22.446753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T05:19:22.423762Z digest=sha256:c0941c6d49bc1b5c2a31b0ee95d516bfb28feeff94dc4de1373e09af31f96661

Observation 4eaf4965-7af9-469c-bf9b-d742213b2150 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T05:19:22.453887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T05:19:22.423762Z digest=sha256:451d32fb7ccbefb372dd356ae2c72be2287cb2d9c8f918f8ecd55901ba35208a

Observation 46622577-559c-4ea9-afda-31666510f829 · outbound

This paper cites Making LLaMA SEE and Draw with SEED Tokenizer.

LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models Making LLaMA SEE and Draw with SEED Tokenizer

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T05:19:22.459326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T05:19:22.423762Z digest=sha256:371ebfbb7d1a75edcf53bc5ccbfc0820a95a9a3d7423322a97ef7001747ce5b8

Observation 61f9941a-f4a7-4bd6-99a1-144fd10c0c72 · outbound

This paper cites A Diagram Is Worth A Dozen Images.

LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models A Diagram Is Worth A Dozen Images

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T05:19:22.463601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T05:19:22.423762Z digest=sha256:b9107b3e48495a87e2f045a42dcc1b1fec516ae74691745cef18a7a1119627cc

Observation da6a5df2-38fc-4d4c-9498-99235737c321 · outbound

This paper cites Coresets for Data-efficient Training of Machine Learning Models.

LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models Coresets for Data-efficient Training of Machine Learning Models

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T05:19:22.468240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T05:19:22.423762Z digest=sha256:fbaa11133d05e515aa0064202a07a5338388544136de7ece9ba7505bd5167af1

Observation a858ae38-1d42-40f8-a4b0-022b15c8c58e · outbound

This paper cites tinyBenchmarks: evaluating LLMs with fewer examples.

LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models tinyBenchmarks: evaluating LLMs with fewer examples

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T05:19:22.473153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T05:19:22.423762Z digest=sha256:f20b8ceedb6270f6b710733690c3f3e67346a102369d0382ce3168773c827bb1

Observation 99525b92-fa84-4b7b-b28d-7c7188b2ea32 · outbound

This paper cites What are the key points in this news story?.

LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models What are the key points in this news story?

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T05:19:22.476018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T05:19:22.423762Z digest=sha256:b2f62ffd62c594f05caddcef95d4d75bb76ed0bef65e70f7ed12b944d15b7500

Observation f0f44743-f5e5-43c0-8f27-41cc25480f27 · outbound

This paper cites What are the factors that led to this event?.

LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models What are the factors that led to this event?

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T05:19:22.478515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T05:19:22.423762Z digest=sha256:687cc9eb1da54c91cb178625810767a0bd6f766503a22fab04df6207faab32ad

Observation 3a96d9ad-e6f5-4fc5-b812-2740fede8ed1 · outbound

This paper cites How could you create a new headline that captures the essence of the event differently?.

LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models How could you create a new headline that captures the essence of the event differently?

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T05:19:22.481198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T05:19:22.423762Z digest=sha256:4f203cf18fea186c2041cb440273d515d8a53d7c94f4e7317afba15cf9f05510

Observation 93846a10-9fcd-4507-87b2-4941b3b630a5 · outbound

This paper cites Please present this news in Arabic and output it in markdown format.

LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models Please present this news in Arabic and output it in markdown format

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T05:19:22.483602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T05:19:22.423762Z digest=sha256:1b93f44b08ecafa1cffd4271f5ccf9c3fcd8903d8d13d5f6e49f77c0471cc62a

Observation c2f125d4-5c36-471f-a3ce-7f7e6e922df5 · outbound

This paper cites an unresolved cited work.

LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-05-17T05:19:22.485651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T05:19:22.423762Z digest=sha256:f242a35e68d88fe3c2a91aa2d506cd50469f16640b1b278889a0b82904ec3919

Observation 407d8bbd-82cb-46f0-b4c0-b6721acb82b8 · outbound

This paper cites However, you should not change the original question's subtask unless the original subtask is not one of these five.

LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models However, you should not change the original question's subtask unless the original subtask is not one of these five

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T05:19:22.488064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T05:19:22.423762Z digest=sha256:fe4dbf9983755a6da0e81b4d4824c29849e926d473ab34102c6754d1a3635894

Observation 5fd5be42-25f9-49ad-9ba7-988414396e6c · outbound

This paper cites an unresolved cited work.

LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-05-17T05:19:22.489979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T05:19:22.423762Z digest=sha256:23e8ec1203930b61b05924fb2b3843306a093c6916d95b7ed4522da167b0d132

Observation 97641f0d-0d6d-48e0-9837-50841ceb6d36 · outbound

This paper cites But don't use python−like format.

LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models But don't use python−like format

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T05:19:22.491916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T05:19:22.423762Z digest=sha256:4308625d20feccbe3850873cfa1a50ac6df26306e8a9b77138dca5a531362801

Observation b1524ccc-abb8-46ab-90db-695baa3ae091 · outbound

This paper cites an unresolved cited work.

LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-05-17T05:19:22.494384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T05:19:22.423762Z digest=sha256:d5fab7814f2381e71141fbab3883bb0c8e71083730eec094802d3f6ad9a56e76

Observation 9c4ebe4c-e8df-49a9-9a55-24e8c8704085 · outbound

This paper cites But if you think some words should be in other language, you can keep it in that language.

LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models But if you think some words should be in other language, you can keep it in that language

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T05:19:22.497093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T05:19:22.423762Z digest=sha256:0c2d0072429fa15192825d0866e58e2dd1fe3edd8926cad866a7d82d597d610c

Observation afde4456-461d-4bb1-bd93-9931c860ad57 · outbound

This paper cites an unresolved cited work.

LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-05-17T05:19:22.499178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T05:19:22.423762Z digest=sha256:211a468d6dd1d5ee8b8bbb577f07c52ddc74df9bef9dc58cd25a2bdc181e4d0f

Observation 38f1f065-e613-4cca-8aef-84769a6ff0c5 · outbound

This paper cites an unresolved cited work.

LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-05-17T05:19:22.502067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T05:19:22.423762Z digest=sha256:5dd0ffc56522ccfdc9c2ee271a818357dc3a76582c4f4556e6448079f8cd0b2b

Observation 085a2ac6-cff5-4219-98d2-f4d184ffc838 · outbound

This paper cites an unresolved cited work.

LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-05-17T05:19:22.504065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T05:19:22.423762Z digest=sha256:b16db7418c2ccea1977a61ad8dce70742e24027258673e9130a49eedf64a81a8

Observation bda072aa-ad67-4127-b2c9-a50814044507 · outbound

This paper cites Some tips.

LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models Some tips

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T05:19:22.506116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T05:19:22.423762Z digest=sha256:f3a04215d8500c850d87e24f9990288f5ab82d59e96b924f5071d411337ea788

Observation e406d34c-a761-4cdd-82ff-faf783f8f100 · outbound

This paper cites In such cases, you can relax the criteria slightly.

LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models In such cases, you can relax the criteria slightly

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T05:19:22.508398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T05:19:22.423762Z digest=sha256:011e7e2edfe048aba2aa1589b3a3d8b4968dd1e87b83592346dd691ef5c6d194

Observation 6762720f-b6fc-4931-9d4f-e91c876f3a3c · outbound

This paper cites Explanation.

LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models Explanation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T05:19:22.510562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T05:19:22.423762Z digest=sha256:5f31cf168fbf74c4a1bd299458889bd426bde58eb73ebdf873e68cea99ecb84c

Observation 68a4ab78-6909-40d3-9419-bd66816e0b28 · outbound

This paper cites This kind of systemic shift often results in skills becoming obsolete, leading to higher unemployment among professionals who cannot quickly adapt to new technological paradigms.

LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models This kind of systemic shift often results in skills becoming obsolete, leading to higher unemployment among professionals who cannot quickly adapt to new technological paradigms

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T05:19:22.512802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T05:19:22.423762Z digest=sha256:3d1144a88ca55f52bbcb72df26477646a5f7cf3af4cefd3e00b45f10250df3de

Pith citing papers

Observation edaf3815-d97d-40ae-83b6-143507c5c40f · inbound

LLaVA-OneVision: Easy Visual Task Transfer cites this paper.

LLaVA-OneVision: Easy Visual Task Transfer LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 161

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:19:22.513527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:c244895771c47b627550639ca4d9f5000f8d7127518c4c59f3b55346f7753b0e

Observation fcf13dd4-e536-4b33-b2be-418a60ae91cd · inbound

LLaVA-Video: Video Instruction Tuning With Synthetic Data cites this paper.

LLaVA-Video: Video Instruction Tuning With Synthetic Data LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 165

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T05:19:22.513527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:af0b70fa5e68019a1d9a25d2d3d52810d16f57bd8878c465e0d56aeedd5cb686

Observation 7f8d7f30-ef47-4ce8-b292-ded9b4ed6754 · inbound

Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces cites this paper.

Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 100

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:27:44.058528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T09:27:43.919941Z digest=sha256:40c58fcf583864f4989b46b09ba51f4dfa180be8244d1cc5b63a76b09a480758

Observation 31d738d5-6381-4128-8300-1446b21f4267 · inbound

Feedback-Driven Vision-Language Alignment with Minimal Human Supervision cites this paper.

Feedback-Driven Vision-Language Alignment with Minimal Human Supervision LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 113

Resolution
unresolved
no resolver link, observed 2026-08-10T21:34:35.657096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:34:35.657096Z digest=sha256:ab98bd4cbd5bb76969af2e2679bce018d2d578ef0433f61042a57854bb7355ab

Observation 86ea2dec-3ff0-4369-8491-984c2e3211bb · inbound

Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding cites this paper.

Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-10T18:53:21.329427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:53:21.329427Z digest=sha256:890bea707a314440e4325e994b87dc02c7e321fc6bebbe66ab2cbe1ddfa71852

Observation 8525a6ab-88cd-4869-90a4-a9bd1a23a7bd · inbound

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos cites this paper.

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:19:22.513527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T00:32:41.059558Z digest=sha256:495a51fd8364d5b3876704ea3131128d71714d00b8fc5b84f3b6365eb133ea9d

Observation 5b9854d1-d8a6-4e5e-b9a3-3af0b15ddcbb · inbound

Temporal Preference Optimization for Long-Form Video Understanding cites this paper.

Temporal Preference Optimization for Long-Form Video Understanding LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-10T15:35:30.301738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:35:30.301738Z digest=sha256:2f8bced25c97e5884b5c5b8af72737aef9b9b04240a6517a6d2ebc8a5b01b473

Observation 22a8c01b-60d0-4d79-9f99-e02db2c23309 · inbound

Mordal: Automated Pretrained Model Selection for Vision Language Models cites this paper.

Mordal: Automated Pretrained Model Selection for Vision Language Models LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-09T19:46:19.974377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:46:19.974377Z digest=sha256:ad0174eab78ef57d323cb23c98b11f086a615469eebc37f79accb9ec9f7219ac

Observation 6dde0cbd-6c00-4a84-a005-0ae1a1f197c0 · inbound

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models cites this paper.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:56.014437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:56.014437Z digest=sha256:141fc03d1f5e86fe150ce3a09879899828d443ea85a16f8c78ca649638fef264

Observation ace4d7a1-6410-4a39-b4fc-15bec714d7f4 · inbound

From Visuals to Vocabulary: Establishing Equivalence Between Image and Text Token Through Autoregressive Pre-training in MLLMs cites this paper.

From Visuals to Vocabulary: Establishing Equivalence Between Image and Text Token Through Autoregressive Pre-training in MLLMs LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T22:43:59.496502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T22:43:59.496502Z digest=sha256:03a855589319552cf92a625c732196c86e084fb10cae28f7447fa1ac78cb1f2a

Observation c1624128-ec7a-482d-8661-7f01e487293e · inbound

Granite Vision: a lightweight, open-source multimodal model for enterprise Intelligence cites this paper.

Granite Vision: a lightweight, open-source multimodal model for enterprise Intelligence LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 103

Resolution
unresolved
no resolver link, observed 2026-08-07T20:06:00.974895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T20:06:00.974895Z digest=sha256:56bc7e55de5205112dd2f9bb74d5729ad10f25b62c06466dbb8fbe17d66cf471

Observation 9959d856-1575-44aa-aef0-cd26b4beca69 · inbound

LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL cites this paper.

LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 91

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:19:22.513527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T15:15:46.255296Z digest=sha256:98ae4e3f1fedd7127cf8845527b663d1ce7f814ac0cd1139f81301ffddd9f819

Observation cdfe6a61-b8c6-4b66-8526-ef453138e929 · inbound

Clapper: Compact Learning and Video Representation in VLMs cites this paper.

Clapper: Compact Learning and Video Representation in VLMs LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:48.289056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:48.289056Z digest=sha256:3676a5268724de16cf1d761a82860f075a4c6edf113350487b5c0ff8d9065291

Observation a1a68158-024f-4428-8b46-b41501f1b8a0 · inbound

QuickVideo: Real-Time Long Video Understanding with System Algorithm Co-Design cites this paper.

QuickVideo: Real-Time Long Video Understanding with System Algorithm Co-Design LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:02.207599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:02.207599Z digest=sha256:490773169414c1df4b72429f9d53d13d572cb959d967f17cee7e8d8f83a796ce

Observation 541fab43-345b-4c5e-af52-0c0dcbfe5899 · inbound

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning cites this paper.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 112

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:19:22.513527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:eedb346e2b6fd12e0baaced662ac0425f332458317a7c260aadb697bb79f98dc

Observation 665b20c0-642a-4749-adce-bb09dbca8641 · inbound

Inference Compute-Optimal Video Vision Language Models cites this paper.

Inference Compute-Optimal Video Vision Language Models LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:45.026613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:27:45.026613Z digest=sha256:55c3aecd3784b22af63595aa6812c50d08c84884e71b245313a941952fa1f55d

Observation 94ea1965-e3a7-4fd4-9eec-5b6d1d628f21 · inbound

AdaTP: Attention-Debiased Token Pruning for Video Large Language Models cites this paper.

AdaTP: Attention-Debiased Token Pruning for Video Large Language Models LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:08:10.490334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:08:10.490334Z digest=sha256:2439a40316e9766a95af1045edbbf20af336bdd853d2e4ba7e782aa229c3bd90

Observation b4a2bfdf-d0be-4e4b-bf84-6059ee7d55e2 · inbound

ID-Align: RoPE-Conscious Position Remapping for Dynamic High-Resolution Adaptation in Vision-Language Models cites this paper.

ID-Align: RoPE-Conscious Position Remapping for Dynamic High-Resolution Adaptation in Vision-Language Models LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:58.047950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:58.047950Z digest=sha256:ba9437c4f7c613c6ecdce5ecf4915a8f93932418616c8a4b84f255be871a9492

Observation 154b2e67-dfd9-49d3-8247-306637013fb2 · inbound

Fostering Video Reasoning via Next-Event Prediction cites this paper.

Fostering Video Reasoning via Next-Event Prediction LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:45.965589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:45.965589Z digest=sha256:bc6c2d7af0ac5b307e79ca327e1b23938bcd2240f814a7e7cc7edfeaabd21633

Observation c1de9b23-f463-4c5c-a580-d41614ee40de · inbound

Spoken question answering for visual queries cites this paper.

Spoken question answering for visual queries LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T12:52:43.377377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:52:43.377377Z digest=sha256:2e4ed802fb33b80d152eb9e110cda9ec61c10a7aed8b381917a7ff2305fde0ff

Observation cb348a27-c885-4a0a-9667-9fb0ed96adda · inbound

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models cites this paper.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:50.123928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:50.123928Z digest=sha256:49bce17ee58782e2cbeb99dc245e03b464eeb28e4854acdb78c50464ce09e50e

Observation 743b0234-2b6f-4c89-a45c-ca3607796007 · inbound

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning cites this paper.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T12:19:33.786626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:19:33.786626Z digest=sha256:4793c60dd710fa69e9e4bcd2c7026ee10e3e679fc53509c597189ec122bfc294

Observation f780789c-d041-484f-bd98-f944e3e6b659 · inbound

EffiVLM-BENCH: A Comprehensive Benchmark for Evaluating Training-Free Acceleration in Large Vision-Language Models cites this paper.

EffiVLM-BENCH: A Comprehensive Benchmark for Evaluating Training-Free Acceleration in Large Vision-Language Models LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:10.660378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:09:10.660378Z digest=sha256:104b8d3434add5f613cb898ed143b194f32ef734ba7164cf12b0c36557efb11d

Observation 8826dfff-bcbe-4a51-859c-962a6a42fd2b · inbound

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding cites this paper.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:10.390042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:10.390042Z digest=sha256:3da2bdca981f8b01d7fc563889bf01c4196535335b1fcf3acb8b3c01849956ba

Observation 399f54ee-1809-460e-a0d9-df04827b9876 · inbound

MiMo-VL Technical Report cites this paper.

MiMo-VL Technical Report LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.284762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.284762Z digest=sha256:b411ec4baa2fb3165864c6b89501fa1470a0f0752148dd1bd0c0f1b9afcc822a

Observation 29db3ad9-d287-4e59-8754-7cbb96c444d0 · inbound

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding cites this paper.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T10:56:05.407052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:56:05.407052Z digest=sha256:ac72b8622d55f1401fd40c03d677cf7e64c91ac3f0e011029642bb3e00cd683d

Observation 1b5cad25-5993-4d34-a0d1-5e6058d88cee · inbound

VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in Videos cites this paper.

VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in Videos LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T10:25:36.338942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:25:36.338942Z digest=sha256:2a5005aa64879b3392d2f1f223d0e9a735a08afdd2be1762bce4ae3117c6230d

Observation aa89350e-5ea6-4487-b41b-6fccf2ed5fd6 · inbound

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing cites this paper.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:53.641537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:53.641537Z digest=sha256:621cdf0296351ecf77e110ed0535fb20ffe76a3ef7a9005685bb3bcd1b6e4a48

Observation e976c28f-cb3f-4be5-afc5-2c4129af7882 · inbound

Vision-EKIPL: External Knowledge-Infused Policy Learning for Visual Reasoning cites this paper.

Vision-EKIPL: External Knowledge-Infused Policy Learning for Visual Reasoning LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 66

Resolution
verified exact
local_arxiv, observed 2026-05-19T10:37:14.950898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T10:34:48.849524Z digest=sha256:7784d9404d63c30aa14ec31a9499fd9303f3d3aaf1200b4475023b39dff7d7df

Observation 0b2d9593-f06e-45e2-a843-0a5d394ef7e1 · inbound

FlagEvalMM: A Flexible Framework for Comprehensive Multimodal Model Evaluation cites this paper.

FlagEvalMM: A Flexible Framework for Comprehensive Multimodal Model Evaluation LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T05:14:45.044853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:14:45.044853Z digest=sha256:0ca2cc62d82daa24607a9597250a3673fd0c2cc31c019d949631e8d643d24d16

Observation 15bd45bc-a1b8-4c77-8dd8-83626e44a128 · inbound

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning cites this paper.

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:19:22.513527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T00:33:50.471804Z digest=sha256:c35c54d2123a08ef4a2bc09971d952a6b4391d03ee07f75c54c4637d130df27c

Observation 1b0053e0-2ad7-40d6-9e37-341dea810c36 · inbound

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs cites this paper.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:41.016343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:41.016343Z digest=sha256:f1293f0f577c6e9dab97510330271674b8258d424439385afbf51d0a0c1597f5

Observation 842f749f-59b3-4d5a-96e5-45f2ff031e20 · inbound

VGR: Visual Grounded Reasoning cites this paper.

VGR: Visual Grounded Reasoning LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-05-19T09:12:14.504288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T09:11:00.295700Z digest=sha256:3f5d2ae85406217c2cf0b306effff241765070a90fd64f29e72858a4d4891410

Observation 1805c7fa-a03b-4a40-92f1-0c7d95361a3f · inbound

GreedyPrune: Retenting Critical Visual Token Set for Large Vision Language Models cites this paper.

GreedyPrune: Retenting Critical Visual Token Set for Large Vision Language Models LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T00:42:53.728937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:42:53.728937Z digest=sha256:701bf728300477a884e3e7078679bb901ff7903e4e645a03c7c593248d75b61f

Observation ed01f621-792c-4699-afb1-f54e0c1f7f6d · inbound

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation cites this paper.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.967404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.967404Z digest=sha256:fb3d83b446fdf5595b0fe7070b14236a6094ccfbd5284fc745a3affdb65927f3

Observation 3ca77f38-1b04-454e-b51f-3d1e2a2bf5e2 · inbound

MMSearch-R1: Incentivizing LMMs to Search cites this paper.

MMSearch-R1: Incentivizing LMMs to Search LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:19:22.513527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T15:27:04.228144Z digest=sha256:159a3922d65fba67cf0f97f2e4db4c0556d994e5f5a39dbf249fa4a5961e7abe

Observation 5d985fbf-4b73-481f-b255-2b737e68829e · inbound

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs cites this paper.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:04.686714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:04.686714Z digest=sha256:101b7806f71b7b9992ddfe31fc97c437bc669f273c3d0c14c65f151fcbebae7c

Observation 94ef0cc7-fe5a-4dab-97df-7275881dcd82 · inbound

VisionThink: Smart and Efficient Vision Language Model via Reinforcement Learning cites this paper.

VisionThink: Smart and Efficient Vision Language Model via Reinforcement Learning LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-06T16:34:00.823567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:34:00.823567Z digest=sha256:04cca014a414ff813add90dacaa4130b109a0f1cd2850203e7bf8b4d8428174c

Observation 20eedbe1-69b8-46a6-adc3-00dd90e4df93 · inbound

EgoPrune: Efficient Token Pruning for Egomotion Video Reasoning in Embodied Agent cites this paper.

EgoPrune: Efficient Token Pruning for Egomotion Video Reasoning in Embodied Agent LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T15:38:45.376369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:38:45.376369Z digest=sha256:c557bf8b856b02f0872cc499617d78b2c0863c0290ad658f50f2761d32e5c998

Observation 99f4e4f6-d6ea-447e-b0c2-7f198124978f · inbound

Object-centric Video Question Answering with Visual Grounding and Referring cites this paper.

Object-centric Video Question Answering with Visual Grounding and Referring LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:39.360265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:39.360265Z digest=sha256:55db6fcd0dc04c77c01add3bb6a32c7d02a23416bef6490ff48e9e660a5993fa

Observation 932c4614-d28b-4ef2-a0e8-434d0e20d706 · inbound

A Glimpse to Compress: Dynamic Visual Token Pruning for Large Vision-Language Models cites this paper.

A Glimpse to Compress: Dynamic Visual Token Pruning for Large Vision-Language Models LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T05:37:35.543694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:37:35.543694Z digest=sha256:6daba25e0d4af27b16bd6aefac475f752dd8dbdcd3b8bce748c7c275f239adf7

Observation 76b552cb-916f-41ca-8550-ba8edc57693f · inbound

Fourier Compressor: Frequency-Domain Visual Token Compression for Vision-Language Models cites this paper.

Fourier Compressor: Frequency-Domain Visual Token Compression for Vision-Language Models LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 20

Resolution
metadata mismatch
local_arxiv, observed 2026-05-21T22:40:43.085060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T22:40:39.892802Z digest=sha256:bb8d7cc093f05844206c4d6668d586453363c2f91d8e2c6c2400ac8d6a1b987c

Observation 63e2b82f-acc5-4c00-8289-ba68f89b3760 · inbound

Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs cites this paper.

Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T13:27:58.447481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:27:58.447481Z digest=sha256:b3733de6b08b6265239041d7c68f6ddd1d555ec805837b26935ae53415f05023

Observation bf9bcb6f-9776-4003-bce1-4caf79ab00a9 · inbound

CARES: Context-Aware Resolution Selector for VLMs cites this paper.

CARES: Context-Aware Resolution Selector for VLMs LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T08:44:42.346164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:44:42.346164Z digest=sha256:5c338c16e76b4fefe154b98edd6e0c646f79701137074fbedd30a12a07e32611

Observation 28787986-459a-4880-a23e-3f02b73298fa · inbound

Beyond Reasoning Gains: Mitigating General-Capability Forgetting in Large Reasoning Models cites this paper.

Beyond Reasoning Gains: Mitigating General-Capability Forgetting in Large Reasoning Models LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 113

Resolution
unresolved
no resolver link, observed 2026-08-04T08:15:59.137719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:15:59.137719Z digest=sha256:29a7a4495ba1a296966a50b9bf6c0bd32fd661e49b44dc97677fb482667fa9d1

Observation c767cd13-4442-4d46-b57d-9660c9126f67 · inbound

Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously cites this paper.

Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-02T18:23:40.241140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:23:40.241140Z digest=sha256:d7272e86e19ec0e7c2705b808bd0b49750285e7ba810768ff4902a22a214bd2c

Observation 7105cd42-cab5-4b37-be49-c3e4845aee6e · inbound

Saliency-R1: Enforcing Interpretable and Faithful Vision-language Reasoning via Saliency-map Alignment Reward cites this paper.

Saliency-R1: Enforcing Interpretable and Faithful Vision-language Reasoning via Saliency-map Alignment Reward LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 89

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:19:22.513527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T19:59:19.379119Z digest=sha256:857a1eb7b169a94fc2d4f8dafb412fb12e034327a92d840985ca1d5266ae9aba

Observation a8e958b3-6d7f-43fb-b5f6-152514e0a3a3 · inbound

Q-Zoom: Query-Aware Adaptive Perception for Efficient Multimodal Large Language Models cites this paper.

Q-Zoom: Query-Aware Adaptive Perception for Efficient Multimodal Large Language Models LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:19:22.513527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T18:46:26.869644Z digest=sha256:f1887412f9c039c9fbe218e6c693fa0a6b28a7b6834b2d71dd59e8f9ff86c4ff

Observation 021f215b-5ac0-45ff-8bbe-e86e3050de45 · inbound

POINTS-Long: Adaptive Dual-Mode Visual Reasoning in MLLMs cites this paper.

POINTS-Long: Adaptive Dual-Mode Visual Reasoning in MLLMs LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 113

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:19:22.513527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T15:23:08.671342Z digest=sha256:9a023bfc817be4d7417ad776c92556e9ecfcbd1e3e549a6900eeb6c268995fb9

Observation 2b15b9c2-f787-4df8-89e7-779273ebab8d · inbound

Latent Denoising Improves Visual Alignment in Large Multimodal Models cites this paper.

Latent Denoising Improves Visual Alignment in Large Multimodal Models LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 98

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T05:19:22.513527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T23:07:54.806529Z digest=sha256:7fdadd7cb2ec3ef7e36b854eb36da8e40f0fd45689778523cb94237cc2f32379

Observation 26688c78-7c09-4b62-989d-fbb5ea9ea4f3 · inbound

Make Your LVLM KV Cache More Lightweight cites this paper.

Make Your LVLM KV Cache More Lightweight LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 104

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:19:22.513527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-09T19:04:48.058450Z digest=sha256:9d1bbbfab077e38bf128fb205eb5140e112f2a584140fbbf08e20ccaf6e33fcf

Observation bf4e9218-4dad-4245-9199-d5f63af42156 · inbound

MACS: Modality-Aware Capacity Scaling for Efficient Multimodal MoE Inference cites this paper.

MACS: Modality-Aware Capacity Scaling for Efficient Multimodal MoE Inference LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T05:19:22.513527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-11T01:29:26.298131Z digest=sha256:de4d651b128a8e249dde5d4106144522e910ff603eb262ae4b128163a294aae9

Observation cf12b96f-50c3-4149-9f93-35282714eba7 · inbound

TTF: Temporal Token Fusion for Efficient Video-Language Model cites this paper.

TTF: Temporal Token Fusion for Efficient Video-Language Model LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:19:22.513527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-11T01:55:32.319344Z digest=sha256:1555d6d267c4fee41bac0afbe3a0cec6c4752bf6a89dcacc4c90467feb92432e

Observation b425d45b-25d9-4bef-9270-e7c6144f21a2 · inbound

Make Each Token Count: Towards Improving Long-Context Performance with KV Cache Eviction cites this paper.

Make Each Token Count: Towards Improving Long-Context Performance with KV Cache Eviction LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:19:22.513527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T05:02:25.513351Z digest=sha256:22495bdf3bb580fb3c10ec0d4ca910e3f66434f9afdb1c00ccf590c3fbd217aa

Observation 2e62e037-cf04-4ccb-a348-883acdead27f · inbound

LDDR: Linear-DPP-Based Dynamic-Resolution Frame Sampling for Video MLLMs cites this paper.

LDDR: Linear-DPP-Based Dynamic-Resolution Frame Sampling for Video MLLMs LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:19:22.513527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T03:20:42.684286Z digest=sha256:0c9545ee6ceca08e6c7cb14fe481eb4198729bdc7e043626142dab072ce5a5aa

Observation d05303f5-c25a-488e-a908-46996715e73e · inbound

Mitigating Mask Prior Drift and Positional Attention Collapse in Large Diffusion Vision-Language Models cites this paper.

Mitigating Mask Prior Drift and Positional Attention Collapse in Large Diffusion Vision-Language Models LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T05:19:22.513527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T02:14:59.487486Z digest=sha256:013da105589b2812443dc2a47897f09557a5451092807c99d180f985e65dd494

Observation e017fd1d-6aff-4637-954b-26390dba5d66 · inbound

Mitigating Mask Prior Drift and Positional Attention Collapse in Large Diffusion Vision-Language Models cites this paper.

Mitigating Mask Prior Drift and Positional Attention Collapse in Large Diffusion Vision-Language Models LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 24

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T21:49:05.337919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T21:47:32.112057Z digest=sha256:cf57d69ae791c51cb5d91710d737b870db88089c63e65797af59338b54e2c619

Observation 7988e842-18e9-4f63-b1ae-ed6ec22bf443 · inbound

HEED: Density-Weighted Residual Alignment for Hybrid Vision-Language Model Distillation cites this paper.

HEED: Density-Weighted Residual Alignment for Hybrid Vision-Language Model Distillation LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-05-20T15:38:26.650745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T15:37:09.899037Z digest=sha256:c0d900b122e83c831041e1f96f1530e4d836c1090bb1c164672b77fd14c4058a

Observation 9027b91b-4234-4f59-8c08-7ef4be21959f · inbound

MGVQ: Synergizing Multi-dimensional Sensitivity-Aware and Gradient-Hessian Fusion for Vector Quantization cites this paper.

MGVQ: Synergizing Multi-dimensional Sensitivity-Aware and Gradient-Hessian Fusion for Vector Quantization LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 16

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T15:05:48.308600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T17:36:45.807397Z digest=sha256:8b9eac057337901a6acea66b91ce91a35fc439953b0240424187f92c40bc8c20

Observation 7554c83d-76c5-4e2f-b847-a6e4214fc8af · inbound

EarlyTom: Early Token Compression Completes Fast Video Understanding cites this paper.

EarlyTom: Early Token Compression Completes Fast Video Understanding LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-06-29T08:23:15.627356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T08:16:02.536341Z digest=sha256:c601b8b9253541b3f1b054e69f85b203a8c7614b93428f13b5a86effe8223636

Observation b0d9029e-a6ab-4bb5-8edb-ee27f65f0453 · inbound

Constitutional On-Policy Safe Distillation cites this paper.

Constitutional On-Policy Safe Distillation LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-07-02T01:36:25.475965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T11:47:14.135793Z digest=sha256:e8e2141944c63ea720db351f4116f3c8e443234237e4337100253a31a29f6d56

Observation b0c5ba65-a2a4-4f37-85f6-555d4489fb76 · inbound

Benchmarking Visual State Tracking in Multimodal Video Understanding cites this paper.

Benchmarking Visual State Tracking in Multimodal Video Understanding LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 62

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:46:27.940892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T10:43:06.811228Z digest=sha256:0acbc2f291131818675dcbcb9a9a9f91371e636cb8c8be267bed0687782bd932

Observation 25dea9a0-2b8d-43d4-bc26-20d6e5324640 · inbound

Closed-Form Spectral Regularization for Multi-Task Model Merging cites this paper.

Closed-Form Spectral Regularization for Multi-Task Model Merging LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-07-02T16:27:09.395340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T22:40:00.510742Z digest=sha256:f2f81c8ee35af7c7c4a50e69be70c6e9589b1dcd6eb195814eb49578b50e725c

Observation f58c7104-7043-41d5-b609-a74fcb6bb294 · inbound

AlloSpatial: Agentic Harness Framework for Spatial Reasoning in Foundation Models cites this paper.

AlloSpatial: Agentic Harness Framework for Spatial Reasoning in Foundation Models LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-07-03T00:27:29.711529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T17:08:22.609095Z digest=sha256:0a15f8bfbeb4422465d944a61237d9f015cee2b6d73a154bddbafe2ba544a280

Observation 513f496d-5795-4876-94d8-c4824a7a44ef · inbound

Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models cites this paper.

Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 110

Resolution
unresolved
no resolver link, observed 2026-07-14T18:07:09.018997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T18:07:09.018997Z digest=sha256:996739580fb2cefa3ee0c5aeca3202d4d2bd50765fe848867203f685e08e2f57

Observation 3b30af51-9153-4c97-b9f3-d65627d76d8b · inbound

Natural-Language Temporal Grounding in Hour-Long Videos is a Search Problem: A Benchmark and Empirical Decomposition cites this paper.

Natural-Language Temporal Grounding in Hour-Long Videos is a Search Problem: A Benchmark and Empirical Decomposition LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-07-03T10:27:56.654197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-27T10:01:09.933880Z digest=sha256:27640a7651110d67f973958abfa9ae01cdd1421ad327e5d3ed53638e71e7dfa4

Observation d84ed359-c4f1-4ecf-b189-b9f9f51441fe · inbound

DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack cites this paper.

DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:18:56.562509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T01:28:59.297634Z digest=sha256:7a8ad46896aa58f75a017eb9cfe726e19d9521f44a9a1bc156c37b4d6f0515ad

Observation ea526200-d7d3-4e92-ab09-a14509d6a1b8 · inbound

Continuous Audio Thinking for Large Audio Language Models cites this paper.

Continuous Audio Thinking for Large Audio Language Models LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-07-02T17:47:18.020004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T21:52:49.839901Z digest=sha256:44e8d7b9a2b1b28239fcdda0275679c5b4df00e24f1eae047d6c0e9c15a6829d

Observation 9dade975-ceed-40d6-bbad-1505b43310fe · inbound

Spectral Query-Key Product Weight Steering for Training-Free VLM Hallucination Mitigation cites this paper.

Spectral Query-Key Product Weight Steering for Training-Free VLM Hallucination Mitigation LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 20

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T03:19:30.594837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-26T18:14:49.763342Z digest=sha256:e34096e896a9b8a192f1eba15ed00b45338e69a91a7869449c5432c2a8c00502

Observation f951c9da-9086-4aab-ad72-8ca7c16096a0 · inbound

Combating Textual Noise and Redundancy: Entropy-Aware Dense Visual Token Pruning cites this paper.

Combating Textual Noise and Redundancy: Entropy-Aware Dense Visual Token Pruning LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 71

Resolution
verified exact
local_arxiv, observed 2026-07-03T14:48:32.411859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-03T14:47:35.377391Z digest=sha256:9e788ab4f330cff683f235dae22229878db51eb8a09f392b10559e24b565269a

Observation cc30a3b0-f8c9-4a67-a617-7fb151d37661 · inbound

MentalThink: Shaping Thoughts in Mental SVG World cites this paper.

MentalThink: Shaping Thoughts in Mental SVG World LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 163

Resolution
unresolved
no resolver link, observed 2026-07-12T01:50:59.184754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T01:50:59.184754Z digest=sha256:3cf953b08f4d161f6b57ea1e49592bf109f00c6c87ce6187c0b9f796f34dda52

Observation a9d82b80-c7ac-41c8-8d1e-595ae626c567 · inbound

TimeThink: Reasoning with Time for Video LLMs cites this paper.

TimeThink: Reasoning with Time for Video LLMs LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 64

Resolution
unresolved
no resolver link, observed 2026-07-11T08:59:46.244502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T08:59:46.244502Z digest=sha256:fb4c8b0208fdad4943196db8802862084f318c552fc0ad8085ffe28c52643cd5

Observation ff51fd54-0a23-4d99-b4b1-10d00fe0095c · inbound

SigLIP-HD by Fine-to-Coarse Supervision cites this paper.

SigLIP-HD by Fine-to-Coarse Supervision LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-07-13T02:39:58.984385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T02:39:58.984385Z digest=sha256:c79fee10c5d229da7fa0bb26e33544bcea6a8ec2791b9d9aa49ae5818fd3d53d

Observation 5c055d66-f9a4-428b-9b63-1a7351e0fd91 · inbound

MIRROR: Learning from the Other View for Multi-Modal Reasoning cites this paper.

MIRROR: Learning from the Other View for Multi-Modal Reasoning LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-01T07:15:04.870137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:15:04.870137Z digest=sha256:8de4cf5f38bbec023a3c2f1d4959eb296ae0a8fb4860588233751a42b0e00f00

Observation 40bbf5e9-7530-4ceb-950b-f498f214554d · inbound

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models cites this paper.

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-02T11:58:31.210800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:58:31.210800Z digest=sha256:d9716a4182ff52d93851600f66e93578763baad14b90b31d79eed6e81268027d

Observation 6b8813a0-05ca-4779-8278-e01d2d41057c · inbound

Allocation Before Ranking: Decoupled Token Compression for OmniLLMs cites this paper.

Allocation Before Ranking: Decoupled Token Compression for OmniLLMs LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T23:23:38.908542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:23:38.908542Z digest=sha256:700859416ffeb18476b869c5547387533b5d13b3b1e417af42640a5e8eca5228