Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-17T05:19:22.423762Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 76 inbound Pith citation observations for arXiv:2407.12772.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-17T05:19:22.423762Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-10T21:34:35.657096Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
24 of 24 outbound references displayed
External citation measurements
1
pith, observed 2026-08-05T02:28:24.338817Z
Observation 41a065a0-33ba-4505-96d7-c12348c473d8 · outbound
LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ac237b95-9dcc-430d-8a2e-a64a6f60f2d4 · outbound
LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4eaf4965-7af9-469c-bf9b-d742213b2150 · outbound
LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 46622577-559c-4ea9-afda-31666510f829 · outbound
LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models Making LLaMA SEE and Draw with SEED Tokenizer
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 61f9941a-f4a7-4bd6-99a1-144fd10c0c72 · outbound
LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models A Diagram Is Worth A Dozen Images
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation da6a5df2-38fc-4d4c-9498-99235737c321 · outbound
LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models Coresets for Data-efficient Training of Machine Learning Models
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a858ae38-1d42-40f8-a4b0-022b15c8c58e · outbound
LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models tinyBenchmarks: evaluating LLMs with fewer examples
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 99525b92-fa84-4b7b-b28d-7c7188b2ea32 · outbound
LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models What are the key points in this news story?
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f0f44743-f5e5-43c0-8f27-41cc25480f27 · outbound
LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models What are the factors that led to this event?
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3a96d9ad-e6f5-4fc5-b812-2740fede8ed1 · outbound
LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models How could you create a new headline that captures the essence of the event differently?
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 93846a10-9fcd-4507-87b2-4941b3b630a5 · outbound
LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models Please present this news in Arabic and output it in markdown format
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c2f125d4-5c36-471f-a3ce-7f7e6e922df5 · outbound
LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models Unresolved cited work
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 407d8bbd-82cb-46f0-b4c0-b6721acb82b8 · outbound
LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models However, you should not change the original question's subtask unless the original subtask is not one of these five
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5fd5be42-25f9-49ad-9ba7-988414396e6c · outbound
LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models Unresolved cited work
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 97641f0d-0d6d-48e0-9837-50841ceb6d36 · outbound
LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models But don't use python−like format
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b1524ccc-abb8-46ab-90db-695baa3ae091 · outbound
LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models Unresolved cited work
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9c4ebe4c-e8df-49a9-9a55-24e8c8704085 · outbound
LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models But if you think some words should be in other language, you can keep it in that language
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation afde4456-461d-4bb1-bd93-9931c860ad57 · outbound
LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models Unresolved cited work
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 38f1f065-e613-4cca-8aef-84769a6ff0c5 · outbound
LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models Unresolved cited work
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 085a2ac6-cff5-4219-98d2-f4d184ffc838 · outbound
LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models Unresolved cited work
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation bda072aa-ad67-4127-b2c9-a50814044507 · outbound
LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models Some tips
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e406d34c-a761-4cdd-82ff-faf783f8f100 · outbound
LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models In such cases, you can relax the criteria slightly
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6762720f-b6fc-4931-9d4f-e91c876f3a3c · outbound
LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models Explanation
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 68a4ab78-6909-40d3-9419-bd66816e0b28 · outbound
LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models This kind of systemic shift often results in skills becoming obsolete, leading to higher unemployment among professionals who cannot quickly adapt to new technological paradigms
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation edaf3815-d97d-40ae-83b6-143507c5c40f · inbound
LLaVA-OneVision: Easy Visual Task Transfer LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Reference 161
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation fcf13dd4-e536-4b33-b2be-418a60ae91cd · inbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Reference 165
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7f8d7f30-ef47-4ce8-b292-ded9b4ed6754 · inbound
Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 31d738d5-6381-4128-8300-1446b21f4267 · inbound
Feedback-Driven Vision-Language Alignment with Minimal Human Supervision LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Reference 113
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86ea2dec-3ff0-4369-8491-984c2e3211bb · inbound
Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8525a6ab-88cd-4869-90a4-a9bd1a23a7bd · inbound
Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5b9854d1-d8a6-4e5e-b9a3-3af0b15ddcbb · inbound
Temporal Preference Optimization for Long-Form Video Understanding LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22a8c01b-60d0-4d79-9f99-e02db2c23309 · inbound
Mordal: Automated Pretrained Model Selection for Vision Language Models LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6dde0cbd-6c00-4a84-a005-0ae1a1f197c0 · inbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Reference 100
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ace4d7a1-6410-4a39-b4fc-15bec714d7f4 · inbound
From Visuals to Vocabulary: Establishing Equivalence Between Image and Text Token Through Autoregressive Pre-training in MLLMs LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1624128-ec7a-482d-8661-7f01e487293e · inbound
Granite Vision: a lightweight, open-source multimodal model for enterprise Intelligence LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Reference 103
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9959d856-1575-44aa-aef0-cd26b4beca69 · inbound
LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation cdfe6a61-b8c6-4b66-8526-ef453138e929 · inbound
Clapper: Compact Learning and Video Representation in VLMs LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1a68158-024f-4428-8b46-b41501f1b8a0 · inbound
QuickVideo: Real-Time Long Video Understanding with System Algorithm Co-Design LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 541fab43-345b-4c5e-af52-0c0dcbfe5899 · inbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Reference 112
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 665b20c0-642a-4749-adce-bb09dbca8641 · inbound
Inference Compute-Optimal Video Vision Language Models LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94ea1965-e3a7-4fd4-9eec-5b6d1d628f21 · inbound
AdaTP: Attention-Debiased Token Pruning for Video Large Language Models LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4a2bfdf-d0be-4e4b-bf84-6059ee7d55e2 · inbound
ID-Align: RoPE-Conscious Position Remapping for Dynamic High-Resolution Adaptation in Vision-Language Models LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 154b2e67-dfd9-49d3-8247-306637013fb2 · inbound
Fostering Video Reasoning via Next-Event Prediction LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1de9b23-f463-4c5c-a580-d41614ee40de · inbound
Spoken question answering for visual queries LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb348a27-c885-4a0a-9667-9fb0ed96adda · inbound
Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 743b0234-2b6f-4c89-a45c-ca3607796007 · inbound
MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f780789c-d041-484f-bd98-f944e3e6b659 · inbound
EffiVLM-BENCH: A Comprehensive Benchmark for Evaluating Training-Free Acceleration in Large Vision-Language Models LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8826dfff-bcbe-4a51-859c-962a6a42fd2b · inbound
FlexSelect: Flexible Token Selection for Efficient Long Video Understanding LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 399f54ee-1809-460e-a0d9-df04827b9876 · inbound
MiMo-VL Technical Report LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29db3ad9-d287-4e59-8754-7cbb96c444d0 · inbound
DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b5cad25-5993-4d34-a0d1-5e6058d88cee · inbound
VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in Videos LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa89350e-5ea6-4487-b41b-6fccf2ed5fd6 · inbound
SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e976c28f-cb3f-4be5-afc5-2c4129af7882 · inbound
Vision-EKIPL: External Knowledge-Infused Policy Learning for Visual Reasoning LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0b2d9593-f06e-45e2-a843-0a5d394ef7e1 · inbound
FlagEvalMM: A Flexible Framework for Comprehensive Multimodal Model Evaluation LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15bd45bc-a1b8-4c77-8dd8-83626e44a128 · inbound
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1b0053e0-2ad7-40d6-9e37-341dea810c36 · inbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 842f749f-59b3-4d5a-96e5-45f2ff031e20 · inbound
VGR: Visual Grounded Reasoning LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1805c7fa-a03b-4a40-92f1-0c7d95361a3f · inbound
GreedyPrune: Retenting Critical Visual Token Set for Large Vision Language Models LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed01f621-792c-4699-afb1-f54e0c1f7f6d · inbound
LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ca77f38-1b04-454e-b51f-3d1e2a2bf5e2 · inbound
MMSearch-R1: Incentivizing LMMs to Search LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5d985fbf-4b73-481f-b255-2b737e68829e · inbound
Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94ef0cc7-fe5a-4dab-97df-7275881dcd82 · inbound
VisionThink: Smart and Efficient Vision Language Model via Reinforcement Learning LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20eedbe1-69b8-46a6-adc3-00dd90e4df93 · inbound
EgoPrune: Efficient Token Pruning for Egomotion Video Reasoning in Embodied Agent LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99f4e4f6-d6ea-447e-b0c2-7f198124978f · inbound
Object-centric Video Question Answering with Visual Grounding and Referring LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 932c4614-d28b-4ef2-a0e8-434d0e20d706 · inbound
A Glimpse to Compress: Dynamic Visual Token Pruning for Large Vision-Language Models LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76b552cb-916f-41ca-8550-ba8edc57693f · inbound
Fourier Compressor: Frequency-Domain Visual Token Compression for Vision-Language Models LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 63e2b82f-acc5-4c00-8289-ba68f89b3760 · inbound
Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf9bcb6f-9776-4003-bce1-4caf79ab00a9 · inbound
CARES: Context-Aware Resolution Selector for VLMs LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28787986-459a-4880-a23e-3f02b73298fa · inbound
Beyond Reasoning Gains: Mitigating General-Capability Forgetting in Large Reasoning Models LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Reference 113
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c767cd13-4442-4d46-b57d-9660c9126f67 · inbound
Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7105cd42-cab5-4b37-be49-c3e4845aee6e · inbound
Saliency-R1: Enforcing Interpretable and Faithful Vision-language Reasoning via Saliency-map Alignment Reward LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a8e958b3-6d7f-43fb-b5f6-152514e0a3a3 · inbound
Q-Zoom: Query-Aware Adaptive Perception for Efficient Multimodal Large Language Models LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 021f215b-5ac0-45ff-8bbe-e86e3050de45 · inbound
POINTS-Long: Adaptive Dual-Mode Visual Reasoning in MLLMs LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Reference 113
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2b15b9c2-f787-4df8-89e7-779273ebab8d · inbound
Latent Denoising Improves Visual Alignment in Large Multimodal Models LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Reference 98
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 26688c78-7c09-4b62-989d-fbb5ea9ea4f3 · inbound
Make Your LVLM KV Cache More Lightweight LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Reference 104
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation bf4e9218-4dad-4245-9199-d5f63af42156 · inbound
MACS: Modality-Aware Capacity Scaling for Efficient Multimodal MoE Inference LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation cf12b96f-50c3-4149-9f93-35282714eba7 · inbound
TTF: Temporal Token Fusion for Efficient Video-Language Model LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b425d45b-25d9-4bef-9270-e7c6144f21a2 · inbound
Make Each Token Count: Towards Improving Long-Context Performance with KV Cache Eviction LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2e62e037-cf04-4ccb-a348-883acdead27f · inbound
LDDR: Linear-DPP-Based Dynamic-Resolution Frame Sampling for Video MLLMs LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d05303f5-c25a-488e-a908-46996715e73e · inbound
Mitigating Mask Prior Drift and Positional Attention Collapse in Large Diffusion Vision-Language Models LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e017fd1d-6aff-4637-954b-26390dba5d66 · inbound
Mitigating Mask Prior Drift and Positional Attention Collapse in Large Diffusion Vision-Language Models LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7988e842-18e9-4f63-b1ae-ed6ec22bf443 · inbound
HEED: Density-Weighted Residual Alignment for Hybrid Vision-Language Model Distillation LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9027b91b-4234-4f59-8c08-7ef4be21959f · inbound
MGVQ: Synergizing Multi-dimensional Sensitivity-Aware and Gradient-Hessian Fusion for Vector Quantization LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7554c83d-76c5-4e2f-b847-a6e4214fc8af · inbound
EarlyTom: Early Token Compression Completes Fast Video Understanding LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b0d9029e-a6ab-4bb5-8edb-ee27f65f0453 · inbound
Constitutional On-Policy Safe Distillation LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b0c5ba65-a2a4-4f37-85f6-555d4489fb76 · inbound
Benchmarking Visual State Tracking in Multimodal Video Understanding LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 25dea9a0-2b8d-43d4-bc26-20d6e5324640 · inbound
Closed-Form Spectral Regularization for Multi-Task Model Merging LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f58c7104-7043-41d5-b609-a74fcb6bb294 · inbound
AlloSpatial: Agentic Harness Framework for Spatial Reasoning in Foundation Models LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 513f496d-5795-4876-94d8-c4824a7a44ef · inbound
Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Reference 110
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b30af51-9153-4c97-b9f3-d65627d76d8b · inbound
Natural-Language Temporal Grounding in Hour-Long Videos is a Search Problem: A Benchmark and Empirical Decomposition LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d84ed359-c4f1-4ecf-b189-b9f9f51441fe · inbound
DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ea526200-d7d3-4e92-ab09-a14509d6a1b8 · inbound
Continuous Audio Thinking for Large Audio Language Models LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9dade975-ceed-40d6-bbad-1505b43310fe · inbound
Spectral Query-Key Product Weight Steering for Training-Free VLM Hallucination Mitigation LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f951c9da-9086-4aab-ad72-8ca7c16096a0 · inbound
Combating Textual Noise and Redundancy: Entropy-Aware Dense Visual Token Pruning LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation cc30a3b0-f8c9-4a67-a617-7fb151d37661 · inbound
MentalThink: Shaping Thoughts in Mental SVG World LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Reference 163
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9d82b80-c7ac-41c8-8d1e-595ae626c567 · inbound
TimeThink: Reasoning with Time for Video LLMs LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff51fd54-0a23-4d99-b4b1-10d00fe0095c · inbound
SigLIP-HD by Fine-to-Coarse Supervision LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c055d66-f9a4-428b-9b63-1a7351e0fd91 · inbound
MIRROR: Learning from the Other View for Multi-Modal Reasoning LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40bbf5e9-7530-4ceb-950b-f498f214554d · inbound
MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b8813a0-05ca-4779-8278-e01d2d41057c · inbound
Allocation Before Ranking: Decoupled Token Compression for OmniLLMs LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.