Pith. sign in

Paper Citation Record · LEDGER

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

As of 11 August 2026, this Paper Citation Record lists 100 of 300 outbound references and 100 inbound Pith citation observations for arXiv:2412.05271.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.05271 v5

Coverage vector

measured 100 of 300 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T13:23:57.588851Z

measured 200 of 200 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 100 of 585 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T23:07:14.402030Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

100 of 300 outbound references displayed

  • verified exact46
  • verified fuzzy53
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

14
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation f6c45375-722c-4cfa-ac61-f9e4c0d07f36 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-10T20:19:27.408543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:6a31d62c0ba3fb257511d2ba0798c4f03b6df070b3e47d4008634f06bb1e7e10

Observation 354b7464-ff37-4753-ad03-5c415cbf0441 · outbound

This paper cites Tallyqa: Answering complex counting questions.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Tallyqa: Answering complex counting questions

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T13:23:58.360590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:964553e90469302203274b2867d56794c9a52559b227915de0a222db23386070

Observation 93d80056-95b1-492c-94ad-be8acfa3505f · outbound

This paper cites GPT-4 Technical Report.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling GPT-4 Technical Report

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-10T13:23:58.304469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:2c10bd6fcdde33454b99791fdf0e633a653050f6a89244f5246730db4d19f0f3

Observation 9fba97ac-8164-4ec5-ab38-540fb7d35e16 · outbound

This paper cites Wave-ui.https://huggingface.co/datasets/agentsea/wave-ui.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Wave-ui.https://huggingface.co/datasets/agentsea/wave-ui

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T13:23:58.363040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:9c20d9b055f07c39b319ff5ecc211dc01d68bcd40f36280a1bac115c25a37616

Observation 323706d6-fa24-480b-ac17-bda0a7906393 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Flamingo: a visual language model for few-shot learning

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T13:23:58.365608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:268fa7d41408f3453bcc4ff7b8795f639c5a58fb22149bde1998b9c5c40f01cb

Observation 44c6580b-d82f-468d-8fd1-3c31b82425bb · outbound

This paper cites MathQA: Towards Interpretable Math Word Problem Solving with Operation-Based Formalisms.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling MathQA: Towards Interpretable Math Word Problem Solving with Operation-Based Formalisms

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:23:57.979866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:565b867dfef75c4b591227395fb41ba075261d2b137d57bfa8cccf6bb0521f82

Observation e0e95554-2034-4cac-9471-b5bcb336e11c · outbound

This paper cites CG-bench: Clue-grounded question answering benchmark for long video understanding.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling CG-bench: Clue-grounded question answering benchmark for long video understanding

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T13:23:58.368059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:f95ef82d1f6ad29fa59b842b724e8ab25eebf2e4d5c9f8cab4e6830a35e1c162

Observation 46d8af48-b448-47f6-9c96-adc88e7ff32c · outbound

This paper cites The claude 3 model family: Opus, sonnet, haiku.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling The claude 3 model family: Opus, sonnet, haiku

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T13:23:58.370526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:d934aeefaafe21a07866a1218b4df18ffa60745a38fef14efed1c894607b7a9c

Observation d8f74ace-34ed-467e-aa61-059efcbdbd07 · outbound

This paper cites Program Synthesis with Large Language Models.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Program Synthesis with Large Language Models

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-10T13:23:57.990455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:c1b7e615a8890f60ba210af89deb2ed12804db1fc18ed6520bb7d16ad49d421f

Observation 624abb5d-089f-4e24-be61-b1e657ff8799 · outbound

This paper cites Scanqa: 3d question answering for spatial scene understanding.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Scanqa: 3d question answering for spatial scene understanding

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T13:23:58.373458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:7b1436a1d30e6c1ba2a4da7741544486cb98841b9b709a772f3bae424e252953

Observation fa35076e-1a37-480f-a103-61cddbd93720 · outbound

This paper cites Layer Normalization.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Layer Normalization

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-10T13:23:58.350691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:8992ab054830f1a6a501641ef7bef490cabe92ca9f52856bbc67a8422affe6db

Observation 4714b379-b938-4275-8412-a754b91415ce · outbound

This paper cites UIBert: Learning Generic Multimodal Representations for UI Understanding.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling UIBert: Learning Generic Multimodal Representations for UI Understanding

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:23:57.822404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:6e689c917f1fe97637b27ae4f0fae58ac5beeaef09eab035f4daa26398e52076

Observation 09e339bd-a6fd-402b-8182-6e3b6b40d1ff · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-10T13:23:57.826373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:a52fd69aa7cd974ead371e8204ac0782b88a740b0ca358125744004f3c7bac6c

Observation 8e6f0162-5275-409f-b560-26a0c7f33847 · outbound

This paper cites Beijing Anjie Zhihe Technology Co.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Beijing Anjie Zhihe Technology Co

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T13:23:58.375954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:f391730b4d34e4116a1b85346623f02bce8b02bdac145aec020c8de3e316a94f

Observation 7f6d035e-ec60-4786-9d2c-f490c62ea316 · outbound

This paper cites Vqa-med: Overview of the medical visual question answering task at imageclef 2019.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Vqa-med: Overview of the medical visual question answering task at imageclef 2019

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T13:23:58.378510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:33cd7a631c3da28d6e92fe49e007eda67da83c2524a28c1cc8bf3a8b7766877b

Observation ea95d6a2-7514-4ea7-a452-c8045913625b · outbound

This paper cites Are we done with ImageNet?.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Are we done with ImageNet?

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:23:58.188150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:c8405e20bd54eceb9247e44e4a7855aa76a03bd63f4fb25543fbe09c98939c7b

Observation 0bb7bda7-fd34-4aba-becf-74f2a1b0f875 · outbound

This paper cites Scene text visual question answering.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Scene text visual question answering

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T13:23:58.381204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:725fffcf28b7bee618903f08c8d6611769cd2d8b354fe5707d0306bb9aa45206

Observation 7121e350-55b6-465f-8688-93b65c912694 · outbound

This paper cites Coco-stuff: Thing and stuff classes in context.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Coco-stuff: Thing and stuff classes in context

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T13:23:58.383753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:037528a20d09bca806d01edc9280ea5e5ea7b3535627f30819f344e2b8f9cd6e

Observation e101ffbc-5196-4ad1-923a-711163632721 · outbound

This paper cites InternLM2 Technical Report.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling InternLM2 Technical Report

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-15T11:44:38.670888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:effd7b03d28937402a46ea3347ece90ee73ac2cfd3d1c090e220ab28a645009d

Observation a4363e8f-6050-4b32-8e09-e2d4012c3416 · outbound

This paper cites An augmented benchmark dataset for geometric question answering through dual parallel text encoding.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling An augmented benchmark dataset for geometric question answering through dual parallel text encoding

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T13:23:58.386224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:39bc53297d20d540c77f268385d05972de39cd17fbf7c957fedddb15d5dac02f

Observation dd3ab7b6-9103-480c-94e6-fa775efba9b0 · outbound

This paper cites openai summarize tldr dataset.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling openai summarize tldr dataset

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T13:23:58.388534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:0ba627333bfc6c784e07719a0d0a7f5404b28c554bcf5af36958215239d8e00b

Observation 6f6b7ac7-e5ca-4054-b417-548da7b31f1f · outbound

This paper cites AMEX: Android Multi-annotation Expo Dataset for Mobile GUI Agents.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling AMEX: Android Multi-annotation Expo Dataset for Mobile GUI Agents

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:23:57.830230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:385f04e9e4ecfd39bfe9e47684e5470f59a26881ed1ea2efb89a9b4c29f3cb6e

Observation fb483837-30a6-4fa4-9425-392ee7ade3e7 · outbound

This paper cites MapQA: A Dataset for Question Answering on Choropleth Maps.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling MapQA: A Dataset for Question Answering on Choropleth Maps

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:23:57.834590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:4701a9a1b7d075fa932a2d5e8c0afefe1dda3a91bcd38e0c233e01458102b0c5

Observation c53053c5-01e5-4ccf-8047-9d75d2bca10e · outbound

This paper cites GUI-World: A Video Benchmark and Dataset for Multimodal GUI-oriented Understanding.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling GUI-World: A Video Benchmark and Dataset for Multimodal GUI-oriented Understanding

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:23:57.890681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:417a01a286d11ad158bc6c01e93935fa14d46e45f2b0090d43edc13480573a1b

Observation 60562b2e-a27c-4d56-b8c2-6b73887b1e0f · outbound

This paper cites ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-23T22:20:22.000193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:0fa65311002f902c896767b6862241f8091c113b8b17c752612f0485fc1b9082

Observation 659889a3-271d-4631-9648-860b79fec7d1 · outbound

This paper cites UniGeo: Unifying Geometry Logical Reasoning via Reformulating Mathematical Expression.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling UniGeo: Unifying Geometry Logical Reasoning via Reformulating Mathematical Expression

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:23:57.903870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:0cfb76e4996c5fcebce9213af0412701cda6faa3f034a2a2300c85c1fac359ab

Observation 127d7c43-50e8-49af-bbac-9ac9e201ae31 · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-12T15:52:36.167809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:530d7514aba5c0ad2425bb303491c95802d25478c503385894bcdfa1cd180091

Observation 1f6485f8-3e12-4f5e-97e5-41ac988bff8e · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-12T19:41:44.612219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:e1669de21d1a762d917b6d8b8b6c7a97505ed56b2ad2f9c81bcaf1e32902247c

Observation f3a0cecc-bea8-4593-a9bb-c07b62f3935f · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-13T17:08:13.182567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:d0d7a4078475b3fb8d8f74eb8f60b90403986fcb48cdbac43d5a08349f7b9358

Observation 235ff406-5120-49e3-9bd1-921fedb6817a · outbound

This paper cites ShareGPT4Video: Improving Video Understanding and Generation with Better Captions.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:23:57.967610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:66fc138860a967bc5388c89132cecfe5efd57576d631cd4ae712cc409ce0f191

Observation 7993ecb8-ec6e-47c0-b96e-09826a1b3e5c · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Evaluating Large Language Models Trained on Code

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-10T13:23:57.971121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:cefd230e96aea500c94331224a8916c053b0dd4e580782ec1dc3671b80b5cb83

Observation 22fcdd29-2720-4847-99da-eb49e1a98a7b · outbound

This paper cites A simple framework for contrastive learning of visual representations.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling A simple framework for contrastive learning of visual representations

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T13:23:58.391100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:c0aa840d558a789aa2c7d0e83d9f9890ef0aed1f70d97e9194446fea0288bc36

Observation ee10c80f-3f0a-4ca0-b821-35c3987f5ae8 · outbound

This paper cites Theoremqa: A theorem-driven question answering dataset.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Theoremqa: A theorem-driven question answering dataset

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T13:23:58.393419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:d6ebb2396a3952609706b394e53c712afd0ee00e698c6337a028c2edad7adbbe

Observation 8e3386e5-9654-44b0-99f9-8f203cd8f7e5 · outbound

This paper cites LongLoRA: Efficient Fine-tuning of Long-Context Large Language Models.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling LongLoRA: Efficient Fine-tuning of Long-Context Large Language Models

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:23:57.999994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:74ac6b3e23f9cde7569bfe295fa820403296b2a0afd68f341cb110cd8ade1047

Observation 3d589368-0d9b-4329-bab3-856bd93828e0 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-12T20:58:59.653229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:3678e7c2ea2b012435ab0fa0c29b16366672655b991c394033821f896f05b177

Observation 32d5ffd0-f0f2-439d-ba2e-b0117d08fd83 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T13:23:58.395678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:bd0bdcc592b04d21f82b287b911f32dc4e43a5b3cf1085e934c1109565eb1677

Observation b8a88a99-2a9f-45e9-ab5a-1e92c7b096b4 · outbound

This paper cites SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-17T10:09:46.778093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:01d8d270169dd79dd20e3dccb7197cf71e4b94d2122c2d223f2d0a614ce12d34

Observation 40bb2c53-e9d1-49b9-855f-a2b08239daea · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-11T02:44:53.738794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:9289d650c164be5a94f710905412b61e9b6f8e66a7f6925d9b353a35c23e8664

Observation 77e3eb98-8e06-4249-a49c-28d457da9b53 · outbound

This paper cites Complicated Table Structure Recognition.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Complicated Table Structure Recognition

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:23:58.297348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:1729dfe10ff1713220c996a56c63be339270e961d184f9b1227cfb06ea93797b

Observation 2072faca-4cc8-4def-8c17-493f25100f25 · outbound

This paper cites Icdar2019 robust reading challenge on arbitrary-shaped text-rrc-art.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Icdar2019 robust reading challenge on arbitrary-shaped text-rrc-art

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T13:23:58.398154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:02f0e43b4cf93d9dd20975dbcf6deb504273f8bc2347297829b33c8bc427957f

Observation 8442830c-d1a9-448f-9194-58572393f992 · outbound

This paper cites Scaling instruction-finetuned language models.Journal of Machine Learning Research, 25(70):1–53.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Scaling instruction-finetuned language models.Journal of Machine Learning Research, 25(70):1–53

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T13:23:58.400595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:f7137b3f65417c8a0573f7a694ef1eaf65966338162f1ef7c066526c98e62769

Observation 5d768e4e-9c07-435f-8abe-2c3758eae32e · outbound

This paper cites Simple and effective multi-paragraph reading comprehension.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Simple and effective multi-paragraph reading comprehension

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T13:23:58.402827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:51df0b0aa46add0538f1be679cf4dc99edc00763d26cf89277f76b58ef1a7bc6

Observation 44bf3c12-1781-47b2-af9c-6c485ee89c51 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Training Verifiers to Solve Math Word Problems

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-05-10T13:23:57.818171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:87f75072446c440478b6871f128ac4c23dac879dd0ad41a81bf2fee13ea7eda0

Observation cf974630-4a7d-480c-bae4-e9fdbb75f9d2 · outbound

This paper cites Free dolly: Introducing the world’s first truly open instruction-tuned llm.Company Blog of Databricks.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Free dolly: Introducing the world’s first truly open instruction-tuned llm.Company Blog of Databricks

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T13:23:58.405052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:f4e84c11795971ddec9d4f32bb98ccd7d7b07db34bc5e81a6341c7e9650fc894

Observation 95ba00ff-1a82-44d4-8d0d-2b02e1fddd1a · outbound

This paper cites Mmsegmentation: Openmmlab semantic segmentation toolbox and benchmark.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Mmsegmentation: Openmmlab semantic segmentation toolbox and benchmark

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T13:23:58.407306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:d30a2d7e7f6f39d1e1389464e1adcb353cd80457dd793433d77e751f9149d836

Observation 47acdf29-cdb4-4dd7-a7e4-4fa9d700e13e · outbound

This paper cites Opencompass: A universal evaluation platform for foundation models.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Opencompass: A universal evaluation platform for foundation models

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T13:23:58.409684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:82fbf5612e053b5202461268934065031a528e3c4695cb24a11e4b34a078d275

Observation 55afcede-c500-4d5a-831c-724fe53fda83 · outbound

This paper cites Grok-1.5 vision preview: Connecting the digital and physical worlds with our first multimodal model.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Grok-1.5 vision preview: Connecting the digital and physical worlds with our first multimodal model

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T13:23:58.412165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:cd7d0de914e907630f96d5393d8943f883e18426c16228aba72fab92c38d2a0b

Observation 10522019-f25c-41d5-9fb1-e405a54f695c · outbound

This paper cites UltraFeedback: Boosting Language Models with Scaled AI Feedback.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling UltraFeedback: Boosting Language Models with Scaled AI Feedback

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-17T16:41:29.205174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:a59f6d2139fbc74435aacfa282a1c9bf802e8507b5b0d01045c25d2590dc1c57

Observation f3cd8071-a568-4027-827a-498bba347186 · outbound

This paper cites 15M Multimodal Facial Image-Text Dataset.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling 15M Multimodal Facial Image-Text Dataset

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:23:57.868217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:50254f350970b08ddd15d51d6d1c39d21c34f40dfa44fc00fc08bb6338fba1dc

Observation e43d73ea-e6e2-4505-95b4-506290f133e8 · outbound

This paper cites NVLM: Open Frontier-Class Multimodal LLMs.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling NVLM: Open Frontier-Class Multimodal LLMs

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:23:57.872990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:f11ceb98c6c8c4e21e30773bd7388a40a824a25080e2d8c38f58c5b2cdd79f0c

Observation d9215db8-6b1a-4f96-94c6-4b1a1308751d · outbound

This paper cites Visual dialog.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Visual dialog

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T13:23:58.414638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:48bcde56dbb0952ad0fbc5d9deaa050e80f3f9afeda5f6a0afd3edf27941e250

Observation 1f38ce5d-6752-4d25-a716-95b6b5740df3 · outbound

This paper cites Deep visual template-free form parsing.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Deep visual template-free form parsing

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T13:23:58.417043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:5ae2de22ad38a4b12ceda23c7d97734cf948a8704c9719929521a0d19eff0051

Observation 5e330b7a-74b4-4e11-bf4b-817f6e0fb2ab · outbound

This paper cites Scaling vision transformers to 22 billion parameters.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Scaling vision transformers to 22 billion parameters

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T13:23:58.419210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:93b16f0b47a6f306e59ff09b31ae74f14e7092664499e2f4e19a8ed8133beeb1

Observation 7286b563-8a32-4138-aa2c-422486401088 · outbound

This paper cites Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-15T01:55:12.888993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:fa2470c40176c80daf125b2beb9d55abd8f77e2010d5c6f94cfeae6d58b7c7f2

Observation e83a9c2e-b407-44d3-8d3c-e83bb8c36cc8 · outbound

This paper cites Rico: A mobile app dataset for building data-driven design applications.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Rico: A mobile app dataset for building data-driven design applications

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T13:23:58.421594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:4f26b378cda15589618afd65dceea5d27f0b25ec1edd21d86cd976f04f77f34b

Observation d7f5bb5b-eab9-4755-9c7b-5da408ac3dd0 · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Imagenet: A large-scale hierarchical image database

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T13:23:58.424129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:cfaf4cf502f29e393750236f870cd8e83fec900ef3a37ebb7c7bb9f4c911a145

Observation bed1a26e-b107-425b-8654-6f8d9aaae828 · outbound

This paper cites Mind2web: Towards a generalist agent for the web.Advances in Neural Information Processing Systems, 36.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Mind2web: Towards a generalist agent for the web.Advances in Neural Information Processing Systems, 36

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T13:23:58.426651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:ea40376876e3930a0045be1e57eefa1bcafa515212f926fad962615a54f37a41

Observation 1fdc4fd8-47a4-4c8e-8316-1f1e03ad236e · outbound

This paper cites Enhancing Chat Language Models by Scaling High-quality Instructional Conversations.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Enhancing Chat Language Models by Scaling High-quality Instructional Conversations

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:25:08.436185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:bdf59cf5b78a5d56c1b37b9d0bc3efaa6c95cdd283ed21b6b640599215e9fc37

Observation 6310e289-b803-4cb2-b21a-858594f9b6ad · outbound

This paper cites Vintern-1B: An Efficient Multimodal Large Language Model for Vietnamese.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Vintern-1B: An Efficient Multimodal Large Language Model for Vietnamese

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:23:57.959889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:6e4d78764e419830cceb2dd374329f5b702c0385386a87e9d92a7c739fd643cf

Observation caf1b279-98f3-42ab-99e1-0add0841ad00 · outbound

This paper cites InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:23:57.963396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:c2902872e2465b4f4d413e20d53279a25876ac66ec3c2a40182404c82222818c

Observation 9a81c0a4-de17-496c-b52b-09e5ae113c23 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling An image is worth 16x16 words: Transformers for image recognition at scale

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T13:23:58.429366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:7eddedcb82c692c72905f9732fe0b27dd6be2e9fe51f60ee2a785317ddf3f492

Observation a2a7b757-325b-4b77-bb96-7764b6a85952 · outbound

This paper cites Glm: General language model pretraining with autoregressive blank infilling.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Glm: General language model pretraining with autoregressive blank infilling

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T13:23:58.432075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:63ab09c60d1b24832e2531b28a8121f98a1834569cce7dadeafc0e843e8cf3f7

Observation 6dbff460-5d56-49d6-b8d8-fab692dc5b19 · outbound

This paper cites Vlmevalkit: An open-source toolkit for evaluating large multi-modality models.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Vlmevalkit: An open-source toolkit for evaluating large multi-modality models

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T13:23:58.434722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:dd1e71ad0fb16aa18e7501982d7724c4d50d3d54272e9fc551f079c43998149c

Observation 3170f94b-7a23-46b3-b90b-751fc514e2d5 · outbound

This paper cites The Llama 3 Herd of Models.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling The Llama 3 Herd of Models

Reference 64

Resolution
verified exact
local_arxiv, observed 2026-05-10T13:23:57.983459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:d915e68768a99e50caa3577c49dd17f45a2f4c51bc7a3f97aacc19344bf15398

Observation 0bbc87ab-4cc4-4e82-851a-e02b4159ce21 · outbound

This paper cites MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:23:57.987120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:47675100a2157beac0b68d6dc95b6618d4b85e1363b393e119525ca416385f64

Observation 66de3b3f-90be-48f2-8765-9a3ed77d7fe2 · outbound

This paper cites Eva: Exploring the limits of masked visual representation learning at scale.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Eva: Exploring the limits of masked visual representation learning at scale

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T13:23:58.437245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:e4ff5f94e2750d4ab6924a54a40aa4de4f4e5211ec582ece4d5d688e9ea901d0

Observation 62238b76-6a20-4a0d-8d56-4e96b07e5dc9 · outbound

This paper cites Finevideo.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Finevideo

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T13:23:58.439617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:8fff6b561b7934c4a8f1743ade4fdeb5f9fcebbafd122bf65ec25746faab7204

Observation e638a4f1-5d1e-4f00-9214-025ed6f888cc · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-10T20:25:34.870358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:865ecfbf0ce1ee5367c99655faa40d915a0a4b2b3b6d1702c40ac05ff4a71bc2

Observation 4ba4a556-a5f4-46d4-b39d-05c6a1cae324 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:58:42.298802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:cfcd719a8f9157a4a258a06defc7e7b0089c12b539616166662df32cee836af3

Observation b97e81f7-0bde-406c-b183-6b7e09787a8c · outbound

This paper cites BLINK: Multimodal Large Language Models Can See but Not Perceive.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:18:15.850579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:b7cc65933c33256cf5678095df8f81a221717a1db0875e1f143eedbc272aa3f2

Observation 7fedb350-47bb-4f9a-9359-3913c8e6be31 · outbound

This paper cites Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:23:58.094306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:3a409b711659fc8f53a8291f41f02d7f0dd5962ecca00f7f3c7edc4391d1004a

Observation 5339a34a-c7b7-4878-b74e-45b760824bde · outbound

This paper cites Overview of the imageclef 2015 medical classification task.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Overview of the imageclef 2015 medical classification task

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T13:23:58.441847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:0e329e9e5fc9ba0c90116c75277b9a4bd3e7bfd1f17705703aca9a0509ceef7c

Observation fc442fde-599a-4e73-85dc-09b7afc08aa7 · outbound

This paper cites Glaive code assistant v3 dataset.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Glaive code assistant v3 dataset

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T13:23:58.444174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:e8d3866feeb84d961d7006cfd709844621853ad788d8af9dcf4cfaefc240a6ca

Observation a38aa9de-01fc-437f-a34c-796538a9b56b · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answering.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Making the v in vqa matter: Elevating the role of image understanding in visual question answering

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T13:23:58.447140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:d45bbeff5985c7016ac2fae11456169b8319e52d9821183ec256f0ac8b78cc00

Observation 7e44df2f-7b5f-4361-abe4-6e2430872b61 · outbound

This paper cites Wukong: A 100 million large-scale chinese cross-modal pre-training benchmark.Advances in Neural Information Processing Systems, 35:26418–26431.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Wukong: A 100 million large-scale chinese cross-modal pre-training benchmark.Advances in Neural Information Processing Systems, 35:26418–26431

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T13:23:58.450331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:e232143d3c726777d522a47ffa4dc030eac55e8e87649576739d28deaf61bbb2

Observation 5fae675e-72b8-49f3-8c13-71fc57915129 · outbound

This paper cites Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:23:58.280731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:5321e1a5904dcb6731ae2d1ef55a222b56fbe3b0219be360f0d8357337fd5342

Observation a3d66bc5-602d-4f6e-833d-809142c9cb8b · outbound

This paper cites HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models

Reference 77

Resolution
verified exact
arxiv_id, observed 2026-05-17T01:22:04.360724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:9cdd6b556d636eb438d7d407eeddc98485e5cdd6027d4c062223b2dce4e9ce48

Observation 85960aff-c758-4087-be01-70127bcd6fee · outbound

This paper cites Eaten: Entity-aware attention for single shot visual text extraction.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Eaten: Entity-aware attention for single shot visual text extraction

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T13:23:58.452832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:d412d6f1c911033c7c807b1eb951a31c34e7a50f127dd49158eeafe544c22a2f

Observation 67a82fb9-b2d9-4c93-a750-623edba794ac · outbound

This paper cites Synthetic data for text localisation in natural images.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Synthetic data for text localisation in natural images

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T13:23:58.455076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:78d03f18eed37f49b4bfc2e5f4ddca9b9abdf85160299840907d95d621d97f9b

Observation 699a2106-755f-48b2-9e04-9a632f4f974e · outbound

This paper cites OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 80

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:38:21.105975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:0eb5bebf832a9d4e4c5d0b7a6218b5d4cc0f64d6a35ea9fe8d67c5bf255bd9a4

Observation e6bfa9cf-62c2-40fc-91be-52d0081facb4 · outbound

This paper cites WanJuan: A Comprehensive Multimodal Dataset for Advancing English and Chinese Large Models.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling WanJuan: A Comprehensive Multimodal Dataset for Advancing English and Chinese Large Models

Reference 81

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:23:58.347493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:bbfa240b3864df9cd581985d2195ee270f90be75bbd4da804105b3003dba8e23

Observation d0e8539c-9254-4489-90ac-7d2fd41c2461 · outbound

This paper cites Icpr2018 contest on robust reading for multi-type web images.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Icpr2018 contest on robust reading for multi-type web images

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T13:23:58.457412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:abca7a5a902733cfd4a700a6c964492a9eb7bd29c97e9662b409bcf9af5d5c00

Observation 4bf4431e-000a-445a-8ce5-748ec76f7c19 · outbound

This paper cites PathVQA: 30000+ Questions for Medical Visual Question Answering.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling PathVQA: 30000+ Questions for Medical Visual Question Answering

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:14:19.084463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:8306fc9ed60e87120fd70d30b551bca62996eedcc2275dc22baa6b6e5fcbe20f

Observation a5520a15-8b69-4963-bf98-1098b9fc63d0 · outbound

This paper cites The many faces of robustness: A critical analysis of out-of-distribution generalization.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling The many faces of robustness: A critical analysis of out-of-distribution generalization

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T13:23:58.460242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:fc727db1446158835e893e33ae835f66a3250c4d06aa399a6bc5c83daced199c

Observation b92d0459-6b89-4b34-8e2f-4f09117617e7 · outbound

This paper cites Measuring massive multitask language understanding.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Measuring massive multitask language understanding

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T13:23:58.462920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:10b4e4c915e294b698008af47559237edc40fe8349631165a580a5398d94128f

Observation df2856e6-2c71-4da3-992e-ebb5d925fba4 · outbound

This paper cites Measuring mathematical problem solving with the MATH dataset.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Measuring mathematical problem solving with the MATH dataset

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T13:23:58.465637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:619d03c99d56f8f5e726651b98d317b33419c92f53fc36ac4ee94e0e7042a996

Observation cfd386e3-f8ed-4be4-aef3-019c4f0c951b · outbound

This paper cites Natural adversarial examples.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Natural adversarial examples

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T13:23:58.469535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:8552c06dce0cd998572c5009499186361771ba9b2449742f355fce6dff043057

Observation 1866826c-40cb-464d-baf8-ad6367a71bad · outbound

This paper cites understanding.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling understanding

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T13:23:58.473631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:25642acde664c941ddd2d4276b5749c96abcc41e89a897d09d7e7368f3ea9a98

Observation d695a355-98d6-48cb-b35b-c526d39faaae · outbound

This paper cites Parsynth-ocr-200k.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Parsynth-ocr-200k

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T13:23:58.475824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:e6eddcefdd18d253723d656b33be70c4dc48069df7318ecf79641ddd752fded3

Observation 9c88df7b-8560-497e-a806-0816a2ed83c7 · outbound

This paper cites Unnatural Instructions: Tuning Language Models with (Almost) No Human Labor.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Unnatural Instructions: Tuning Language Models with (Almost) No Human Labor

Reference 90

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:23:57.839380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:24a759357a25f4ca81ca6d919ab14745c8e531a0e3f9451bc034bb4aa57c832a

Observation 7c2ade04-1f1a-4b33-9f6c-9ddf43dfba83 · outbound

This paper cites Koniq-10k: An ecologically valid database for deep learning of blind image quality assessment.IEEE Transactions on Image Processing, 29:4041–4056.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Koniq-10k: An ecologically valid database for deep learning of blind image quality assessment.IEEE Transactions on Image Processing, 29:4041–4056

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T00:12:54.670390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:08a7aac12dfba5478c1ff8f5aaa2a6a3a2ba64a919f38afc4f1305ebd3858acb

Observation f9cda3eb-b097-4f41-814d-e2911563940b · outbound

This paper cites ScreenQA: Large-Scale Question-Answer Pairs over Mobile App Screenshots.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling ScreenQA: Large-Scale Question-Answer Pairs over Mobile App Screenshots

Reference 92

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:23:57.851045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:631372a543b7a9e8acc11edbdb12f20dba5f21b937fb0c10ecf2ef780349bd0c

Observation 36eb7878-7b97-440b-86f7-8d44812fc120 · outbound

This paper cites mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

Reference 93

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:23:57.855416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:e3c79c2a60808a0e8334a29b84250ba898de71ca72cefaa44fba640534bcdfa3

Observation 15dd39bf-a818-44c5-916a-6b13c7ab6f3c · outbound

This paper cites Medical-diff-vqa: a large-scale medical dataset for difference visual question answering on chest x-ray images.PhysioNet.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Medical-diff-vqa: a large-scale medical dataset for difference visual question answering on chest x-ray images.PhysioNet

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T00:12:54.736851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:a3d4d960bfafeda98a0966ac4c612f8c9ffd8bb38560b41d53a7aeb24cdc635b

Observation 6f964294-f194-4733-88d5-be647956b2a8 · outbound

This paper cites Movienet: A holistic dataset for movie understanding.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Movienet: A holistic dataset for movie understanding

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T00:12:54.824691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:a4cf0af3dc1d609c861e711625155cbc82a623b4189a64bd28eb8dba216aaf8f

Observation 2f3cc5e8-6445-4223-bfea-2f528fcd8dd9 · outbound

This paper cites C-eval: A multi-level multi-discipline chinese evaluation suite for foundation models.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling C-eval: A multi-level multi-discipline chinese evaluation suite for foundation models

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T00:12:54.801783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:80cf73ad33ddb6a3a5d392dcc76b2bdc8aaf74a3e717ac1c5af0eb9fc52f0386

Observation 21468330-e7e7-4cca-9bb7-6c2f79fa3cef · outbound

This paper cites Icdar2019 competition on scanned receipt ocr and information extraction.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Icdar2019 competition on scanned receipt ocr and information extraction

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T00:12:54.615637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:28414fe47b087f72a9f36769cb17f8eeae5c5696ddfebd0fe9d74312f531ee5d

Observation cab861df-f609-4b41-9152-7f9c9c61b937 · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T00:12:54.516591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:a0da21161770cee5203ccf87e683d3f9c00f6c2985cbce4464bf2511a0fafad9

Observation 6ef5c445-9ec3-44d0-bcea-6883d0311021 · outbound

This paper cites Egotaskqa: Understanding human tasks in egocentric videos.Advances in Neural Information Processing Systems, 35:3343–3360.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Egotaskqa: Understanding human tasks in egocentric videos.Advances in Neural Information Processing Systems, 35:3343–3360

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T00:12:54.666807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:0ebb3caebf9b4a05edb93de037fec0236029bc097501dc66b78a512bf882d592

Observation d3cb90da-ce14-4dc8-937c-e425142bc3da · outbound

This paper cites MANTIS: Interleaved Multi-Image Instruction Tuning.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling MANTIS: Interleaved Multi-Image Instruction Tuning

Reference 100

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T13:23:57.921274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:bbbb724183da74229a42ba4cd2f57150f0a840edc754a80317f756ea93a48d4f

Pith citing papers

Observation dbf3dc5f-04ef-452f-b47e-d9e69ed2731a · inbound

Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction cites this paper.

Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 73

Resolution
verified exact
local_arxiv, observed 2026-05-18T04:09:41.571937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-18T04:09:41.494136Z digest=sha256:366c2b59f598ca22fc0a6823523719c27437e240adfd096c71580a9393b26475

Observation 4a4005a2-c553-4284-93bf-4d199e027499 · inbound

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model cites this paper.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T23:07:14.402030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:07:14.402030Z digest=sha256:3a14c2cdaa18906ab31d0f07d8fe57c33d0ded87e4ae231f7292bddb98521fe2

Observation fe5de7f2-d8da-4945-9db7-4404a8086643 · inbound

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning cites this paper.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:33:26.755761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:f5a401118622adc55a27daadc546e247c2625fa7def8b0f8b2534ede67bc7a39

Observation 4b69a1b7-2758-4655-b24f-0ec49c7eaa31 · inbound

CaReBench: A Fine-Grained Benchmark for Video Captioning and Retrieval cites this paper.

CaReBench: A Fine-Grained Benchmark for Video Captioning and Retrieval Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:00.747316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:00.747316Z digest=sha256:a5dcfe6546471db4a250c954b53245560f5cdf2f2dbe4054b71e74e0f2b09f01

Observation db4489fd-e6fb-47cb-98ff-2ea6d77c59d0 · inbound

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling cites this paper.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-18T04:02:43.692536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:610d3eba9d7c3aaa478fbaccdced8bfc0abb7f022270bc7bb7910bb826f1dfd2

Observation d5cdff0d-953d-440d-997d-61b082553ac0 · inbound

HLV-1K: A Large-scale Hour-Long Video Benchmark for Time-Specific Long Video Understanding cites this paper.

HLV-1K: A Large-scale Hour-Long Video Benchmark for Time-Specific Long Video Understanding Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:56.494682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:56.494682Z digest=sha256:f718c6754e8ca30c538c8ab87c3b4506a18fc61749a482e67bf6e8f766dad8ba

Observation fdfc9d99-3ca6-4c53-8d07-d375493f4258 · inbound

Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos cites this paper.

Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-16T11:39:22.419329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T11:39:22.340737Z digest=sha256:be8cf58385c2a8b691af6cae3dad046d4a8e44d697609d228800a13fb0ffaaad

Observation 087cd7be-d6e1-4c05-b617-a14ff3b46f4a · inbound

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs cites this paper.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:15.878718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:31:15.878718Z digest=sha256:88b9b72b35a4cb29a9d438926c12af6e78d8790cf6ad5d40cd433db51c6091aa

Observation c85b112d-5f3d-4c60-b896-0c86dd59248e · inbound

LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding cites this paper.

LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-23T06:02:37.386976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T06:01:00.775721Z digest=sha256:08132bf67da33baf9fe932c97bde6fabda54d16906bc87fb7875bce4fa66e28f

Observation e14bc7da-57d3-437f-930e-d95c1e01be9e · inbound

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark cites this paper.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:59.495233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:59.495233Z digest=sha256:a54e106c5a678f9b5bef0169b5e44da02c604e88708caa5f9212a6b4260b32ce

Observation 07d3c61f-97c8-43a9-b894-6fb073d2041b · inbound

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design cites this paper.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:18.883573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:18.883573Z digest=sha256:fac0da90482c95ec61e55cd35bc688f67058cc02dcab8f8bc7ef3cf11b953ee1

Observation f56fa86e-db31-4618-a2a3-65e0feffd7b9 · inbound

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding cites this paper.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:35.338553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:35.338553Z digest=sha256:6a17429ae22113407bd22c66de086ac4fa236e6949e34daabda66a40c41eb061

Observation 03ae5b0e-f7eb-4be8-9daf-d4fa9ce6d87b · inbound

Omni-Emotion: Extending Video MLLM with Detailed Face and Audio Modeling for Multimodal Emotion Analysis cites this paper.

Omni-Emotion: Extending Video MLLM with Detailed Face and Audio Modeling for Multimodal Emotion Analysis Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T20:03:12.706160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:03:12.706160Z digest=sha256:a4aede39c7a60605f7d47151f472af50494420966d9d5dd3d78e85fb47d377ee

Observation fdacd6b9-e11a-4829-a704-0402494f65ce · inbound

AdaFV: Rethinking of Visual-Language alignment for VLM acceleration cites this paper.

AdaFV: Rethinking of Visual-Language alignment for VLM acceleration Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T20:00:12.574172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:00:12.574172Z digest=sha256:2e3ec45d95e7f60fc883c42951a04c0b28c2392df99fb7a88571a4e33f8d1f1f

Observation d12f5c2c-075e-46f1-a857-0347cf2ed681 · inbound

A Simple Aerial Detection Baseline of Multimodal Language Models cites this paper.

A Simple Aerial Detection Baseline of Multimodal Language Models Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T19:48:22.445114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:48:22.445114Z digest=sha256:92ac4f9e880e0b4484006abf9e74d5867992db8b9e63fcb34b25a52549e3b508

Observation 8514a535-74bc-4bba-ac64-6b00fd3a5803 · inbound

Is your LLM trapped in a Mental Set? Investigative study on how mental sets affect the reasoning capabilities of LLMs cites this paper.

Is your LLM trapped in a Mental Set? Investigative study on how mental sets affect the reasoning capabilities of LLMs Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T17:51:06.874624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:51:06.874624Z digest=sha256:d89ae6e6c9c311f341f4f4bc211f64e341c76375400ba31749e28eadd5a193ec

Observation fdbc36e7-7bd3-4698-8f73-99efc9412160 · inbound

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model cites this paper.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:40.020319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:40.020319Z digest=sha256:8a8bc485a220d476d33b0f37fd0c437801509cf645e071c356ea11106fa36a46

Observation 9bc5ee28-ccce-48cb-8d24-fbf9ab89ec04 · inbound

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding cites this paper.

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-11T01:19:59.991095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T01:19:59.603343Z digest=sha256:e6a63038724c04bfe188047cb9ca0cec2384c725f39023af616895a0fd9b1ce8

Observation 6d0295fd-ec2e-40ef-b063-7e45bdf771de · inbound

OCSU: Optical Chemical Structure Understanding for Molecule-centric Scientific Discovery cites this paper.

OCSU: Optical Chemical Structure Understanding for Molecule-centric Scientific Discovery Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T14:23:33.436247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:23:33.436247Z digest=sha256:e5812a522fb591ae436645e9499413cfeface8904dd630f91594e6d025527600

Observation bba54c0c-287d-494f-aa6e-d45ed753f2dc · inbound

Ocean-OCR: Towards General OCR Application via a Vision-Language Model cites this paper.

Ocean-OCR: Towards General OCR Application via a Vision-Language Model Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T14:14:54.622105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:14:54.622105Z digest=sha256:2699cf70a25cd22bd0d1f3567080d8dbfc6deef690b4d1f1c9b8c9db085a24b7

Observation 7a099e45-6051-4743-a87f-585d228c25a1 · inbound

AIN: The Arabic INclusive Large Multimodal Model cites this paper.

AIN: The Arabic INclusive Large Multimodal Model Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-09T20:17:40.316184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T20:17:40.316184Z digest=sha256:99e2c4bdffe76aed6537f6da4d68567d3b4ec3936f6ee8a12601992262cdf08f

Observation 43321dff-4be8-48d7-8055-8195a6e15781 · inbound

MRAMG-Bench: A Comprehensive Benchmark for Advancing Multimodal Retrieval-Augmented Multimodal Generation cites this paper.

MRAMG-Bench: A Comprehensive Benchmark for Advancing Multimodal Retrieval-Augmented Multimodal Generation Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T23:20:00.987901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T23:20:00.987901Z digest=sha256:60365420ce6dc5aa5ed4db1b0ac6df529bf51807e7ce9acfeb03b6b0f4b478fa

Observation 2193c28a-8434-4b28-9561-5bd00eefe9ad · inbound

WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs cites this paper.

WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-17T05:53:26.373635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T05:53:26.066674Z digest=sha256:cfd2114bdf0fb018c595ee12bb67b4b7ef24180f904830a242c2d1718e33f4d3

Observation d3d61bf3-5cef-4116-b793-1bcec186b665 · inbound

Ola: Pushing the Frontiers of Omni-Modal Language Model cites this paper.

Ola: Pushing the Frontiers of Omni-Modal Language Model Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.073045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.073045Z digest=sha256:b61e57d39939b63d528b5777c7914ecad1f71b4da6b7d7f338894717ecfdc764

Observation 3e497a2f-7f7b-4020-a666-7df7f524aa32 · inbound

SARChat-Bench-2M: A Multi-Task Vision-Language Benchmark for SAR Image Interpretation cites this paper.

SARChat-Bench-2M: A Multi-Task Vision-Language Benchmark for SAR Image Interpretation Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T10:14:30.118925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:14:30.118925Z digest=sha256:9bf22845f0d93a7844ca8408ec210a6df8c57ae50a705e9e847e1ba30bcc14fd

Observation a0a61174-98ae-4d45-a0c6-b7deacda45be · inbound

EmoAssist: Emotional Assistant for Visual Impairment Community cites this paper.

EmoAssist: Emotional Assistant for Visual Impairment Community Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T22:05:56.815330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T22:05:56.815330Z digest=sha256:e6de774ebd35805c38b7bc3303449cfe3957cdd375b73e91212ff0d05bcdbd77

Observation be40dd94-a9cd-4805-a809-abc0d8a832aa · inbound

Qwen2.5-VL Technical Report cites this paper.

Qwen2.5-VL Technical Report Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-23T02:25:19.071625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T02:25:04.405036Z digest=sha256:41f70e9cde60318f1cdc25d2f1c074d915bd9411ce596a0b8359728172041b2d

Observation 3a059973-2a2b-4636-9b9f-9ab0c3887ebe · inbound

Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs cites this paper.

Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-11T22:22:27.842608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T22:22:27.455361Z digest=sha256:3b01a8a2bf8aa1b1e161dbdaa9a9f4e828f1332d980ebf22f175a7ba0eaed9f3

Observation 4cdad193-8e50-4b67-9043-fa2598f9f8e8 · inbound

AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems cites this paper.

AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-11T15:09:24.633068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T15:09:24.367362Z digest=sha256:08ad696d6d96b2f5aadbbd3db1d0cc85dc72c1e51d74648868cbd736cc7682c4

Observation 27815057-470f-436b-8557-88f2dbea35fa · inbound

AlphaDrive: Unleashing the Power of VLMs in Autonomous Driving via Reinforcement Learning and Reasoning cites this paper.

AlphaDrive: Unleashing the Power of VLMs in Autonomous Driving via Reinforcement Learning and Reasoning Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-16T20:06:27.187089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T20:06:27.136345Z digest=sha256:fac942c0e6614a8ccec1b1a5e104791c6a8fdb660bda68e83d11216180a99343

Observation 4cf32f77-ca1b-417f-aa46-6735afb14994 · inbound

R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization cites this paper.

R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-16T00:19:20.612554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T00:19:20.462455Z digest=sha256:6f67cad613e079e47abbd4fe3506e5fe00bf4865ba79d8071a7a4bcfbc9caf9b

Observation 454d47c7-8dd6-4678-8e23-d5621c470468 · inbound

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey cites this paper.

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 272

Resolution
verified exact
local_arxiv, observed 2026-05-15T17:18:53.605459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T17:18:52.996467Z digest=sha256:7b3d9678eabd009d6bccffbf35f24a94cdb673376259758e6acc3497254cffbf

Observation 715acb0e-a21b-4696-a749-b7d8769192c9 · inbound

Cosmos-Reason1: From Physical Common Sense To Embodied Reasoning cites this paper.

Cosmos-Reason1: From Physical Common Sense To Embodied Reasoning Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-16T12:47:10.200336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T12:47:10.146795Z digest=sha256:924f26c9aa8be8d4ef9701363ee3c7dd6d31f13e8ba26c29ebb9242c8cd9758d

Observation f62fed40-d4c2-4ab0-b2b4-19a60f7d09ad · inbound

MathFlow: Enhancing the Perceptual Flow of MLLMs for Visual Mathematical Problems cites this paper.

MathFlow: Enhancing the Perceptual Flow of MLLMs for Visual Mathematical Problems Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-22T22:57:13.284486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T22:55:34.238427Z digest=sha256:f76c2ab1914902080453855a289595c931477e09a9f08598d6988022ac620bac

Observation d00c3be7-a55c-4580-b867-4b22d5693751 · inbound

OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles cites this paper.

OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-19T06:59:03.365351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T06:59:03.112252Z digest=sha256:c7774c2233ac9081b3b636655ffccb1064459ee874eea0f78f00ea6ef0e2ae72

Observation 413e1df8-f063-46eb-9398-5b712199bbec · inbound

SpaceR: Reinforcing MLLMs in Video Spatial Reasoning cites this paper.

SpaceR: Reinforcing MLLMs in Video Spatial Reasoning Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-15T15:18:43.827202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T15:18:43.724432Z digest=sha256:201d487acf0cdcb54b9aca3155006d49ee4694b031fcc23253a130dc2b50d7cc

Observation f97b9fff-840d-4c74-a7a3-fb38cd2134f6 · inbound

VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model cites this paper.

VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-13T01:13:57.530850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T01:13:57.368874Z digest=sha256:fb3b3e2edfb7f4a555a8664395d21392b5f4bb9c4d2b476b0d55af04d8955fe8

Observation dca35e67-146b-42ea-926a-304bbc25abee · inbound

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models cites this paper.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-10T13:41:08.239742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:7682d1aa954e6ba495ea4da51c6a498c7d592d5bda132e14052185becbb8897f

Observation 2ab87385-289b-4fe2-a7e0-3b51be7dd4cc · inbound

Consensus Entropy: Harnessing Multi-VLM Agreement for Self-Verifying and Self-Improving OCR cites this paper.

Consensus Entropy: Harnessing Multi-VLM Agreement for Self-Verifying and Self-Improving OCR Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-22T20:32:04.684976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T20:31:34.074705Z digest=sha256:f44f05e6c70b8caadbd1be31814c5386abc41090268bcedcbbf0626ae67a6537

Observation 57bd24ad-6817-4225-ace8-5877e4367350 · inbound

Perception Encoder: The best visual embeddings are not at the output of the network cites this paper.

Perception Encoder: The best visual embeddings are not at the output of the network Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-13T22:21:15.752203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T22:21:15.681336Z digest=sha256:39567de29e5c1accb7803169d9230af82f087ad0d1cea3dcd7299faf3ee0b5f5

Observation ec1683fa-ebdf-461b-bfa0-47116c0ab52a · inbound

We'll Fix it in Post: Improving Text-to-Video Generation with Neuro-Symbolic Feedback cites this paper.

We'll Fix it in Post: Improving Text-to-Video Generation with Neuro-Symbolic Feedback Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 12

Resolution
metadata mismatch
local_arxiv, observed 2026-05-22T18:56:58.261190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T18:56:39.735334Z digest=sha256:1f6c262fe109e702d9e2ec36ff14aa040809d86f5a46b0f13609ade02d9951ab

Observation 48eafba0-689d-44b6-a352-d9810ad2d040 · inbound

CAFES: A Collaborative Multi-Agent Framework for Multi-Granular Multimodal Essay Scoring cites this paper.

CAFES: A Collaborative Multi-Agent Framework for Multi-Granular Multimodal Essay Scoring Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:31.399503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:42:31.399503Z digest=sha256:5724584aa85f6283a7685f02cd7dfd49eb6f1cd7e84464bee4632458fbb8d086

Observation fcbff939-49af-4d50-8fd5-c9dad6aed463 · inbound

ShieldVLM: Safeguarding the Multimodal Implicit Toxicity via Deliberative Reasoning with LVLMs cites this paper.

ShieldVLM: Safeguarding the Multimodal Implicit Toxicity via Deliberative Reasoning with LVLMs Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:21.649314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:21.649314Z digest=sha256:d3ffb95113cb4cf59e45c7272bb5f48bef1adfe03f9930890c5558f60b3895a1

Observation 1bb37353-e8cb-4ac9-86bc-d6b8d4ebf0d1 · inbound

Visual Agentic Reinforcement Fine-Tuning cites this paper.

Visual Agentic Reinforcement Fine-Tuning Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:25.897449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:25.897449Z digest=sha256:622602dc9a9e8c8c3c1b9634c876340138d9b1d371fbf94abf4f5df75883d0d2

Observation bce2c6ed-8e9d-463c-b9df-bb33676dc6c4 · inbound

DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning cites this paper.

DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-11T14:42:56.719011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T14:42:56.565621Z digest=sha256:186d97239a88a59354ee2a36f80fcf4abe53856aa4426c055ac2dba431460a8e

Observation b9eae6d5-b575-4723-9821-2e20512be1e8 · inbound

Towards a Foundation Model for Communication Systems cites this paper.

Towards a Foundation Model for Communication Systems Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:56.449803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:36:56.449803Z digest=sha256:489e5ba0581c99e597a654a47125ce3c9890a1ed286719bcd9362540686cc9b4

Observation 5d0872ad-c2e1-40e5-a047-d42e50b38f36 · inbound

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation cites this paper.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:44.518133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:44.518133Z digest=sha256:06b7ad0bc3afbbaff56bf47acdde6d1f3de85c1f17e834d5105f10fd8b353ef2

Observation 79f87b93-d215-452b-a229-07b36248cee7 · inbound

Emerging Properties in Unified Multimodal Pretraining cites this paper.

Emerging Properties in Unified Multimodal Pretraining Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-10T16:23:41.995166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T16:23:41.854132Z digest=sha256:88f75bb3636c603cece3c56a5b0bfccdfe3a2a873929b29ff3a56a819d6c54ae

Observation aeb95aec-7c5b-48a8-9143-5c9c25861d5b · inbound

Clapper: Compact Learning and Video Representation in VLMs cites this paper.

Clapper: Compact Learning and Video Representation in VLMs Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:44.667515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:44.667515Z digest=sha256:6614c6ebc44dd50192e9b79f348b443a58b75230c75d5164bf429b1e92327639

Observation b1d57b38-53f5-49c4-a16d-20c220232ee3 · inbound

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs cites this paper.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:41.300013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:41.300013Z digest=sha256:8de2d2844bd69d58a9e7bfc7b541910f27226e83bb54863205801ff9e5a17c73

Observation 294efcf3-6fe5-4ed8-bcf1-da716dec89ff · inbound

Training-Free Reasoning and Reflection in MLLMs cites this paper.

Training-Free Reasoning and Reflection in MLLMs Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T15:13:57.151656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:13:57.151656Z digest=sha256:1ec9f5741278111f9b2032d2fb4ae064b08f03b9a4e300bb6424dce64abcf68e

Observation cfd80a0d-2535-4e0e-8341-55ac8772c364 · inbound

QuickVideo: Real-Time Long Video Understanding with System Algorithm Co-Design cites this paper.

QuickVideo: Real-Time Long Video Understanding with System Algorithm Co-Design Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T15:08:58.336854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:08:58.336854Z digest=sha256:d8fd894591df47c5153f50973fb2bcbf4284a79befd83c9f3589b5fc06ac3097

Observation f10b55dc-d9f9-484d-b772-14fa6ab0c3ed · inbound

Understanding Generative AI Capabilities in Everyday Image Editing Tasks cites this paper.

Understanding Generative AI Capabilities in Everyday Image Editing Tasks Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:19.079742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:19.079742Z digest=sha256:76b199fa92c840293433c5d975b89dfe0d2e64c5ab0a20fc1f0d72aa567bddc9

Observation a075c7a6-6f15-4551-b397-40d6859966ba · inbound

NTIRE 2025 challenge on Text to Image Generation Model Quality Assessment cites this paper.

NTIRE 2025 challenge on Text to Image Generation Model Quality Assessment Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:42.460938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:42.460938Z digest=sha256:2db1a210c2eaf280376867f9667d7c1bc8aea4a2b95e217b2125ba924adff8c4

Observation 050d2c3b-ca16-4e0d-857b-5540d678d2de · inbound

Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models cites this paper.

Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-22T14:21:39.737877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T14:19:34.622854Z digest=sha256:c0ac3908cce0cd1fe8cb544a9ffd66e0d37724dddb607211298dca6477d007f2

Observation 724ccaf1-5152-4020-b225-fbaa6e536ca7 · inbound

Bridging the Dynamic Perception Gap: Training-Free Draft Chain-of-Thought for Dynamic Multimodal Spatial Reasoning cites this paper.

Bridging the Dynamic Perception Gap: Training-Free Draft Chain-of-Thought for Dynamic Multimodal Spatial Reasoning Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:25.132191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:02:25.132191Z digest=sha256:e2ae3abf0b4a8b7a2bd685f546cd4059c9a0a4e26bb3719c3e1da0aa310fa87c

Observation 13a8bef6-d4c6-4bb7-9e4f-450488504528 · inbound

DetailMaster: Can Your Text-to-Image Model Handle Long Prompts? cites this paper.

DetailMaster: Can Your Text-to-Image Model Handle Long Prompts? Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:25.989364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:25.989364Z digest=sha256:a0f18fe007fc69c4c1a4ca4717e603a69531a45c4f4916411e846530efc2dfe6

Observation df65dd4a-7c41-4556-91b1-3026e5e6bdad · inbound

Backdoor Cleaning without External Guidance in MLLM Fine-tuning cites this paper.

Backdoor Cleaning without External Guidance in MLLM Fine-tuning Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:48.792555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:56:48.792555Z digest=sha256:9103b9e06be37b3cc4af99a2a0539df842ff4076c8327f87a4222bb0e257d9c0

Observation dc7986d5-c5af-4279-bd62-9dae15108a73 · inbound

SpatialScore: Towards Comprehensive Evaluation for Spatial Intelligence cites this paper.

SpatialScore: Towards Comprehensive Evaluation for Spatial Intelligence Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-22T13:11:35.756580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T13:07:11.548885Z digest=sha256:06bbf3ac1eda101670945f5adb67278b7d0e9fda47672a4ca8335d51b3796730

Observation 4aa1239c-a36c-4ac4-8feb-489ff15f1a8f · inbound

ARB: A Comprehensive Arabic Multimodal Reasoning Benchmark cites this paper.

ARB: A Comprehensive Arabic Multimodal Reasoning Benchmark Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:24.001934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:24.001934Z digest=sha256:29e04a187eebfb8a5d36aa90e6e6b67ca29ba5887166ca1f11286f016d1cf706

Observation 917cbe05-127a-47dd-8b36-ba927934ca76 · inbound

FullFront: Benchmarking MLLMs Across the Full Front-End Engineering Workflow cites this paper.

FullFront: Benchmarking MLLMs Across the Full Front-End Engineering Workflow Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:51:18.119095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:51:18.119095Z digest=sha256:5b9855357a092d0339fc75afdae107b4c50244ddfd2911decd92d60f59dad6e2

Observation 822023ef-72e4-4818-a6fb-0a8c6d76f737 · inbound

VIBE: Annotation-Free Video-to-Text Information Bottleneck Evaluation for TL;DR cites this paper.

VIBE: Annotation-Free Video-to-Text Information Bottleneck Evaluation for TL;DR Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T14:51:44.270249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:51:44.270249Z digest=sha256:900136fac5d66dcea739978a4c4390740292df0c9a084f732bf3c213ad3f3401

Observation 83cd6933-ec06-4df0-b1eb-f65ae61f62c2 · inbound

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts cites this paper.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:13.827834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:13.827834Z digest=sha256:4c434919ff2a35e3b527c00deac8b12907e900a91f33f16e53f3f10980b8765c

Observation 84836bc0-651d-46b3-8fbf-c737d4b9694f · inbound

MLLMs are Deeply Affected by Modality Bias cites this paper.

MLLMs are Deeply Affected by Modality Bias Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:21.584728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:21.584728Z digest=sha256:0a16cb0832520b6f541c40c7afe8aa1eb52730eabcbaa8300e03d2d4e1d7e97c

Observation 33c63cf9-3e2e-4215-ada8-85d55d388162 · inbound

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion cites this paper.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:41.747798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:41.747798Z digest=sha256:1d7eebf4b6aed47efee432c4c4b4267d631992511eae1b3ecbf9f886f7586738

Observation ad2e589f-6bcc-4b8b-8565-678b05f362ab · inbound

CardioCoT: Hierarchical Reasoning for Multimodal Survival Analysis cites this paper.

CardioCoT: Hierarchical Reasoning for Multimodal Survival Analysis Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:23.679059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:23.679059Z digest=sha256:b40ec225e24b7130d006048f9adb874ff10c4783f9bec845b9c4f9ec2b6f3d25

Observation f4f7ae01-793e-4ea3-90f5-516f7ba1e176 · inbound

Improving Medical Reasoning with Curriculum-Aware Reinforcement Learning cites this paper.

Improving Medical Reasoning with Curriculum-Aware Reinforcement Learning Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:24:42.366000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:24:42.366000Z digest=sha256:b04bf250212c7d58dd6347714a600af84cca884255f262ba8aa4e1d335144e78

Observation f295fbe9-90b6-48b8-ad51-5d90ac34222a · inbound

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model cites this paper.

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:01.977150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:01.977150Z digest=sha256:73630d2a0d833adfd27042a957f979b5ec6a01e7691166defe5abf90ead30889

Observation 3c97a7b7-3b6a-40a8-959d-3a065bd0e19d · inbound

Benchmarking Multimodal Knowledge Conflict for Large Multimodal Models cites this paper.

Benchmarking Multimodal Knowledge Conflict for Large Multimodal Models Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:55.605857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:55.605857Z digest=sha256:4e8889c539afa558a03885b0d2939fdef95fec73f0dc4c88898f9bb30a1ee123

Observation 2699f0e0-d716-4443-96fa-cfcfafcf8b6a · inbound

Align and Surpass Human Camouflaged Perception: Visual Refocus Reinforcement Fine-Tuning cites this paper.

Align and Surpass Human Camouflaged Perception: Visual Refocus Reinforcement Fine-Tuning Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:54.408709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:13:54.408709Z digest=sha256:781c81e92d3dbc4e6510c46e7508d152da40853852aa3c9f0c0ac64c5cfdf0f1

Observation 037a3a34-ad60-4583-98aa-61a6086261d6 · inbound

VisCRA: A Visual Chain Reasoning Attack for Jailbreaking Multimodal Large Language Models cites this paper.

VisCRA: A Visual Chain Reasoning Attack for Jailbreaking Multimodal Large Language Models Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:16.910969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:13:16.910969Z digest=sha256:5cfa66c8dfefce774da301d576ecaf4094a561083e550991aa4f8c6983f2f8fa

Observation 2bf1845f-f9b1-4086-930c-d38c6ee018da · inbound

MT$^{3}$: Scaling MLLM-based Text Image Machine Translation via Multi-Task Reinforcement Learning cites this paper.

MT$^{3}$: Scaling MLLM-based Text Image Machine Translation via Multi-Task Reinforcement Learning Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:59.949559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:13:59.949559Z digest=sha256:7906874edf5cf7f0e7b647ef7b4571441f5fd109b817b4aa0709b4ea0bc88e54

Observation bc8321d9-4ebe-4ec6-9b16-5f6864a728b2 · inbound

AdaTP: Attention-Debiased Token Pruning for Video Large Language Models cites this paper.

AdaTP: Attention-Debiased Token Pruning for Video Large Language Models Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:08:07.651269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:08:07.651269Z digest=sha256:92e3aeea8d23c03ef66ec766eddc06b71496271e9e4e948a2438afaab5dd5b0f

Observation 373e72bc-1f8a-40de-bfaa-f65fa07922f0 · inbound

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration cites this paper.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:53.381768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:53.381768Z digest=sha256:f6b0131cc8eca618493d5bd5034fd68f9d0394ab8e67cdb023edc53e6d991041

Observation bbb13462-b46a-4cf4-9581-2759a67ac555 · inbound

Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning? cites this paper.

Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning? Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-05-17T05:40:56.057848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T05:40:55.944288Z digest=sha256:54a6311bee48af8fcda899002fe79bec4af0bf8a7789b51990b98416c9406908

Observation 0c8d5d8f-a9db-45a7-b901-2184cca4158c · inbound

AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs cites this paper.

AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T13:40:17.736081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:40:17.736081Z digest=sha256:5bd9a9b500acfb22f463e27101ef9eaefb8548ae4e68e5f521c86033039517fc

Observation 794782d4-e9a5-40f2-8394-4d03513b5c12 · inbound

ID-Align: RoPE-Conscious Position Remapping for Dynamic High-Resolution Adaptation in Vision-Language Models cites this paper.

ID-Align: RoPE-Conscious Position Remapping for Dynamic High-Resolution Adaptation in Vision-Language Models Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:53.834178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:53.834178Z digest=sha256:2dca5840b41d6b70e9835af0375c6885b17372d46c8301ebbac96bec96332523

Observation e9790918-9649-4cc5-ab71-331c48752194 · inbound

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning cites this paper.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:49.709222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:49.709222Z digest=sha256:e7ec9f59788c3109dede130e094fd2808d7d7015bbb02abd3bea411718a28e2d

Observation ab1fe195-2f7d-4b5c-bb0e-a44cd8a1818e · inbound

Reinforced Reasoning for Embodied Planning cites this paper.

Reinforced Reasoning for Embodied Planning Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:21.610900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:21.610900Z digest=sha256:4e8c392480dc27909bc673c7918db1be7d42809b896aa4a3ab10099e757f400f

Observation 7e203eb4-52f9-4748-a37d-59c946346027 · inbound

SridBench: Benchmark of Scientific Research Illustration Drawing of Image Generation Model cites this paper.

SridBench: Benchmark of Scientific Research Illustration Drawing of Image Generation Model Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T13:19:04.332212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:19:04.332212Z digest=sha256:a9a7f9e0d328c4b2f2c51ef7e25f7a94083a8befd4ff77a4c29618c3327dbf0b

Observation 616cbabc-ee87-4b4a-81fc-493b07965263 · inbound

Zooming from Context to Cue: Hierarchical Preference Optimization for Multi-Image MLLMs cites this paper.

Zooming from Context to Cue: Hierarchical Preference Optimization for Multi-Image MLLMs Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T13:14:09.473336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:14:09.473336Z digest=sha256:2f2338dfacf925326b7c2aa58a378294cbeb8b831bd9cf328d09b05f249c6c39

Observation 233a67b4-7300-4094-bd33-e442cae16063 · inbound

Scaling-up Perceptual Video Quality Assessment cites this paper.

Scaling-up Perceptual Video Quality Assessment Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:05.938818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:05.938818Z digest=sha256:ff5da603280744bf893c0449dcb9a813df01efb1ab043e6f68fb4a4666e96da5

Observation c51cfa79-1d7a-46af-bebb-b92b42132eb1 · inbound

Universal Visuo-Tactile Video Understanding for Embodied Interaction cites this paper.

Universal Visuo-Tactile Video Understanding for Embodied Interaction Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:26.367593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:26.367593Z digest=sha256:6bbc3a9fa613000f2be28ee483d4493b51dbb6fed749b16a2c3cb2b1f8c58fff

Observation 50f5e6e4-7b4e-42e3-81e1-9c05c4c47877 · inbound

Synthetic Document Question Answering in Hungarian cites this paper.

Synthetic Document Question Answering in Hungarian Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:24.706741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:24.706741Z digest=sha256:3eb6e8a8d26f1d2593df8ef5880ba3f454431ea577b769d562047960f8dd7053

Observation 1633aeeb-f418-4f81-ad87-d005d7b4126e · inbound

Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation cites this paper.

Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T13:00:57.690899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:00:57.690899Z digest=sha256:18de13b6e01118d8103cb90cb5229911be5ce60c7f73456cb23fd954dba95c3f

Observation 064d151c-f6d6-48f8-abe0-6684ee452d39 · inbound

Infi-MMR: Curriculum-based Unlocking Multimodal Reasoning via Phased Reinforcement Learning in Multimodal Small Language Models cites this paper.

Infi-MMR: Curriculum-based Unlocking Multimodal Reasoning via Phased Reinforcement Learning in Multimodal Small Language Models Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:22.902200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:22.902200Z digest=sha256:8174a1e93d7b6694c97597616938289e3d8e575dca244880bfaa176e3256894f

Observation 25d24c42-a878-4a4f-a6eb-056edf6a179e · inbound

ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding cites this paper.

ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:40:43.917741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:40:43.917741Z digest=sha256:22be8625975f5d97c68fa939f7072f645e0bd0e179b2d66f2f0a32c15efae9c3

Observation 82dd815f-cc4c-4aac-a85e-a3dd84af6992 · inbound

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models cites this paper.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:43.644290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:43.644290Z digest=sha256:c1a99e1e5872229523f01fddc076e9e0b6e94b909b971902965cb2f9609612c2

Observation 24f747fc-509d-416c-80e3-8e1d9704fa43 · inbound

Seeing is Not Reasoning: MVPBench for Graph-based Evaluation of Multi-path Visual Physical CoT cites this paper.

Seeing is Not Reasoning: MVPBench for Graph-based Evaluation of Multi-path Visual Physical CoT Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T12:36:24.031956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:36:24.031956Z digest=sha256:32f2e39c55c764bf8e18af9cfeb6d8af27ace4f2b5dbb98332520bfc426c9568

Observation 538b5115-2b18-4501-8090-17fa6c5d4c27 · inbound

DisTime: Distribution-based Time Representation for Video Large Language Models cites this paper.

DisTime: Distribution-based Time Representation for Video Large Language Models Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:50.466205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:50.466205Z digest=sha256:a7a13e1822984ada58e5d9eba172126505d7deacae6aa07d7a691a312b8b52c5

Observation 87aac395-fd51-415d-992b-53fc65b37f08 · inbound

Unifying Language Agent Algorithms with Graph-based Orchestration Engine for Reproducible Agent Research cites this paper.

Unifying Language Agent Algorithms with Graph-based Orchestration Engine for Reproducible Agent Research Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:29:17.939296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:29:17.939296Z digest=sha256:c652c44fe8638759b824d934bc73a89b551e9211f147d6ca2f62d522536a54d3

Observation c113cf7c-5adc-4af2-af5c-8140a0ea177c · inbound

MMAFFBen: A Multilingual and Multimodal Affective Analysis Benchmark for Evaluating LLMs and VLMs cites this paper.

MMAFFBen: A Multilingual and Multimodal Affective Analysis Benchmark for Evaluating LLMs and VLMs Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:27:31.646610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:27:31.646610Z digest=sha256:5f0f2e9ff0a7035ec6eec803584fb9a094044b25188a633194404e74a7d682b2

Observation 082977d1-44db-463c-a5ba-4a4e62f6ed8b · inbound

SORCE: Small Object Retrieval in Complex Environments cites this paper.

SORCE: Small Object Retrieval in Complex Environments Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:27:28.026719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:27:28.026719Z digest=sha256:053d373434a6bc7d26f35e88cf112ef8b8236089ecd512c3fc813b244c1691cf

Observation faf511ce-2e13-437b-8378-3041049502fe · inbound

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces cites this paper.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:09.508644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:09.508644Z digest=sha256:2602cc2285f29e18d7d6c4d130171008967ef1a17656ba0467848d1fb5b4b9f6

Observation 45b1c30d-60b4-4e2d-9249-f678a1b45626 · inbound

EffiVLM-BENCH: A Comprehensive Benchmark for Evaluating Training-Free Acceleration in Large Vision-Language Models cites this paper.

EffiVLM-BENCH: A Comprehensive Benchmark for Evaluating Training-Free Acceleration in Large Vision-Language Models Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:03.480920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:09:03.480920Z digest=sha256:872814bbba34fb667d2bbb248f23b3cca09f9914f6a12d6272747dc98f9aca18

Observation b01b931d-9111-4054-a48e-48d08084632b · inbound

MedBookVQA: A Systematic and Comprehensive Medical Benchmark Derived from Open-Access Book cites this paper.

MedBookVQA: A Systematic and Comprehensive Medical Benchmark Derived from Open-Access Book Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:02.994137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:02.994137Z digest=sha256:2fe2325d95816c5e971f897ebd4855d27762907e44fe976c5e2bd5251d95a5d3

Observation 765a8170-44c3-4b5f-8f69-cb30a2eebca8 · inbound

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding cites this paper.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:06.001743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:06.001743Z digest=sha256:e75ab4b5bd4caa56cf8267705535a67c1cc424da31a66aff80fdddc825bb8f8d

Observation c1723077-1fe6-4428-b60f-2d9391758425 · inbound

NavBench: Probing Multimodal Large Language Models for Embodied Navigation cites this paper.

NavBench: Probing Multimodal Large Language Models for Embodied Navigation Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:34.368716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:34.368716Z digest=sha256:68a382ccce09f7fd6aa0e25d7d0bec0c3acb87b84099b51d58b7e11b0e35504e

Observation 0e6fe2f7-205c-4d23-bd7a-9b9e7cd2f583 · inbound

Abstractive Visual Understanding of Multi-modal Structured Knowledge: A New Perspective for MLLM Evaluation cites this paper.

Abstractive Visual Understanding of Multi-modal Structured Knowledge: A New Perspective for MLLM Evaluation Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T11:51:14.988403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:51:14.988403Z digest=sha256:1b6dcdfe2c9f4e9c160482c449be9f6497efbcc966ba055e7dfb8fbf1b55e827

Observation b65b9093-4a76-49dc-8d33-1b46a11cc01e · inbound

ReAgent-V: A Reward-Driven Multi-Agent Framework for Video Understanding cites this paper.

ReAgent-V: A Reward-Driven Multi-Agent Framework for Video Understanding Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:51:59.500410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:51:59.500410Z digest=sha256:accb72740e9e4223cd7444f6c7579300bfa57b3d97eec2a844f72b8eb893613a