Pith. sign in

Paper Citation Record · LEDGER

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models

As of 7 August 2026, this Paper Citation Record lists 100 of 137 outbound references and 2 inbound Pith citation observations for arXiv:2507.12566.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.12566 v1

Coverage vector

measured 100 of 137 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T16:50:05.568237Z

measured 102 of 102 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-13T05:12:37.339084Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-13T05:17:18.750827Z

Reference resolution

100 of 137 outbound references displayed

  • verified exact2
  • verified fuzzy0
  • unresolved98
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 235b395d-36c4-4a87-86a8-8c3e66eb72d8 · outbound

This paper cites GPT-4 Technical Report.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:55.260899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:55.260899Z digest=sha256:362b57ec5018ba58d04eec2ac9dc635b546a723aa308405f49b46d92dc7a5362

Observation d9774078-184d-4460-83d9-23b4956496b9 · outbound

This paper cites Qwen Technical Report.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Qwen Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:55.431651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:55.431651Z digest=sha256:7611f5b883bf7134301c4c4ae00ca746a617cf5477cf2843fa48fc449de24783

Observation b0b86d26-1227-476b-8a16-ea8ab9088e16 · outbound

This paper cites InternLM2 Technical Report.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models InternLM2 Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:55.597995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:55.597995Z digest=sha256:950085ee050ff2a4f3256f085edb80923314aa2c12ec8d1cd0c4d37088067674

Observation c183d979-6b75-4a40-bd9a-0dafa73d6666 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Learning transferable visual models from natural language supervision,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:55.726783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:55.726783Z digest=sha256:47cadf0b4d7cffc4a69ced89aaf265ca44dde0182c91734218812031e09f52b6

Observation e0901f71-40b2-4641-908e-c290254308de · outbound

This paper cites Visual instruction tuning,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Visual instruction tuning,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:55.834849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:55.834849Z digest=sha256:20dd8e5bb00990665ca82bca56c9698ac644f62b1cf9655161a679a19488eaad

Observation 3ae42cdb-16e2-4436-966e-725c985a3d53 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:55.905181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:55.905181Z digest=sha256:b12fa10777d9eff82c6c75c38a15b19e311126a39e0ccad26549896e1a9da2c0

Observation ac57a490-84dc-4135-9514-b86f724a33e8 · outbound

This paper cites BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:56.005313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:56.005313Z digest=sha256:5d2b3eb75e8a24dcd45158533cb280f5081a2580262c733b91eb2123b0fc9f1d

Observation 9c7033a1-d62d-4e9a-9189-9f5e541158b9 · outbound

This paper cites Introducing our multimodal models,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Introducing our multimodal models,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:56.084702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:56.084702Z digest=sha256:2317c7583e6a08a459fe0d9f4f895bfdce5e84853e80348e9a89d514bc7ece7d

Observation 31e4b70e-3842-4761-9ffd-58b30ea63fba · outbound

This paper cites Unveiling Encoder-Free Vision-Language Models.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Unveiling Encoder-Free Vision-Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:56.218152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:56.218152Z digest=sha256:02906f58232096847ba349dbd92486683f7f797779887bd0ca733b1b5d5d80db

Observation 1c525bd1-5ec4-4b4f-9005-06ab4929f066 · outbound

This paper cites SOLO: A Single Transformer for Scalable Vision-Language Modeling.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models SOLO: A Single Transformer for Scalable Vision-Language Modeling

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-08-06T16:50:10.351624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:49:56.373596Z digest=sha256:c2a539a4df2a0c9e94f079a45c6e13e9f71a19d8e6682a1c63ac37a248ce1c5d

Observation 23b62d2c-270b-4f8c-8d6d-96db9f90abcd · outbound

This paper cites Mono-internvl: Pushing the boundaries of monolithic multimodal large language models with endogenous visual pre-training,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Mono-internvl: Pushing the boundaries of monolithic multimodal large language models with endogenous visual pre-training,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:56.460558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:56.460558Z digest=sha256:9f8f9dde54c0fe41c1fd2370fedd77a640ab93875bfd672867aaf57f18625d23

Observation d4dfeccd-4672-44a9-8db0-229275532d35 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:56.572596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:56.572596Z digest=sha256:4fb8c116f50893daeed4674137d05279b3b80e9d150ca288cc7813c75274882d

Observation df089f74-f30b-44f3-a10f-2de6df635cdf · outbound

This paper cites Investigating the Catastrophic Forgetting in Multimodal Large Language Models.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Investigating the Catastrophic Forgetting in Multimodal Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:56.665670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:56.665670Z digest=sha256:f5950c858eb60cf67f7ee6b154778a043810ec1bcb26fcd1c54935cb3a3bd22c

Observation 2cffd4b2-6780-4103-aa7f-85eca601f8e6 · outbound

This paper cites Delta Tuning: A Comprehensive Study of Parameter Efficient Methods for Pre-trained Language Models.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Delta Tuning: A Comprehensive Study of Parameter Efficient Methods for Pre-trained Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:56.785510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:56.785510Z digest=sha256:3aeee1f50b87055634c7b4aa7be94dd781ab4c5fd87d5466f2553674e7e0935d

Observation 73aee5ed-f683-4926-a808-79f3ecb4aa9b · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:56.896912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:56.896912Z digest=sha256:3d124ce9b776c01a5e2c6fbe71c84f057e051129478e13647ceee4d0c24fed6a

Observation 93868ee2-dba0-452c-8fd9-2d19ae9d9976 · outbound

This paper cites Lima: Less is more for alignment,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Lima: Less is more for alignment,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:56.999030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:56.999030Z digest=sha256:0a3f32af3a44ee022b62597a61ce11eb66e3e92e296f4397fe2205e977bcc1f5

Observation 27ad4242-fc85-46d9-a397-83c596813eed · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Emu3: Next-Token Prediction is All You Need

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:57.069428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:57.069428Z digest=sha256:3fc175ea5e0b83f65861ca850c1fd7369bffc4758c961e0c8c698de82589a68b

Observation 901c8ced-76a7-47a9-93e0-abdc860ade0e · outbound

This paper cites The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision).

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision)

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:57.153253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:57.153253Z digest=sha256:a413b21c2ecbfbb7bd444798ac41e0db92f8a8cdfcec10ed398a1c57d3b4de67

Observation 4c338e29-89b5-40fe-980e-4969d4907720 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Gemini: A Family of Highly Capable Multimodal Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:57.248846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:57.248846Z digest=sha256:8b3f15cc63fd1fb78ae267cc8f695d5fc1f0b50382a1248de01e3704766dd2c3

Observation 220026a1-5a60-44c4-8099-462afadca072 · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Improved Baselines with Visual Instruction Tuning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:57.353167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:57.353167Z digest=sha256:62969d5afac06207681856e019227032737d9c9bf915457d0b559b258bad4b47

Observation 0f625dc8-cdfc-42c1-9448-9668a4a65cfd · outbound

This paper cites Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:57.463091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:57.463091Z digest=sha256:f0824d1c35c0ebe627b2882ca8753b8428c00bc465f26471eeca38253172214d

Observation cc75399e-c724-46e9-8338-dbd706c97401 · outbound

This paper cites Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-06T16:50:10.051759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:49:57.683627Z digest=sha256:88008221288b63e7777d93f0b7dc3d9a83a536f5c48dd5cf543609b9c978d62c

Observation 718b4b77-42d6-4f16-baa8-b702afb4b3c8 · outbound

This paper cites SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:57.751825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:57.751825Z digest=sha256:9afab051b10c54511f687bd947ddd843cb5bfd600175f78b0ad2f8ea41bbf97d

Observation f5a625a0-61af-40c6-b69e-ed10033673fa · outbound

This paper cites BLIP: bootstrapping language-image pre-training for unified vision-language understanding and generation,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models BLIP: bootstrapping language-image pre-training for unified vision-language understanding and generation,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:57.862972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:57.862972Z digest=sha256:5b25162807f65c03a3ed5521abd199744fb22fa672f18df6a32c739d4b56b78c

Observation b6ab71c8-d911-40a6-8d89-4ef5fba23175 · outbound

This paper cites Instructblip: Towards general-purpose vision-language models with instruction tuning,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Instructblip: Towards general-purpose vision-language models with instruction tuning,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:57.985266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:57.985266Z digest=sha256:8467b8dec5d63cac0753539ed4328f9ea93263953efe1a695c50ab6faa2c0105

Observation dab453f1-892f-41b9-9259-186e59988c55 · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Llava-next: Improved reasoning, ocr, and world knowledge,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:58.074139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:58.074139Z digest=sha256:202d64e290f3c8f24064b4204e9b473ae071ebea1cc51c03623918dcbf5e7f8b

Observation f68fdac2-7f06-4dfa-a0b8-6f5c56f104dd · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:58.168347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:58.168347Z digest=sha256:c50289270edd2a98fa10de6d184100ae2dbe98d22474a92efa4b11b91b466492

Observation df391df8-6293-45c1-b131-2b8d35f1e3a7 · outbound

This paper cites Qwen2.5-VL Technical Report.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Qwen2.5-VL Technical Report

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:58.311277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:58.311277Z digest=sha256:ac9791658a1f5ba4d78da01fe79ca5e5c855b3be442271d31b618710d5d821ec

Observation 202f1580-fc12-439a-8af0-3029d2195b70 · outbound

This paper cites InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:58.390094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:58.390094Z digest=sha256:a3f924b5fae97d584060c5666257c00eda7a86a66b335bfd6cee2ff08f05735b

Observation 871361c5-bde1-4422-8f1b-ebb842a6a0b1 · outbound

This paper cites Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:58.597304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:58.597304Z digest=sha256:df857cde1fdf8c143ee2ece021abb13c2c78cf58335d2c6edf2ca0242e2f771d

Observation 8eba6607-61ef-454b-a467-b2be0c2d9381 · outbound

This paper cites Mini-internvl: a flexible-transfer pocket multi-modal model with 5% parameters and 90% performance,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Mini-internvl: a flexible-transfer pocket multi-modal model with 5% parameters and 90% performance,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:58.697634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:58.697634Z digest=sha256:db6529866ad7f46827bcb026967a313c6b32c6863ba4c793f11d31b61b94fd82

Observation 63a1e2a6-3313-4851-9e28-bdd64ec2e4a3 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models LLaMA: Open and Efficient Foundation Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:58.800042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:58.800042Z digest=sha256:c2cb40510b1d14f9e1267cdf76947eaf1b8a9e55838cf535df9547f79cf9c125

Observation 730f0235-ff65-47e0-b3e4-822e238e42da · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:58.881346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:58.881346Z digest=sha256:27d392dc47936a0a51b1e5f0d2a90885640ff820292bdb531a6c05ecf471415d

Observation 0244e7a9-46d2-49e3-86a4-682e461f0b46 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models An image is worth 16x16 words: Transformers for image recognition at scale,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:58.977983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:58.977983Z digest=sha256:ce15c728dbea8916bdcd195e85d386215a4da52551d9f2223bdcfec8e045dd32

Observation 7e7d78cb-6aef-4d4a-8d55-1fea0f4b1312 · outbound

This paper cites Show-o: One Single Transformer to Unify Multimodal Understanding and Generation.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:59.055806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:59.055806Z digest=sha256:148a4a0341658b3d0fa10641c33d59172ce695298be47f3222494fad32d477b0

Observation bed0b7ad-391e-4c43-bca4-ac6c35dda569 · outbound

This paper cites Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:59.135515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:59.135515Z digest=sha256:9cc4ed6b721f8b914d792ec144279b2bb46b88021e0eeda33723f21cee49bdb1

Observation 2dde65db-d26c-4439-b13c-a083fc6b3b8d · outbound

This paper cites Vlmo: Unified vision-language pre-training with mixture-of-modality-experts,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Vlmo: Unified vision-language pre-training with mixture-of-modality-experts,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:59.224871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:59.224871Z digest=sha256:59e247996f59a31ca50a9a817326c2bfb0cc774af26cb10baf0b25e6f9519414

Observation 1fe92c26-84cd-40aa-97f4-5e7386379e26 · outbound

This paper cites Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:59.336128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:59.336128Z digest=sha256:7362a72390baca00f9cdf6095b73ec711fda62260d732b021a5df86cb68e9228

Observation 3e3f5af5-d0c2-4d93-89e5-8a1fe41310a5 · outbound

This paper cites Scaling Vision-Language Models with Sparse Mixture of Experts.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Scaling Vision-Language Models with Sparse Mixture of Experts

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:59.424967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:59.424967Z digest=sha256:c8e88ab932badfb9afa2959e6a76cddf6cf370f0f15200e407df728afdc364d0

Observation f5284dae-5a57-4d0b-9d5d-eb743de78905 · outbound

This paper cites Twenty years of mixture of experts,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Twenty years of mixture of experts,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:59.548683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:59.548683Z digest=sha256:3c7c6ce2620c26488a3ace42244cb7c785e9226e741eb9abd8c4247a26e45a4c

Observation 0a5d6262-037d-431e-83b5-cee3b08144d0 · outbound

This paper cites MoMa: Efficient Early-Fusion Pre-training with Mixture of Modality-Aware Experts.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models MoMa: Efficient Early-Fusion Pre-training with Mixture of Modality-Aware Experts

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:59.656307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:59.656307Z digest=sha256:e1799cc1ec7160dbbde84e71e7fb850907593c33cf6c3a63dd163f6c27b60a7e

Observation fc70db5e-76e7-4e1b-af06-22a874c99c41 · outbound

This paper cites Mixture-of-Depths: Dynamically allocating compute in transformer-based language models.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Mixture-of-Depths: Dynamically allocating compute in transformer-based language models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:59.760100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:59.760100Z digest=sha256:07cb482e77023cb4632800cf00f25ec81e151d77b18526cb90ed0c3612416f5c

Observation 5512c647-879c-48fd-b611-5cd6e6d1ced4 · outbound

This paper cites Aria: An Open Multimodal Native Mixture-of-Experts Model.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:59.827141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:59.827141Z digest=sha256:ee9dc8d8ac8fc77590d22f60cd6d3ba4b2f88c7778498c036ad7aed25a17074b

Observation 855a26ff-2ca3-4c63-a529-07a4d654281f · outbound

This paper cites Attention is all you need,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Attention is all you need,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:59.936912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:59.936912Z digest=sha256:6f14a7c3f6522f3ecd7c6c3fce23546ecc328dc5133ee0b849995724a7caabf0

Observation eb7f29c9-3b3a-430d-98aa-32245d0eae88 · outbound

This paper cites Root mean square layer normalization,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Root mean square layer normalization,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:00.080685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:00.080685Z digest=sha256:3d260df1711690475d2c6fb8b1b82d5b2650053bf65acb8856b5ee7c966d3de9

Observation 3cec4baf-2c06-4943-b6c5-51583007169e · outbound

This paper cites Laion-5b: An open large-scale dataset for training next generation image-text models,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Laion-5b: An open large-scale dataset for training next generation image-text models,

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:00.195147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:00.195147Z digest=sha256:849e5c0cd5337a8341b8a8a148209e9520f41758daff5737981790bda05d59d1

Observation 7cdd1300-3cee-46af-9e71-46005c45e6b0 · outbound

This paper cites Coyo-700m: Image-text pair dataset,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Coyo-700m: Image-text pair dataset,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:00.302668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:00.302668Z digest=sha256:6c4a8ad761b6901097ca2529a8dfc3b8fb9e6303831f251a93d7c5a0b7ca40f6

Observation 55bf65fe-a0b4-424d-a3f3-aa0da6d66bbf · outbound

This paper cites Segment Anything.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Segment Anything

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:00.388194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:00.388194Z digest=sha256:0d5924d2ec4f320a589a1614216fe0ffc2037aae0c1237e0b6282fa8436b6085

Observation b8d07cf3-fba4-4f8d-9360-55cb2ec9d4a8 · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:00.517198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:00.517198Z digest=sha256:78289d83688429e52fb2f499ba300c91c8af51b647ac5b35cfa7e9f13922f5f2

Observation ffe36186-4645-44ad-8511-f3eb6f815ab9 · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:00.619189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:00.619189Z digest=sha256:8a1bc6abbb1181593924d2d542039a13ac609f2e76704b222ee72714438e26df

Observation 712be2aa-2719-4a1b-b345-2678d7220474 · outbound

This paper cites Textcaps: A dataset for image captioning with reading comprehension,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Textcaps: A dataset for image captioning with reading comprehension,

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:00.734000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:00.734000Z digest=sha256:cd436668494a4275498d215aee47eac3aee1ade646404275a49d9a99972cbfcf

Observation 86908a5d-546f-45a5-9e28-ddc473bd2a39 · outbound

This paper cites Objects365: A large-scale, high-quality dataset for object detection,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Objects365: A large-scale, high-quality dataset for object detection,

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:00.805492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:00.805492Z digest=sha256:a1b21d8a45243d1f80b02c6443ffa2983a3b0b983867eacb6a343c758e1b8871

Observation 776abb4f-f5de-465b-9e53-745ae1e26aa0 · outbound

This paper cites The all-seeing project: Towards panoptic visual recognition and understanding of the open world,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models The all-seeing project: Towards panoptic visual recognition and understanding of the open world,

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:00.874286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:00.874286Z digest=sha256:052b06ad0273c58bf0a11c5050b6a08f57eba2529bc0258a0aecf595c0ee1411

Observation ba30b0f4-3e61-476c-b798-e2e44d305b27 · outbound

This paper cites Wukong: A 100 million large-scale chinese cross-modal pre-training benchmark,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Wukong: A 100 million large-scale chinese cross-modal pre-training benchmark,

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:00.947847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:00.947847Z digest=sha256:281069267600190dac77c45f42a34cdb36008cd7398bbb7869f3f44c0959ca12

Observation a0b97b9e-f3a1-473d-bb14-7c57e43644a3 · outbound

This paper cites Laion coco: 600m synthetic captions from laion2b-en.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Laion coco: 600m synthetic captions from laion2b-en

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:01.031594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:01.031594Z digest=sha256:3c5828067b7ee12a900108a4264066b8c0a7cbc403f3cc6f77c22d6c41bbc058

Observation 4ea5e301-bf72-41f3-ab76-9789c6925236 · outbound

This paper cites MMC: Advancing Multimodal Chart Understanding with Large-scale Instruction Tuning.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models MMC: Advancing Multimodal Chart Understanding with Large-scale Instruction Tuning

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:01.162563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:01.162563Z digest=sha256:b8f551e935abd933e15be1e7619a9d653e2e6546ee95b385d55004b5d0648b2d

Observation d2e3ed08-d8f3-43f3-a7ca-f6d5c4d79193 · outbound

This paper cites Icdar 2019 competition on large-scale street view text with partial labeling-rrc-lsvt,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Icdar 2019 competition on large-scale street view text with partial labeling-rrc-lsvt,

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:01.240559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:01.240559Z digest=sha256:46cea1cae8d1e5bc2de7b836dfa157372eb0bde590cf6c07df89eccb30fea029

Observation 4912a34c-18da-4fa0-97a8-d82f53569c69 · outbound

This paper cites Scene text visual question answering,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Scene text visual question answering,

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:01.315416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:01.315416Z digest=sha256:d23e455edcca229f4d81e2567bf0e91c5f196cca193169d8a4c3a78d4e5e56a5

Observation c7493250-004c-4ba9-b345-df73855b3bcb · outbound

This paper cites Icdar2017 competition on reading chinese text in the wild (rctw-17),.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Icdar2017 competition on reading chinese text in the wild (rctw-17),

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:01.390445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:01.390445Z digest=sha256:3ba796eff7925e6417804a2ae1187c2009b279ad03cbd3df344d8699715101ac

Observation 66f3aa88-0aee-4f2e-bd2b-d4dffbf4ddec · outbound

This paper cites Icdar 2019 robust reading challenge on reading chinese text on signboard,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Icdar 2019 robust reading challenge on reading chinese text on signboard,

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:01.495683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:01.495683Z digest=sha256:4a2851695fc0ad95513cb57565341e10aa55659ca1b72bfb9beb7ad1fe7b8ac2

Observation b330ee7f-0072-4b50-895e-b9ee0691b6ea · outbound

This paper cites Icdar2019 robust reading challenge on arbitrary- shaped text-rrc-art,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Icdar2019 robust reading challenge on arbitrary- shaped text-rrc-art,

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:01.578838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:01.578838Z digest=sha256:6289ec04e1718be51f1ae36aca18a45d6db8f6ad95068d0f1f208389d6611f50

Observation 3756c26d-5fd7-4b25-929a-db601c9dbe4f · outbound

This paper cites Ocr-free document understanding transformer,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Ocr-free document understanding transformer,

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:01.712181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:01.712181Z digest=sha256:ed14a4035a68b2af5a2b3e65db15a47ca5f5fc384e0105bef94615294a3709ce

Observation 0ec19cce-6c0a-4f41-be1e-fd52bd161f61 · outbound

This paper cites COCO-Text: Dataset and Benchmark for Text Detection and Recognition in Natural Images.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models COCO-Text: Dataset and Benchmark for Text Detection and Recognition in Natural Images

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:01.817911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:01.817911Z digest=sha256:7616711d520ec6edb9370f2c634cefc391675b62998e360f142038b5de84fc1d

Observation 7c8d6861-01c7-4da6-9a3c-1890fc2068bc · outbound

This paper cites Chartqa: A benchmark for question answering about charts with visual and logical reasoning,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Chartqa: A benchmark for question answering about charts with visual and logical reasoning,

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:01.924520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:01.924520Z digest=sha256:aa53af5dcdcb0af2cf42562b17ff35db0276487cee90a6a75c4164af2db35d2f

Observation 61e6366c-974c-407f-8fcb-4b8263ade4fe · outbound

This paper cites A large chinese text dataset in the wild,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models A large chinese text dataset in the wild,

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:02.002148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:02.002148Z digest=sha256:efb779830b48e8b78fb95614ceed11484beedfa313ce9c47fcf60d0d758fc59f

Observation d043d554-ecdf-429c-88da-cf42cc6088d7 · outbound

This paper cites Simple and effective multi-paragraph reading comprehension,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Simple and effective multi-paragraph reading comprehension,

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:02.099560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:02.099560Z digest=sha256:30ce9c1f2f1dabcb0504c9eb653dde9697fc3ea98e09dd9b512cfa0115970b4a

Observation 0b0c7140-c968-45e5-a601-f6060c883f63 · outbound

This paper cites Textocr: Towards large-scale end-to-end reasoning for arbitrary-shaped scene text,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Textocr: Towards large-scale end-to-end reasoning for arbitrary-shaped scene text,

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:02.203648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:02.203648Z digest=sha256:ffdc25f80b34a3b2776826bf0314504f88cc74bdc281f0df64f9a03a3cf70c4a

Observation 07cde5a0-af22-42d9-a8cd-9c43ab5f7c76 · outbound

This paper cites Plotqa: Reasoning over scientific plots,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Plotqa: Reasoning over scientific plots,

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:02.320729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:02.320729Z digest=sha256:6644c5853cd3d89163f14ec9a66f5fb2961e9374edfec73e681bbe7f7ba07bf3

Observation c83be487-e7ae-427b-bf7a-9f520ec9f738 · outbound

This paper cites Infographicvqa,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Infographicvqa,

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:02.423626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:02.423626Z digest=sha256:e32ed482c9af1e85e8131921fedf9f2fce16ee78afb50f64fcda6424d30a17ce

Observation 66675f87-dc47-41fa-b6ab-4ab661215fed · outbound

This paper cites Making the V in VQA matter: Elevating the role of image understanding in visual question answering,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Making the V in VQA matter: Elevating the role of image understanding in visual question answering,

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:02.536884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:02.536884Z digest=sha256:8cb383358a8d651bcb96b5d54a7151fc1394020fd63bf4b0c108ea857a819add

Observation 546425af-4490-472f-8619-4c156080f53f · outbound

This paper cites GQA: A new dataset for real-world visual reasoning and compositional question answering,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models GQA: A new dataset for real-world visual reasoning and compositional question answering,

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:02.641238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:02.641238Z digest=sha256:d0da97d1e1d78ca4f78db8f8c7acafd363273bfea409c97fb168a24a83db620d

Observation f900c643-6726-47c9-aa9f-549eb9f80e8f · outbound

This paper cites Ok-vqa: A visual question answering benchmark requiring external knowledge,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Ok-vqa: A visual question answering benchmark requiring external knowledge,

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:02.746998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:02.746998Z digest=sha256:0fc4f609d0d5963c3e0b568bbc3d720e52c9fbdee209cd18a94a9989231dfc27

Observation 2bc39bd7-4222-4de2-9972-5cd6c418f854 · outbound

This paper cites Visual spatial reasoning,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Visual spatial reasoning,

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:02.852857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:02.852857Z digest=sha256:a905a3dd84aecd945f484ae2e7ca485a2db9d1d1f99e2aac430e6a419d3c5678

Observation d9264504-9076-436e-8b1e-bc1716e767a3 · outbound

This paper cites Visual dialog,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Visual dialog,

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:02.963424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:02.963424Z digest=sha256:416745494c1c5fd139b5e26f91dcac0d86041de350857a1d44ab3f02597dc655

Observation eeba70c5-126f-4a87-bc5f-c70595bee6fd · outbound

This paper cites A diagram is worth a dozen images,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models A diagram is worth a dozen images,

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:03.027579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:03.027579Z digest=sha256:60486c40c9b9beaebbfd2e7f9f8b7df42bea3e0e617217a64a1889d9f6dc39a5

Observation c19cef38-137e-4204-beaa-d73fa8976722 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Learn to explain: Multimodal reasoning via thought chains for science question answering,

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:03.153797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:03.153797Z digest=sha256:7cd7dd6d1b321edd7446a9ca87c700f7daa15490f4f60a831daaca0ff104bd89

Observation 4b270660-b3fa-4df9-b4fd-4cbb7a7e6ffd · outbound

This paper cites Are you smarter than a sixth grader? textbook question answering for multimodal machine comprehension,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Are you smarter than a sixth grader? textbook question answering for multimodal machine comprehension,

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:03.236776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:03.236776Z digest=sha256:974ea35b812c2dc8da2d7e67eb7a22b9ad34ffe7b570c5a682139965d3279af9

Observation 28539a3e-357e-43f1-a2e4-1a229e7f10a5 · outbound

This paper cites Dvqa: Understanding data visualizations via question answering,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Dvqa: Understanding data visualizations via question answering,

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:03.384359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:03.384359Z digest=sha256:30cd50e87e184a1ad462031127788e4cbd2c236f570baaa9478b4c38f02a978c

Observation 3f8cb091-aeeb-40d2-97cc-477d8c8b0a9f · outbound

This paper cites Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:03.529613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:03.529613Z digest=sha256:839f837a14f7bb6e62b79aabd77be3609d52c9921410de2c57549c7811831324

Observation 7382f2c5-4d5a-431e-bb17-b95211031b07 · outbound

This paper cites An augmented benchmark dataset for geometric question answering through dual parallel text encoding,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models An augmented benchmark dataset for geometric question answering through dual parallel text encoding,

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:03.654147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:03.654147Z digest=sha256:b6c957cbde62888f3e4f1d3c2375ba0f2f2eef71feb49796f962eb00caf60edf

Observation f8b413f7-7d5a-45e3-8655-2a7e0c828785 · outbound

This paper cites Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:03.783517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:03.783517Z digest=sha256:ce4d27c0b64f6f9cd29c7f5017b36e62424d590dbee95d8924a25fbdca0c2541

Observation 33c867b7-c587-4172-a55b-d3a85925e403 · outbound

This paper cites MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:03.926291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:03.926291Z digest=sha256:d58a586c3252c863f53ee1acde55d4bfd151d5fd66ac7e10829d249384c7d727

Observation 65f50570-1ed8-4434-a038-0a613d7da31b · outbound

This paper cites CLEVR-Math: A Dataset for Compositional Language, Visual and Mathematical Reasoning.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models CLEVR-Math: A Dataset for Compositional Language, Visual and Mathematical Reasoning

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:04.072385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:04.072385Z digest=sha256:38a06fdbc6701e1c4e99947c0524261c87a62fb01265866a63c99ef96381985c

Observation 26aebb84-91f1-4696-a69d-57638124a7e8 · outbound

This paper cites Super-clevr: A virtual benchmark to diagnose domain robustness in visual reasoning,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Super-clevr: A virtual benchmark to diagnose domain robustness in visual reasoning,

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:04.211430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:04.211430Z digest=sha256:eca29babcdcb44d0721c8c6696c430559c6ebb217ffb02fda1717c786774e25c

Observation 3eaa7f89-402a-4e8d-bff3-c1d1320a9647 · outbound

This paper cites Inter-GPS: Interpretable Geometry Problem Solving with Formal Language and Symbolic Reasoning.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Inter-GPS: Interpretable Geometry Problem Solving with Formal Language and Symbolic Reasoning

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:04.330336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:04.330336Z digest=sha256:57248f4001c0c63b8b7efe6e40238d0a069745f84bcd564f3d5fa71acfdb81c2

Observation 9408ddf6-13f9-46d8-9640-f0bca390a041 · outbound

This paper cites Kvqa: Knowledge- aware visual question answering,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Kvqa: Knowledge- aware visual question answering,

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:04.443260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:04.443260Z digest=sha256:b2483a9e7151a1378566c66ac0be734073f49a5ce1f0c83c3adca4166b6e637d

Observation e2063723-1016-463b-a2ca-3d5ac2691921 · outbound

This paper cites A-okvqa: A benchmark for visual question answering using world knowledge,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models A-okvqa: A benchmark for visual question answering using world knowledge,

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:04.560616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:04.560616Z digest=sha256:e4707d5371dd6c2ce2aaa338d63037de0bcb1caceeff8cd26fcc5e0e84deb324

Observation 197b7174-653d-4cb5-9496-32b7ee46a3a6 · outbound

This paper cites Viquae, a dataset for knowledge- based visual question answering about named entities,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Viquae, a dataset for knowledge- based visual question answering about named entities,

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:04.660448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:04.660448Z digest=sha256:f0332073cb9ef19b55ca82ce62022853b5ce8c8016af0974e7ab9cc9490ce88e

Observation 3d17bb68-d896-410f-921f-b62d7d2c94aa · outbound

This paper cites WanJuan: A Comprehensive Multimodal Dataset for Advancing English and Chinese Large Models.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models WanJuan: A Comprehensive Multimodal Dataset for Advancing English and Chinese Large Models

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:04.722409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:04.722409Z digest=sha256:5aa9409a46b67e9f866c6743afac335b3f7419835267ed9e25575d515bf1cfef

Observation 7cb3cda1-178f-4ead-a61f-9f8600364ab3 · outbound

This paper cites Ocr-vqa: Visual question answering by reading text in images,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Ocr-vqa: Visual question answering by reading text in images,

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:04.785021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:04.785021Z digest=sha256:caef25dc15f1659e03ad7de6dcf028958b9a0349997ae559a4d41ec1754d1a40

Observation b50417c2-8ed8-42df-b13a-983ad4fd2128 · outbound

This paper cites Towards VQA models that can read,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Towards VQA models that can read,

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:04.861787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:04.861787Z digest=sha256:83e602e5b770dc9986ecda480cabe8e958be6a090d51bdd39376002af3e140c0

Observation 581d54b1-e307-452a-bfe9-6335283eaa70 · outbound

This paper cites Modeling context in referring expressions,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Modeling context in referring expressions,

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:04.943013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:04.943013Z digest=sha256:6e61bc9c420a5f715db80c0e82f019f08ce471156353c084b89eb5cf87f06cdd

Observation 8f850fd4-aafe-4cb7-9acd-c1616efcf848 · outbound

This paper cites Generation and comprehension of unambiguous object descriptions,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Generation and comprehension of unambiguous object descriptions,

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:05.027353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:05.027353Z digest=sha256:c514f9f57e231c0f6be4745ea27fee6afd62235a608bf38fb04ed3b22f26e85f

Observation ed4402f5-07e4-43c4-bc10-7c2124596bd3 · outbound

This paper cites Visual genome: Connecting language and vision using crowdsourced dense image annotations,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Visual genome: Connecting language and vision using crowdsourced dense image annotations,

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:05.105073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:05.105073Z digest=sha256:59bd20bfe2ca2108630113ea881f7908bcbdf639ba51fbbcaf674ead231667f3

Observation b970ecbc-9311-4a1b-bd3c-f83feadd74b4 · outbound

This paper cites To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:05.174737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:05.174737Z digest=sha256:e6d1f952acb0bf616f2c2b6c1090b08b0598dc96c2c77087b9aaa032f04b6a73

Observation 244696a9-e355-4869-9b6a-18fa1cf73075 · outbound

This paper cites ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:05.268162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:05.268162Z digest=sha256:05bab3a984b03198d007f0a818c15cc1c53ecb7b7e8573378d7c16b111031206

Observation 65664252-de60-4c8d-b58a-736bf1bdcfb6 · outbound

This paper cites Gpt-4v dataset,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Gpt-4v dataset,

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:05.335533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:05.335533Z digest=sha256:cff8d937997c532bee0d0553764a06fc84446f2a24620a8e2ad5d368972c282d

Observation 6d38fdb5-0f7b-4210-b379-69b149ae4558 · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Judging llm-as-a-judge with mt-bench and chatbot arena,

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:05.406294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:05.406294Z digest=sha256:8c377cae19ab80dd2a454a3488353479b78f9e5004ac9db3c5da619ee9b12419

Observation d86c7702-3db9-44f1-ab53-65377a19e076 · outbound

This paper cites SVIT: Scaling up Visual Instruction Tuning.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models SVIT: Scaling up Visual Instruction Tuning

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:05.494156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:05.494156Z digest=sha256:8c2cad6626c8020495a4e551e47f5fb284b90b4e1b64df67b758e5a3fbaa95c2

Observation 24e6065c-6e91-4ccd-ade8-cada6d4633bf · outbound

This paper cites Openhermes 2.5: An open dataset of synthetic data for generalist llm assistants,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Openhermes 2.5: An open dataset of synthetic data for generalist llm assistants,

Reference 101

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:05.568237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:05.568237Z digest=sha256:c194f17f0351d7a44b76e80cdb341079504d7e1b9c5c7995bf7d6390a1d1acaf

Pith citing papers

Observation fe67a2ae-8c0f-4042-abb3-5246f5c6fee4 · inbound

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency cites this paper.

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:58:59.202761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T11:58:58.660564Z digest=sha256:014eea322dceb0e8f7183e6241e5c9eba90697888701deb40a3d279fcf163a4d

Observation 85ce1315-55db-4859-9f18-b28e63813223 · inbound

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture cites this paper.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models

Reference 88

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:17:18.752838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:e0ec923e5f7b9633e169ec3b416b63358df82fdda42808268d33a6fc0577af5f