Pith. sign in

Paper Citation Record · LEDGER

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models

As of 14 August 2026, this Paper Citation Record lists 100 of 137 outbound references and 2 inbound Pith citation observations for arXiv:2507.12566.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.12566 v1

Coverage vector

measured 100 of 137 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T16:50:05.568237Z

measured 102 of 102 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-13T05:12:37.339084Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-13T05:17:18.750827Z

Reference resolution

100 of 137 outbound references displayed

  • verified exact2
  • verified fuzzy0
  • unresolved98
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 235b395d-36c4-4a87-86a8-8c3e66eb72d8 · outbound

This paper cites GPT-4 Technical Report.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:55.260899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:55.260899Z digest=sha256:b21c45bcd70bd04fe61d36b213603179e85572f131ce8a0ac93fe549111a54ee

Observation d9774078-184d-4460-83d9-23b4956496b9 · outbound

This paper cites Qwen Technical Report.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Qwen Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:55.431651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:55.431651Z digest=sha256:bd2b7b8e6a79b23aa957d70315de368f39f565554e7dfe91c210e11c163e5619

Observation b0b86d26-1227-476b-8a16-ea8ab9088e16 · outbound

This paper cites InternLM2 Technical Report.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models InternLM2 Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:55.597995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:55.597995Z digest=sha256:10ec09ddbdd2ed7ff1c047f0104cdb98202f4ef252d5fb16a6b917b63cca3828

Observation c183d979-6b75-4a40-bd9a-0dafa73d6666 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Learning transferable visual models from natural language supervision,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:55.726783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:55.726783Z digest=sha256:af2e2623a3890b7764d159a2978a09d9b8b006a834aafc348c7ba9472be5e0de

Observation e0901f71-40b2-4641-908e-c290254308de · outbound

This paper cites Visual instruction tuning,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Visual instruction tuning,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:55.834849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:55.834849Z digest=sha256:cac10ea3285a80a8779fd5d7a56d5587831c80fa1d0c4401c8deaf871d65d7c4

Observation 3ae42cdb-16e2-4436-966e-725c985a3d53 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:55.905181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:55.905181Z digest=sha256:294e4cdcdda17af1dde9a5bc8638761bc18a7154c1c96474b0b89a832bf192e5

Observation ac57a490-84dc-4135-9514-b86f724a33e8 · outbound

This paper cites BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:56.005313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:56.005313Z digest=sha256:77717ffbc3564a5907788107def3014db142600fc39eb2e6a72e5474f7aa8de8

Observation 9c7033a1-d62d-4e9a-9189-9f5e541158b9 · outbound

This paper cites Introducing our multimodal models,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Introducing our multimodal models,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:56.084702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:56.084702Z digest=sha256:495dc87abddb3366b05dfb398f5f94d00bcbbc8d25c9acfc114213396b296be3

Observation 31e4b70e-3842-4761-9ffd-58b30ea63fba · outbound

This paper cites Unveiling Encoder-Free Vision-Language Models.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Unveiling Encoder-Free Vision-Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:56.218152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:56.218152Z digest=sha256:5a24756657ac34d66de1f26ed6f7509623aa1356e3e9a8c08b6866d6511708d5

Observation 1c525bd1-5ec4-4b4f-9005-06ab4929f066 · outbound

This paper cites SOLO: A Single Transformer for Scalable Vision-Language Modeling.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models SOLO: A Single Transformer for Scalable Vision-Language Modeling

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-08-06T16:50:10.351624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T16:49:56.373596Z digest=sha256:6d702c781baec3f9e312b41046c6ea49e5c8c116dc3a4f78a3e3ec9549740b4e

Observation 23b62d2c-270b-4f8c-8d6d-96db9f90abcd · outbound

This paper cites Mono-internvl: Pushing the boundaries of monolithic multimodal large language models with endogenous visual pre-training,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Mono-internvl: Pushing the boundaries of monolithic multimodal large language models with endogenous visual pre-training,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:56.460558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:56.460558Z digest=sha256:10be1d4c59a3f0590731f0f2e924d3a870ec95e1b44f1457a837596301c8d4bd

Observation d4dfeccd-4672-44a9-8db0-229275532d35 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:56.572596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:56.572596Z digest=sha256:89a38e8962e95ed894dc3a7967024c3acb865d4be9f635e321b88106440f5265

Observation df089f74-f30b-44f3-a10f-2de6df635cdf · outbound

This paper cites Investigating the Catastrophic Forgetting in Multimodal Large Language Models.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Investigating the Catastrophic Forgetting in Multimodal Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:56.665670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:56.665670Z digest=sha256:a3463d24411de2541d3687ca02e861da8a6cebb0728ca2ed032c76e9161b3256

Observation 2cffd4b2-6780-4103-aa7f-85eca601f8e6 · outbound

This paper cites Delta Tuning: A Comprehensive Study of Parameter Efficient Methods for Pre-trained Language Models.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Delta Tuning: A Comprehensive Study of Parameter Efficient Methods for Pre-trained Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:56.785510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:56.785510Z digest=sha256:8468babb8bae2724f1c8ed7c27b2060119460dfe76195fba87cc9c3a019130c7

Observation 73aee5ed-f683-4926-a808-79f3ecb4aa9b · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:56.896912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:56.896912Z digest=sha256:02d7998d9016e7cc010ee6e0d5e81b5b06c5749b1628b39bc3da22c032a4f201

Observation 93868ee2-dba0-452c-8fd9-2d19ae9d9976 · outbound

This paper cites Lima: Less is more for alignment,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Lima: Less is more for alignment,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:56.999030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:56.999030Z digest=sha256:fc5dcacacc2a396651e9df3ca2b47e1de27881fcbcf12a7d3f68c326e5aabc56

Observation 27ad4242-fc85-46d9-a397-83c596813eed · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Emu3: Next-Token Prediction is All You Need

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:57.069428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:57.069428Z digest=sha256:9e0bfd96cf01f63d80368c4e45f6e3051863e5c4fa6b5b951272ef069e1917a8

Observation 901c8ced-76a7-47a9-93e0-abdc860ade0e · outbound

This paper cites The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision).

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision)

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:57.153253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:57.153253Z digest=sha256:ee171e1f61b5d568c8e1eae0ef17ca3aca1054460e632595dade97fc273c0ba9

Observation 4c338e29-89b5-40fe-980e-4969d4907720 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Gemini: A Family of Highly Capable Multimodal Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:57.248846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:57.248846Z digest=sha256:05b4c0994f79313858305aa33af0c6ca14b8b27f31ef29f968c98e4df4603fbc

Observation 220026a1-5a60-44c4-8099-462afadca072 · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Improved Baselines with Visual Instruction Tuning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:57.353167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:57.353167Z digest=sha256:2cdecd062bf2dcbe3c87561355e196c4ebc8fcc1b1594a11202f6dec5c983ef2

Observation 0f625dc8-cdfc-42c1-9448-9668a4a65cfd · outbound

This paper cites Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:57.463091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:57.463091Z digest=sha256:c8e73fd544ad8d59e41eacb3fcf250fd2c1132e072f017df0dc3de5c2da7b1c7

Observation cc75399e-c724-46e9-8338-dbd706c97401 · outbound

This paper cites Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-06T16:50:10.051759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T16:49:57.683627Z digest=sha256:d03591e82394aad41041e8573c66ebe683984cd6aaebff59f5933e680d79104d

Observation 718b4b77-42d6-4f16-baa8-b702afb4b3c8 · outbound

This paper cites SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:57.751825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:57.751825Z digest=sha256:c5beef51cbb7a641c50e521c462bd3dcda3db6d3673ac14b2925bd5e8d3bec2e

Observation f5a625a0-61af-40c6-b69e-ed10033673fa · outbound

This paper cites BLIP: bootstrapping language-image pre-training for unified vision-language understanding and generation,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models BLIP: bootstrapping language-image pre-training for unified vision-language understanding and generation,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:57.862972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:57.862972Z digest=sha256:fc589eca5315f5fbdccaa6b62dbaf42d8e17a8f4628f5e26cdda9b817f8f45bf

Observation b6ab71c8-d911-40a6-8d89-4ef5fba23175 · outbound

This paper cites Instructblip: Towards general-purpose vision-language models with instruction tuning,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Instructblip: Towards general-purpose vision-language models with instruction tuning,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:57.985266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:57.985266Z digest=sha256:1ddb73712ff58076d70df0bbb1fe0d79fd24deb3ed13ddef73c4d5f62aef1ddb

Observation dab453f1-892f-41b9-9259-186e59988c55 · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Llava-next: Improved reasoning, ocr, and world knowledge,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:58.074139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:58.074139Z digest=sha256:bffa1731157a9c5a8689526ddffdb708f48e5c8a6ddb3a50fd0eec1b75de44e5

Observation f68fdac2-7f06-4dfa-a0b8-6f5c56f104dd · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:58.168347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:58.168347Z digest=sha256:7bf537f384e678691a05c22acf914c2aeb28a24bba3ea35890cdbda6951830bb

Observation df391df8-6293-45c1-b131-2b8d35f1e3a7 · outbound

This paper cites Qwen2.5-VL Technical Report.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Qwen2.5-VL Technical Report

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:58.311277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:58.311277Z digest=sha256:0c27d5d5b4ef9975f87960583c50c8feeb0726c07adfeb1732a1af27341a1e9e

Observation 202f1580-fc12-439a-8af0-3029d2195b70 · outbound

This paper cites InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:58.390094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:58.390094Z digest=sha256:c11688375d164a8ca294279f03f6fb648b40c3db78d865825543b78495b50cb9

Observation 871361c5-bde1-4422-8f1b-ebb842a6a0b1 · outbound

This paper cites Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:58.597304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:58.597304Z digest=sha256:f2f50b584698fa08cb1c4f3900a303b9b04ffa433e5ac7f56ba9fa417bd43ae3

Observation 8eba6607-61ef-454b-a467-b2be0c2d9381 · outbound

This paper cites Mini-internvl: a flexible-transfer pocket multi-modal model with 5% parameters and 90% performance,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Mini-internvl: a flexible-transfer pocket multi-modal model with 5% parameters and 90% performance,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:58.697634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:58.697634Z digest=sha256:8059794bd02fbfb90a526cf413614547b8798ee99817987c97f097357b682da2

Observation 63a1e2a6-3313-4851-9e28-bdd64ec2e4a3 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models LLaMA: Open and Efficient Foundation Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:58.800042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:58.800042Z digest=sha256:a12123e9eded147876612b4ae6ce90727b1bfaf6b9a5bd1991c6e43ef13e0470

Observation 730f0235-ff65-47e0-b3e4-822e238e42da · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:58.881346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:58.881346Z digest=sha256:057371ef43a77bef8013b9bf6385517ab012793b059e6628c97232f6c96e4bef

Observation 0244e7a9-46d2-49e3-86a4-682e461f0b46 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models An image is worth 16x16 words: Transformers for image recognition at scale,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:58.977983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:58.977983Z digest=sha256:cbf2a747fc0718552865fc1e1c3309fb2aec4414f888452c46cfc854bd9b99a5

Observation 7e7d78cb-6aef-4d4a-8d55-1fea0f4b1312 · outbound

This paper cites Show-o: One Single Transformer to Unify Multimodal Understanding and Generation.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:59.055806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:59.055806Z digest=sha256:a77feedfe8b3093fc73fc88ccefde165cd8cafad39791ff686108d1be6b50cea

Observation bed0b7ad-391e-4c43-bca4-ac6c35dda569 · outbound

This paper cites Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:59.135515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:59.135515Z digest=sha256:b60f9d516eef53b7b91912956e11fe59ae6cfdc45e75a1cbae9dbc153528b635

Observation 2dde65db-d26c-4439-b13c-a083fc6b3b8d · outbound

This paper cites Vlmo: Unified vision-language pre-training with mixture-of-modality-experts,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Vlmo: Unified vision-language pre-training with mixture-of-modality-experts,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:59.224871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:59.224871Z digest=sha256:92f7961646ee7c989d59e0b8b7268a46cfeaa3afdfda4572285993fe38f43d16

Observation 1fe92c26-84cd-40aa-97f4-5e7386379e26 · outbound

This paper cites Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:59.336128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:59.336128Z digest=sha256:201d5cd726ddfcd7eb6a6649315f24e2caf6c5121d9f09f0dee0ae8c15a04277

Observation 3e3f5af5-d0c2-4d93-89e5-8a1fe41310a5 · outbound

This paper cites Scaling Vision-Language Models with Sparse Mixture of Experts.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Scaling Vision-Language Models with Sparse Mixture of Experts

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:59.424967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:59.424967Z digest=sha256:8975369628940e654d8b107849974990c3d29faaf6d95807e2d0294663485567

Observation f5284dae-5a57-4d0b-9d5d-eb743de78905 · outbound

This paper cites Twenty years of mixture of experts,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Twenty years of mixture of experts,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:59.548683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:59.548683Z digest=sha256:cfaa8322c5e1eebf47ddd07dbe3888f9208a056d920bbcfc89a46880853ee1e2

Observation 0a5d6262-037d-431e-83b5-cee3b08144d0 · outbound

This paper cites MoMa: Efficient Early-Fusion Pre-training with Mixture of Modality-Aware Experts.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models MoMa: Efficient Early-Fusion Pre-training with Mixture of Modality-Aware Experts

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:59.656307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:59.656307Z digest=sha256:8e831418ceeb5011d26131cd158c0a33fc46a95162a6af7e5cbb0783c2590714

Observation fc70db5e-76e7-4e1b-af06-22a874c99c41 · outbound

This paper cites Mixture-of-Depths: Dynamically allocating compute in transformer-based language models.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Mixture-of-Depths: Dynamically allocating compute in transformer-based language models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:59.760100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:59.760100Z digest=sha256:b0b3e6a12334fcd534ca3a5e9d8842d82e744d301498e51baa9b6b9d70318c93

Observation 5512c647-879c-48fd-b611-5cd6e6d1ced4 · outbound

This paper cites Aria: An Open Multimodal Native Mixture-of-Experts Model.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:59.827141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:59.827141Z digest=sha256:7966f3cfd9d00f7585897fc31c563c37757b170e7ce7c739aad6f11b13d470e5

Observation 855a26ff-2ca3-4c63-a529-07a4d654281f · outbound

This paper cites Attention is all you need,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Attention is all you need,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:59.936912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:59.936912Z digest=sha256:22586d3ce966fbbf5af372f0aa64ffb03d766e1ef7d94acf01419486cc26e6b8

Observation eb7f29c9-3b3a-430d-98aa-32245d0eae88 · outbound

This paper cites Root mean square layer normalization,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Root mean square layer normalization,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:00.080685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:00.080685Z digest=sha256:92f29f6a6a1269f9fd1e7d809846d1cfc89ff3656fc27182eac73645b6fddafa

Observation 3cec4baf-2c06-4943-b6c5-51583007169e · outbound

This paper cites Laion-5b: An open large-scale dataset for training next generation image-text models,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Laion-5b: An open large-scale dataset for training next generation image-text models,

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:00.195147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:00.195147Z digest=sha256:afb84ed698f7f2f8c90bfe2b85b88028a6010a839ea0260e21f007599615d65d

Observation 7cdd1300-3cee-46af-9e71-46005c45e6b0 · outbound

This paper cites Coyo-700m: Image-text pair dataset,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Coyo-700m: Image-text pair dataset,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:00.302668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:00.302668Z digest=sha256:0fada54e6237dfaa137c7dfbfe58425b4487544f10ee385eb9f8a7179398bbdd

Observation 55bf65fe-a0b4-424d-a3f3-aa0da6d66bbf · outbound

This paper cites Segment Anything.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Segment Anything

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:00.388194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:00.388194Z digest=sha256:a893b2f3dc88feead37d1cde7bb934b80b93e8df704814ec5082af9fe16604b9

Observation b8d07cf3-fba4-4f8d-9360-55cb2ec9d4a8 · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:00.517198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:00.517198Z digest=sha256:65e4d8b2e46b09af5e14fd00db48b4443d527eeef713a15b438ec9941823c974

Observation ffe36186-4645-44ad-8511-f3eb6f815ab9 · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:00.619189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:00.619189Z digest=sha256:65c4a0b2fe1d84329293dc8318b774d055c4c24793c1b8c04454d52909fbbd11

Observation 712be2aa-2719-4a1b-b345-2678d7220474 · outbound

This paper cites Textcaps: A dataset for image captioning with reading comprehension,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Textcaps: A dataset for image captioning with reading comprehension,

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:00.734000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:00.734000Z digest=sha256:6ad13c80b29c088036e998df7b11e467715dc37f179aa72631208cc189751470

Observation 86908a5d-546f-45a5-9e28-ddc473bd2a39 · outbound

This paper cites Objects365: A large-scale, high-quality dataset for object detection,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Objects365: A large-scale, high-quality dataset for object detection,

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:00.805492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:00.805492Z digest=sha256:9c5eab23c85a04c7ace060fa00713b3438650f37d6499e5f4186d0861630843f

Observation 776abb4f-f5de-465b-9e53-745ae1e26aa0 · outbound

This paper cites The all-seeing project: Towards panoptic visual recognition and understanding of the open world,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models The all-seeing project: Towards panoptic visual recognition and understanding of the open world,

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:00.874286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:00.874286Z digest=sha256:0b61b8187fbe18668a729ff6f80aa9b6a5794fb79f61702ba2b51a41dbbee221

Observation ba30b0f4-3e61-476c-b798-e2e44d305b27 · outbound

This paper cites Wukong: A 100 million large-scale chinese cross-modal pre-training benchmark,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Wukong: A 100 million large-scale chinese cross-modal pre-training benchmark,

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:00.947847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:00.947847Z digest=sha256:d8b2b7f89660e1327c378af59ac7243e4a37c4d43eefa8445e82fc1df5b3fd54

Observation a0b97b9e-f3a1-473d-bb14-7c57e43644a3 · outbound

This paper cites Laion coco: 600m synthetic captions from laion2b-en.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Laion coco: 600m synthetic captions from laion2b-en

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:01.031594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:01.031594Z digest=sha256:3a77f5e320f0b542a602915969ee32cca9028e7820554438beb4f47f1bedbc5e

Observation 4ea5e301-bf72-41f3-ab76-9789c6925236 · outbound

This paper cites MMC: Advancing Multimodal Chart Understanding with Large-scale Instruction Tuning.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models MMC: Advancing Multimodal Chart Understanding with Large-scale Instruction Tuning

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:01.162563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:01.162563Z digest=sha256:c7121078bdfc4f7aecfd525478709cc72ef7d7cd1db06cf71a078d5fcad666f0

Observation d2e3ed08-d8f3-43f3-a7ca-f6d5c4d79193 · outbound

This paper cites Icdar 2019 competition on large-scale street view text with partial labeling-rrc-lsvt,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Icdar 2019 competition on large-scale street view text with partial labeling-rrc-lsvt,

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:01.240559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:01.240559Z digest=sha256:948cd505a2fb8f6c55bafa3f83a925d6e87f39a0134e79ded250e794d929a725

Observation 4912a34c-18da-4fa0-97a8-d82f53569c69 · outbound

This paper cites Scene text visual question answering,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Scene text visual question answering,

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:01.315416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:01.315416Z digest=sha256:d9fe3fede2e603e138358e1983f7422d46b38166faaaa154a787efa4901ea662

Observation c7493250-004c-4ba9-b345-df73855b3bcb · outbound

This paper cites Icdar2017 competition on reading chinese text in the wild (rctw-17),.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Icdar2017 competition on reading chinese text in the wild (rctw-17),

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:01.390445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:01.390445Z digest=sha256:bff3a5531398586f456bf237d6c6d7d914473fe5af937f67093a16229115f02e

Observation 66f3aa88-0aee-4f2e-bd2b-d4dffbf4ddec · outbound

This paper cites Icdar 2019 robust reading challenge on reading chinese text on signboard,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Icdar 2019 robust reading challenge on reading chinese text on signboard,

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:01.495683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:01.495683Z digest=sha256:9f6ea8f1c036c27c57bc6e299dc422b12e7d6f11786aac8f4fde34c228497fc9

Observation b330ee7f-0072-4b50-895e-b9ee0691b6ea · outbound

This paper cites Icdar2019 robust reading challenge on arbitrary- shaped text-rrc-art,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Icdar2019 robust reading challenge on arbitrary- shaped text-rrc-art,

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:01.578838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:01.578838Z digest=sha256:6dcc3743b95cf74549ec27104470ab8329d360ff8efc646a331b0a974b36e19b

Observation 3756c26d-5fd7-4b25-929a-db601c9dbe4f · outbound

This paper cites Ocr-free document understanding transformer,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Ocr-free document understanding transformer,

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:01.712181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:01.712181Z digest=sha256:7558a202443ad0a759f72518374619e1bd5da07926f42f17820a929518b29c1c

Observation 0ec19cce-6c0a-4f41-be1e-fd52bd161f61 · outbound

This paper cites COCO-Text: Dataset and Benchmark for Text Detection and Recognition in Natural Images.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models COCO-Text: Dataset and Benchmark for Text Detection and Recognition in Natural Images

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:01.817911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:01.817911Z digest=sha256:733e683b538eab8b0c383b163e53b2a498f5f4556c893e94d7a5b1714d5bbb65

Observation 7c8d6861-01c7-4da6-9a3c-1890fc2068bc · outbound

This paper cites Chartqa: A benchmark for question answering about charts with visual and logical reasoning,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Chartqa: A benchmark for question answering about charts with visual and logical reasoning,

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:01.924520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:01.924520Z digest=sha256:70b539e4ed79c6deb8511b1149463dd01cea35e3474e1777f9d0f2d5765325e6

Observation 61e6366c-974c-407f-8fcb-4b8263ade4fe · outbound

This paper cites A large chinese text dataset in the wild,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models A large chinese text dataset in the wild,

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:02.002148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:02.002148Z digest=sha256:e577023889e7a7bd17b8f5fb7674ef64e3b9d604a235e9b1fdbbf4a2d85c043c

Observation d043d554-ecdf-429c-88da-cf42cc6088d7 · outbound

This paper cites Simple and effective multi-paragraph reading comprehension,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Simple and effective multi-paragraph reading comprehension,

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:02.099560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:02.099560Z digest=sha256:dde2440510a4c2ff7b5084de4590b266ac086c2829be370acb06ef9d04043a92

Observation 0b0c7140-c968-45e5-a601-f6060c883f63 · outbound

This paper cites Textocr: Towards large-scale end-to-end reasoning for arbitrary-shaped scene text,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Textocr: Towards large-scale end-to-end reasoning for arbitrary-shaped scene text,

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:02.203648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:02.203648Z digest=sha256:6a6091c97ab4b992ded5660b2cceecea333e7aaa305c3c5aa6fc0d9a521682d2

Observation 07cde5a0-af22-42d9-a8cd-9c43ab5f7c76 · outbound

This paper cites Plotqa: Reasoning over scientific plots,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Plotqa: Reasoning over scientific plots,

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:02.320729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:02.320729Z digest=sha256:a0f9f4568be94993b64db0b2b5b7624cd8223b221bac5c997314b28f3ddb15cf

Observation c83be487-e7ae-427b-bf7a-9f520ec9f738 · outbound

This paper cites Infographicvqa,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Infographicvqa,

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:02.423626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:02.423626Z digest=sha256:7b3bcccf658fc66e6481c8aeabf359ab9fe1a288f9d8344b6597ef03687e5e50

Observation 66675f87-dc47-41fa-b6ab-4ab661215fed · outbound

This paper cites Making the V in VQA matter: Elevating the role of image understanding in visual question answering,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Making the V in VQA matter: Elevating the role of image understanding in visual question answering,

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:02.536884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:02.536884Z digest=sha256:843eb0c6b1e95d3f294cc18e7f20e3c1a340d6a800aba91c5fc49da34c950d8c

Observation 546425af-4490-472f-8619-4c156080f53f · outbound

This paper cites GQA: A new dataset for real-world visual reasoning and compositional question answering,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models GQA: A new dataset for real-world visual reasoning and compositional question answering,

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:02.641238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:02.641238Z digest=sha256:6555a29f0e7124054f11e1062f8c2bc30cc6cbba5132350d56b1780118dea3af

Observation f900c643-6726-47c9-aa9f-549eb9f80e8f · outbound

This paper cites Ok-vqa: A visual question answering benchmark requiring external knowledge,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Ok-vqa: A visual question answering benchmark requiring external knowledge,

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:02.746998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:02.746998Z digest=sha256:f231e8aa7964841cbb33e760a85ffa235942538ce82619e4743c0b56fb407011

Observation 2bc39bd7-4222-4de2-9972-5cd6c418f854 · outbound

This paper cites Visual spatial reasoning,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Visual spatial reasoning,

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:02.852857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:02.852857Z digest=sha256:7eed3240f6379a23a0dd65a5314a3e09e511db182ae70053083a57d2f06f5ced

Observation d9264504-9076-436e-8b1e-bc1716e767a3 · outbound

This paper cites Visual dialog,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Visual dialog,

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:02.963424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:02.963424Z digest=sha256:c1589f8f6de9c45a595348be4e6253aabee64984450c4eb02785d2f043205d7d

Observation eeba70c5-126f-4a87-bc5f-c70595bee6fd · outbound

This paper cites A diagram is worth a dozen images,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models A diagram is worth a dozen images,

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:03.027579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:03.027579Z digest=sha256:8ed48d17aeffea2a76669ac21e486cdfb967dee0fdceaa112242386fa7479e16

Observation c19cef38-137e-4204-beaa-d73fa8976722 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Learn to explain: Multimodal reasoning via thought chains for science question answering,

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:03.153797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:03.153797Z digest=sha256:7e806c2536f3c0f3b96a57b3dea7385926de68801ddefe38ec6344e11d3eca5f

Observation 4b270660-b3fa-4df9-b4fd-4cbb7a7e6ffd · outbound

This paper cites Are you smarter than a sixth grader? textbook question answering for multimodal machine comprehension,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Are you smarter than a sixth grader? textbook question answering for multimodal machine comprehension,

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:03.236776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:03.236776Z digest=sha256:5044c23756d6f2aeabf97770e71609fdb0d47f22cdc86591d152d8246093f7fb

Observation 28539a3e-357e-43f1-a2e4-1a229e7f10a5 · outbound

This paper cites Dvqa: Understanding data visualizations via question answering,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Dvqa: Understanding data visualizations via question answering,

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:03.384359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:03.384359Z digest=sha256:ae05adde0ec96253409e03bd691d59bf026e85b1525afc863fbf347094ae43cd

Observation 3f8cb091-aeeb-40d2-97cc-477d8c8b0a9f · outbound

This paper cites Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:03.529613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:03.529613Z digest=sha256:02a4574de5dac8a6bdf4f2372bf2f67ada217e5c264cb15c7cba6b51c239abee

Observation 7382f2c5-4d5a-431e-bb17-b95211031b07 · outbound

This paper cites An augmented benchmark dataset for geometric question answering through dual parallel text encoding,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models An augmented benchmark dataset for geometric question answering through dual parallel text encoding,

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:03.654147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:03.654147Z digest=sha256:b0c6c15c51af285cac33a15870cdfe7294f89577faf04a38b86aee78eaafe00e

Observation f8b413f7-7d5a-45e3-8655-2a7e0c828785 · outbound

This paper cites Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:03.783517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:03.783517Z digest=sha256:1f675a17dc30d6500d540667af1d15482aa7d28003fac6807bf271098680bf86

Observation 33c867b7-c587-4172-a55b-d3a85925e403 · outbound

This paper cites MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:03.926291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:03.926291Z digest=sha256:f9c1e9d4bf1012dd5a502163068d15829ae669c9a388c2acc6425c2357592ae7

Observation 65f50570-1ed8-4434-a038-0a613d7da31b · outbound

This paper cites CLEVR-Math: A Dataset for Compositional Language, Visual and Mathematical Reasoning.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models CLEVR-Math: A Dataset for Compositional Language, Visual and Mathematical Reasoning

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:04.072385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:04.072385Z digest=sha256:ed6d8ddd8763fd33a82788401854f7be4ef30407c6a41d1ceae3dde5b0f73dea

Observation 26aebb84-91f1-4696-a69d-57638124a7e8 · outbound

This paper cites Super-clevr: A virtual benchmark to diagnose domain robustness in visual reasoning,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Super-clevr: A virtual benchmark to diagnose domain robustness in visual reasoning,

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:04.211430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:04.211430Z digest=sha256:e27cd481e76cdb91a0d88bc1ad5a7d8cf5d5489662041abf93b5b93d7608bd63

Observation 3eaa7f89-402a-4e8d-bff3-c1d1320a9647 · outbound

This paper cites Inter-GPS: Interpretable Geometry Problem Solving with Formal Language and Symbolic Reasoning.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Inter-GPS: Interpretable Geometry Problem Solving with Formal Language and Symbolic Reasoning

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:04.330336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:04.330336Z digest=sha256:e3a822e728bb06eae65460f274fe5bb0a16c5f11564e1c81a73078f33ccda7ae

Observation 9408ddf6-13f9-46d8-9640-f0bca390a041 · outbound

This paper cites Kvqa: Knowledge- aware visual question answering,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Kvqa: Knowledge- aware visual question answering,

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:04.443260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:04.443260Z digest=sha256:f31da448a9bcd421ed76d0f0ad79c622b677017ad1a498fbf7ee7cc6e39fb832

Observation e2063723-1016-463b-a2ca-3d5ac2691921 · outbound

This paper cites A-okvqa: A benchmark for visual question answering using world knowledge,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models A-okvqa: A benchmark for visual question answering using world knowledge,

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:04.560616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:04.560616Z digest=sha256:6ea052804a0c52bf41372653896ab5a5fb274dbb20263008c959f518459eb82c

Observation 197b7174-653d-4cb5-9496-32b7ee46a3a6 · outbound

This paper cites Viquae, a dataset for knowledge- based visual question answering about named entities,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Viquae, a dataset for knowledge- based visual question answering about named entities,

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:04.660448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:04.660448Z digest=sha256:78e7fdd0c7e594ce26104c5d6011092ef8e899522de1c4e6da03c6b115afccbc

Observation 3d17bb68-d896-410f-921f-b62d7d2c94aa · outbound

This paper cites WanJuan: A Comprehensive Multimodal Dataset for Advancing English and Chinese Large Models.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models WanJuan: A Comprehensive Multimodal Dataset for Advancing English and Chinese Large Models

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:04.722409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:04.722409Z digest=sha256:344a7949aef8f4589db6dafd2c618ec09b6ef3b598a5a81fb9b0686bdda4f1b9

Observation 7cb3cda1-178f-4ead-a61f-9f8600364ab3 · outbound

This paper cites Ocr-vqa: Visual question answering by reading text in images,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Ocr-vqa: Visual question answering by reading text in images,

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:04.785021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:04.785021Z digest=sha256:25f4a358a33ec68dee6876481f6e30ca24fa8fa8b8a475978112d6cad3a10ab6

Observation b50417c2-8ed8-42df-b13a-983ad4fd2128 · outbound

This paper cites Towards VQA models that can read,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Towards VQA models that can read,

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:04.861787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:04.861787Z digest=sha256:46bffebbaaecc20b012be35608e1bf94caba9943237b7f97ca1e8de31190a7db

Observation 581d54b1-e307-452a-bfe9-6335283eaa70 · outbound

This paper cites Modeling context in referring expressions,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Modeling context in referring expressions,

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:04.943013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:04.943013Z digest=sha256:c786e5b9bf192f45984aeb15baf28c4946a5b4808872395d8f12cdbdf69eebe7

Observation 8f850fd4-aafe-4cb7-9acd-c1616efcf848 · outbound

This paper cites Generation and comprehension of unambiguous object descriptions,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Generation and comprehension of unambiguous object descriptions,

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:05.027353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:05.027353Z digest=sha256:eeee6bb4a0bc7cd89898544bd05941152cf24a1239891ad3c638b8e57da36926

Observation ed4402f5-07e4-43c4-bc10-7c2124596bd3 · outbound

This paper cites Visual genome: Connecting language and vision using crowdsourced dense image annotations,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Visual genome: Connecting language and vision using crowdsourced dense image annotations,

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:05.105073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:05.105073Z digest=sha256:dcd977453ca5807bbd9a78bed5c2147d5cf795763f917c34f0857775e957d4ff

Observation b970ecbc-9311-4a1b-bd3c-f83feadd74b4 · outbound

This paper cites To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:05.174737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:05.174737Z digest=sha256:e7287fe5f789a02775abf56f9c5260eda6128cf24462c1f987a7b1a2c40556b7

Observation 244696a9-e355-4869-9b6a-18fa1cf73075 · outbound

This paper cites ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:05.268162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:05.268162Z digest=sha256:d89dd1cc11cc18945c6f18fb2bef0c862c90cbf7576c8b2ccb0a3a934863c535

Observation 65664252-de60-4c8d-b58a-736bf1bdcfb6 · outbound

This paper cites Gpt-4v dataset,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Gpt-4v dataset,

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:05.335533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:05.335533Z digest=sha256:3f1cc3ef525238cb91ec59a1c2b6af5742a11f2ad8bb0c6c20405c10d052bbbc

Observation 6d38fdb5-0f7b-4210-b379-69b149ae4558 · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Judging llm-as-a-judge with mt-bench and chatbot arena,

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:05.406294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:05.406294Z digest=sha256:a90b06bd13814f6c0701d454b1a5fe8387d3ea9151af9a25ababe6fe8cdfe555

Observation d86c7702-3db9-44f1-ab53-65377a19e076 · outbound

This paper cites SVIT: Scaling up Visual Instruction Tuning.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models SVIT: Scaling up Visual Instruction Tuning

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:05.494156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:05.494156Z digest=sha256:e0d918aaa8ca249ae756bdfc1867ea72f3691ed605a83f3453417c3d3e2d4aa2

Observation 24e6065c-6e91-4ccd-ade8-cada6d4633bf · outbound

This paper cites Openhermes 2.5: An open dataset of synthetic data for generalist llm assistants,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Openhermes 2.5: An open dataset of synthetic data for generalist llm assistants,

Reference 101

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:05.568237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:05.568237Z digest=sha256:a255691e5fc440ce24af4f164e21a893bf65288fa32281cac540419a665e263b

Pith citing papers

Observation fe67a2ae-8c0f-4042-abb3-5246f5c6fee4 · inbound

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency cites this paper.

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:58:59.202761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T11:58:58.660564Z digest=sha256:ed7ba9c6e24b7a1e4019baed1bdc73e17814290d32e97c92ded1b1b5cb919dbc

Observation 85ce1315-55db-4859-9f18-b28e63813223 · inbound

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture cites this paper.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models

Reference 88

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:17:18.752838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:9e09b966e52a5cb257142bc5baf12278ab6a64d31cb587e04cbb7f2d8e3e14a5