Pith. sign in

Paper Citation Record · LEDGER

Mono-InternVL: Pushing the Boundaries of Monolithic Multimodal Large Language Models with Endogenous Visual Pre-training

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2410.08202.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.08202 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:21:05.372302Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-16T09:16:17.330970Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 9f39929b-a390-4545-91ca-7d44a6c219a5 · inbound

Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization cites this paper.

Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization Mono-InternVL: Pushing the Boundaries of Monolithic Multimodal Large Language Models with Endogenous Visual Pre-training

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:16:17.333274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T09:16:17.150383Z digest=sha256:a212db59d77027f3ee026e700851cda752b9cdd87b2b14d3dbd0e53ab8428a20

Observation ea889e35-28b7-4e61-9495-e975b419072e · inbound

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling cites this paper.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Mono-InternVL: Pushing the Boundaries of Monolithic Multimodal Large Language Models with Endogenous Visual Pre-training

Reference 172

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:23:57.806018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:fb9430a569063db2e0c5f410cbcc2844c482b5f0dc9431c3266b4d8ecd57f539

Observation 13d25a94-f9e4-46f9-ac01-cb6142da07aa · inbound

TimeCausality: Evaluating the Causal Ability in Time Dimension for Vision Language Models cites this paper.

TimeCausality: Evaluating the Causal Ability in Time Dimension for Vision Language Models Mono-InternVL: Pushing the Boundaries of Monolithic Multimodal Large Language Models with Endogenous Visual Pre-training

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T15:21:05.372302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:21:05.372302Z digest=sha256:9e22d884e0362645e6d68cfb76dc8ae007fe7816a09ee1e1dc5861986bbdc627

Observation 8ef13d95-c7ca-4dad-b000-c89cc7399d84 · inbound

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces cites this paper.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Mono-InternVL: Pushing the Boundaries of Monolithic Multimodal Large Language Models with Endogenous Visual Pre-training

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:13.209440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:13.209440Z digest=sha256:076d7c4d1b5c7e6395b504eedf7125c72de523bef9732fada844bc049fdae968

Observation 5c10ab9b-7207-44a0-b710-b1ec91a7e84f · inbound

GOBench: Benchmarking Geometric Optics Generation and Understanding of MLLMs cites this paper.

GOBench: Benchmarking Geometric Optics Generation and Understanding of MLLMs Mono-InternVL: Pushing the Boundaries of Monolithic Multimodal Large Language Models with Endogenous Visual Pre-training

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:23.699416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:23.699416Z digest=sha256:394ea9b51a6132756de0822a3571feb10886ffdef9929ec6fcabbd907eb355b5

Observation 3e05ec7b-f3b2-4b73-9f76-0d1c6efa92ff · inbound

SMAR: Soft Modality-Aware Routing Strategy for MoE-based Multimodal Large Language Models Preserving Language Capabilities cites this paper.

SMAR: Soft Modality-Aware Routing Strategy for MoE-based Multimodal Large Language Models Preserving Language Capabilities Mono-InternVL: Pushing the Boundaries of Monolithic Multimodal Large Language Models with Endogenous Visual Pre-training

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:18.718743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:18.718743Z digest=sha256:c334e6155a8b6ec3974eaf86f0908bef1f4221a056363b620038dda0b1ba329c

Observation d6c1c945-7804-4a3a-abcf-16ce18d90eec · inbound

Breaking Bad Molecules: Are MLLMs Ready for Structure-Level Molecular Detoxification? cites this paper.

Breaking Bad Molecules: Are MLLMs Ready for Structure-Level Molecular Detoxification? Mono-InternVL: Pushing the Boundaries of Monolithic Multimodal Large Language Models with Endogenous Visual Pre-training

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-07T04:17:59.681471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:17:59.681471Z digest=sha256:8861cb12a2663fca311a6cc002b6bfe67dec465168eac4e947835c561f558121

Observation a59232dd-a6e0-4986-affb-a787d45ad1d7 · inbound

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation cites this paper.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Mono-InternVL: Pushing the Boundaries of Monolithic Multimodal Large Language Models with Endogenous Visual Pre-training

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.784958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.784958Z digest=sha256:06daac64a8c740d9c2c1ebb852f4fb0f3136cd21b0af8994b2150557c2ab1799

Observation f0432910-0c1c-432f-b169-d2edd481baa3 · inbound

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs cites this paper.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Mono-InternVL: Pushing the Boundaries of Monolithic Multimodal Large Language Models with Endogenous Visual Pre-training

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:08.054862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:08.054862Z digest=sha256:6e97a6b4a1decbb09b48338390186c39f4d9a12817df34df72898882bd53cca0

Observation cf96d5d1-f44a-4e86-b065-e6a114834a41 · inbound

MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning cites this paper.

MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning Mono-InternVL: Pushing the Boundaries of Monolithic Multimodal Large Language Models with Endogenous Visual Pre-training

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T05:36:19.867128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:36:19.867128Z digest=sha256:9972b550b895862841b4998cf4c8d45fd549821454be90c76602f96373e8f891

Observation dd768d52-3dee-44e1-a7bf-da37ea51c617 · inbound

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency cites this paper.

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency Mono-InternVL: Pushing the Boundaries of Monolithic Multimodal Large Language Models with Endogenous Visual Pre-training

Reference 80

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:58:59.207528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T11:58:58.660564Z digest=sha256:7cd363b59cfe57aeeffa99d03663b1cea51628cdd259144776fd73dd5b6e3742