Pith. sign in

Paper Citation Record · LEDGER

Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 14 inbound Pith citation observations for arXiv:2410.16261.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.16261 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:27:32.357748Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T10:39:44.698268Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 7fedb350-47bb-4f9a-9359-3913c8e6be31 · inbound

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling cites this paper.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:23:58.094306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:d90aedc8e36bf58fbf0da4007b57c8c7bde6fc8cbeabcfcb8fd65e1043112be7

Observation 74a09442-949e-43fa-b0e9-8dcd2b66185e · inbound

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models cites this paper.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:41:08.305987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:63d07ae7424f09bbde03b38a5cc6870f700f3241623806e29d51dd1e3bbeaf2f

Observation 25d78dc3-1e8c-4e0e-b101-df0023bd965b · inbound

MMAFFBen: A Multilingual and Multimodal Affective Analysis Benchmark for Evaluating LLMs and VLMs cites this paper.

MMAFFBen: A Multilingual and Multimodal Affective Analysis Benchmark for Evaluating LLMs and VLMs Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:27:32.357748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:27:32.357748Z digest=sha256:3c5291d849def6c080afd94004c616cec01df583e738cbf7278cd80192d5f05e

Observation fd55f037-e5a8-499c-80cb-7ce337155cc7 · inbound

Challenging Vision-Language Models with Surgical Data: A New Dataset and Broad Benchmarking Study cites this paper.

Challenging Vision-Language Models with Surgical Data: A New Dataset and Broad Benchmarking Study Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T06:01:41.004736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:01:41.004736Z digest=sha256:d0c8e2aa5ab1b9f5b9ffb8b11b414d004703eebbfd13c6002afd8beee64a7393

Observation 05f04817-a770-4ae9-a563-be7b0d121e73 · inbound

Pushing the Limits of Safety: A Technical Report on the ATLAS Challenge 2025 cites this paper.

Pushing the Limits of Safety: A Technical Report on the ATLAS Challenge 2025 Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:13.702701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:13.702701Z digest=sha256:7a14b4bbe328bf9fcc64a1f0339c8b8e04fbfc055ba885e257c107c3d8ba5f98

Observation 7883666e-4ec6-4b91-a616-a090dfa88957 · inbound

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets cites this paper.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T19:45:52.913989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:45:52.913989Z digest=sha256:869977f56479fe606b04be2ab433bee1080da6f021fe022e2abb95ea1ca629f6

Observation b00b1574-7563-444b-bf4e-6c77be7efac6 · inbound

Docopilot: Improving Multimodal Models for Document-Level Understanding cites this paper.

Docopilot: Improving Multimodal Models for Document-Level Understanding Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T15:57:00.817293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:57:00.817293Z digest=sha256:cf4d1d06515852a406ca04ee5b18ca8ae5193c02999cba5be53949c78950e199

Observation 4cd8fa66-5bfa-4f4f-a128-d1cd0d32135b · inbound

In-context Learning of Vision Language Models for Detection of Physical and Digital Attacks against Face Recognition Systems cites this paper.

In-context Learning of Vision Language Models for Detection of Physical and Digital Attacks against Face Recognition Systems Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-06T15:43:16.211195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:43:16.211195Z digest=sha256:6e30f5ef09e0f15578800e55d2b97671e3752a549399ce8733d4b0de5faa5528

Observation 2cb98090-3eb1-4a65-8b1f-1d91e71b65e4 · inbound

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency cites this paper.

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:58:59.097039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T11:58:58.660564Z digest=sha256:e01f2988d94c9ffbe1bfe234186ee9643ca3a2db8895af23de936ffb8f263667

Observation 00a4bd24-b0d4-4288-8287-cf1da5b5f214 · inbound

The Blind Spot of Adaptation: Quantifying and Mitigating Forgetting in Fine-tuned Driving Models cites this paper.

The Blind Spot of Adaptation: Quantifying and Mitigating Forgetting in Fine-tuned Driving Models Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:15:50.128171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T20:04:46.144856Z digest=sha256:4e582314cd03019685cfd03cf08cde34d4bd08c98374da5b3fd31fe2a546e6de

Observation 2764177a-b161-4f0c-a723-9ac3cb1d0fd5 · inbound

A Readiness-Driven Runtime for Pipeline-Parallel Training under Runtime Variability cites this paper.

A Readiness-Driven Runtime for Pipeline-Parallel Training under Runtime Variability Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:38:09.440345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T07:35:32.225708Z digest=sha256:ddcadd63f100436000eeb166802299172d08fd7ab9cd1adfba51194525f72fa5

Observation f8b9de0f-89ce-459f-a406-dd62ab82d8f8 · inbound

EventDrive: Event Cameras for Vision-Language Driving Intelligence cites this paper.

EventDrive: Event Cameras for Vision-Language Driving Intelligence Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:18:57.083781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T01:26:43.752335Z digest=sha256:8aecdb2e920ca35b2c37b56235149be701defd40dd6d3e8688f234245eda491c

Observation 6ef7b290-08a6-4782-9174-8a312e6a2b37 · inbound

LightSTAR: Efficient Visual Document Retrieval via Lightweight Selection with Vision-Adaptive Refinement cites this paper.

LightSTAR: Efficient Visual Document Retrieval via Lightweight Selection with Vision-Adaptive Refinement Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T10:39:44.700307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T08:43:17.902876Z digest=sha256:16b03f9d74236cf3f69b8c30252d25a50a95a87c084fb3e66a81ae8ca77c92dc

Observation 2d5289ad-e642-4ab7-8f38-abb0500efc6f · inbound

Do Video-LLMs Actually Watch? Diagnosing Character-Tracking Failures in Long-Form Video cites this paper.

Do Video-LLMs Actually Watch? Diagnosing Character-Tracking Failures in Long-Form Video Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-14T07:12:52.433947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T07:12:52.433947Z digest=sha256:068c6c40d1ef3d196c40f9e16252cd2840c0476ccdb4173ffb37ee62f9db4fee