Pith. sign in

Paper Citation Record · LEDGER

Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance

As of 19 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 19 inbound Pith citation observations for arXiv:2410.16261.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.16261 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 19 of 19 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:22:52.465714Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T10:39:44.698268Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c5ece020-f293-4eae-85a2-43dcce0cec50 · inbound

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective cites this paper.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T15:39:32.457788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:39:32.457788Z digest=sha256:64c16918105671273c6c4c394376704e7b386bb8f58861f1e5671c18c6870c38

Observation 650d6f95-f6fe-48d0-9e13-5b35b1e0c92a · inbound

Looking Beyond Text: Reducing Language bias in Large Vision-Language Models via Multimodal Dual-Attention and Soft-Image Guidance cites this paper.

Looking Beyond Text: Reducing Language bias in Large Vision-Language Models via Multimodal Dual-Attention and Soft-Image Guidance Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T15:26:16.360622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:26:16.360622Z digest=sha256:14eec8fbf0620dd70ac4fdda6b9b741ce122932139b1fa0f483809b9dd501c38

Observation 7fedb350-47bb-4f9a-9359-3913c8e6be31 · inbound

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling cites this paper.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:23:58.094306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:92c55d9624d6f348dac00ede758de84336b08852854f6d4d34dea7c9f424f38a

Observation da3e5a54-df68-4abb-9546-69fb36b0dee1 · inbound

V2PE: Improving Multimodal Long-Context Capability of Vision-Language Models with Variable Visual Position Encoding cites this paper.

V2PE: Improving Multimodal Long-Context Capability of Vision-Language Models with Variable Visual Position Encoding Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T16:58:03.045413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:58:03.045413Z digest=sha256:7b8e40eae2b0af161aa1c6aa714681cbd9f8ec1fba953a813bb1c1f557878033

Observation b29c2515-6685-4636-81f5-332033a2f65a · inbound

Do Language Models Understand Time? cites this paper.

Do Language Models Understand Time? Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T12:47:17.236616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:47:17.236616Z digest=sha256:41b57aa64255a17dcb3d68ac40881a1734a71094138fade6a644a0fd2688cea1

Observation 74a09442-949e-43fa-b0e9-8dcd2b66185e · inbound

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models cites this paper.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:41:08.305987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:1524966536171d456beba9ec254a97450e6663cba8d543c3676e87d331c16f3b

Observation b33ecc19-b636-4868-807b-f9f156c0d2a3 · inbound

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? cites this paper.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T10:22:52.465714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:22:52.465714Z digest=sha256:98672e8f1752a4b965c3c64d39b4cff9514c3e2c5bdf45c296619c39ef03a1bb

Observation 25d78dc3-1e8c-4e0e-b101-df0023bd965b · inbound

MMAFFBen: A Multilingual and Multimodal Affective Analysis Benchmark for Evaluating LLMs and VLMs cites this paper.

MMAFFBen: A Multilingual and Multimodal Affective Analysis Benchmark for Evaluating LLMs and VLMs Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:27:32.357748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:27:32.357748Z digest=sha256:15c91b2ba25bd2c9103052c5ab601dfc0c5767e655590aaa8ee1ba962cd8d0bc

Observation fd55f037-e5a8-499c-80cb-7ce337155cc7 · inbound

Challenging Vision-Language Models with Surgical Data: A New Dataset and Broad Benchmarking Study cites this paper.

Challenging Vision-Language Models with Surgical Data: A New Dataset and Broad Benchmarking Study Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T06:01:41.004736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:01:41.004736Z digest=sha256:758621a7cfc90a1a41421bb77358ce535ecfb753e3f4486aa4719a0e4117ff3c

Observation 05f04817-a770-4ae9-a563-be7b0d121e73 · inbound

Pushing the Limits of Safety: A Technical Report on the ATLAS Challenge 2025 cites this paper.

Pushing the Limits of Safety: A Technical Report on the ATLAS Challenge 2025 Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:13.702701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:13.702701Z digest=sha256:dc03a5ccab3808c70726e233bcb95ddcf9e4877a68b9aace2c1d2a0995b22267

Observation 7883666e-4ec6-4b91-a616-a090dfa88957 · inbound

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets cites this paper.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T19:45:52.913989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:45:52.913989Z digest=sha256:9525388fdeac1ed83c61bb415a5ab6a4d50efae3a365431699ba72afd1d468dc

Observation b00b1574-7563-444b-bf4e-6c77be7efac6 · inbound

Docopilot: Improving Multimodal Models for Document-Level Understanding cites this paper.

Docopilot: Improving Multimodal Models for Document-Level Understanding Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T15:57:00.817293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:57:00.817293Z digest=sha256:292196f5710eb4a3690dba8798c396c1136301dd001ed48817c591ad8e825484

Observation 4cd8fa66-5bfa-4f4f-a128-d1cd0d32135b · inbound

In-context Learning of Vision Language Models for Detection of Physical and Digital Attacks against Face Recognition Systems cites this paper.

In-context Learning of Vision Language Models for Detection of Physical and Digital Attacks against Face Recognition Systems Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-06T15:43:16.211195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:43:16.211195Z digest=sha256:b85a71d02c4c025ef9a0171c9cec7c9622947b8fbef167f20e8ac54d271ebd09

Observation 2cb98090-3eb1-4a65-8b1f-1d91e71b65e4 · inbound

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency cites this paper.

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:58:59.097039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T11:58:58.660564Z digest=sha256:16fb2927b5cc0f8101799277b36977971acd9d8b497908cdbca8d0c3ec4373f8

Observation 00a4bd24-b0d4-4288-8287-cf1da5b5f214 · inbound

The Blind Spot of Adaptation: Quantifying and Mitigating Forgetting in Fine-tuned Driving Models cites this paper.

The Blind Spot of Adaptation: Quantifying and Mitigating Forgetting in Fine-tuned Driving Models Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:15:50.128171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T20:04:46.144856Z digest=sha256:f666eef3d4c3bbf41a5245869bb353fc7ed28d442e8162576f7448194290440d

Observation 2764177a-b161-4f0c-a723-9ac3cb1d0fd5 · inbound

A Readiness-Driven Runtime for Pipeline-Parallel Training under Runtime Variability cites this paper.

A Readiness-Driven Runtime for Pipeline-Parallel Training under Runtime Variability Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:38:09.440345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T07:35:32.225708Z digest=sha256:d946c52e68eed02822282e1aa9db50bbeea0bd302eeba3c2aa9f2d2f7b3e16af

Observation f8b9de0f-89ce-459f-a406-dd62ab82d8f8 · inbound

EventDrive: Event Cameras for Vision-Language Driving Intelligence cites this paper.

EventDrive: Event Cameras for Vision-Language Driving Intelligence Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:18:57.083781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T01:26:43.752335Z digest=sha256:18eca2ba964824be5ed765b498ee0548677603745a9652fc4e0b2b93ec844684

Observation 6ef7b290-08a6-4782-9174-8a312e6a2b37 · inbound

LightSTAR: Efficient Visual Document Retrieval via Lightweight Selection with Vision-Adaptive Refinement cites this paper.

LightSTAR: Efficient Visual Document Retrieval via Lightweight Selection with Vision-Adaptive Refinement Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T10:39:44.700307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-26T08:43:17.902876Z digest=sha256:991b255cd3871e1daa81fbc5d17b25c4bde1846f64e59383944ffc89288e3e3a

Observation 2d5289ad-e642-4ab7-8f38-abb0500efc6f · inbound

Do Video-LLMs Actually Watch? Diagnosing Character-Tracking Failures in Long-Form Video cites this paper.

Do Video-LLMs Actually Watch? Diagnosing Character-Tracking Failures in Long-Form Video Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-14T07:12:52.433947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T07:12:52.433947Z digest=sha256:59e8432da9967abdf783e8eb37a908287df5e5f8453a4d4a1bc1352efacdb8cd