Pith. sign in

Paper Citation Record · LEDGER

Vision-Flan: Scaling Human-Labeled Tasks in Visual Instruction Tuning

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 14 inbound Pith citation observations for arXiv:2402.11690.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.11690 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:46:32.685395Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T18:33:50.167861Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f13163c8-e6d7-4c9f-a577-0302221d0768 · inbound

SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation cites this paper.

SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation Vision-Flan: Scaling Human-Labeled Tasks in Visual Instruction Tuning

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:48:36.226183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T22:48:36.010306Z digest=sha256:9d813e149216b5ab9c2026ad3baf0e33b10227c3583469f1bf5c8598da081f5e

Observation 172ad1be-4fdb-4bca-990f-31e145282e45 · inbound

LLaVA-OneVision: Easy Visual Task Transfer cites this paper.

LLaVA-OneVision: Easy Visual Task Transfer Vision-Flan: Scaling Human-Labeled Tasks in Visual Instruction Tuning

Reference 145

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:23:49.510641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:02b88db772256593b3af6add38952eb81ea6612c50a30acdcc46741462cd9662

Observation cf6575ec-7d5a-49b1-995c-67ee23ccd43d · inbound

VILA-M3: Enhancing Vision-Language Models with Medical Expert Knowledge cites this paper.

VILA-M3: Enhancing Vision-Language Models with Medical Expert Knowledge Vision-Flan: Scaling Human-Labeled Tasks in Visual Instruction Tuning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T17:08:41.752207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:08:41.752207Z digest=sha256:1cf893f6b4251752bae39e699f428469e65338917c4bf158d5752bfddb24982d

Observation 38c05955-15bf-4a53-862c-588bbcb3612d · inbound

Exploring the Adversarial Vulnerabilities of Vision-Language-Action Models in Robotics cites this paper.

Exploring the Adversarial Vulnerabilities of Vision-Language-Action Models in Robotics Vision-Flan: Scaling Human-Labeled Tasks in Visual Instruction Tuning

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-12T18:51:16.580678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:51:16.580678Z digest=sha256:e528e15b98867ac90fcedb78085f0bed35983cebdfb885919f7396a26d56371c

Observation 79c769cf-075d-44d9-81e3-990c5440e972 · inbound

SMoLoRA: Exploring and Defying Dual Catastrophic Forgetting in Continual Visual Instruction Tuning cites this paper.

SMoLoRA: Exploring and Defying Dual Catastrophic Forgetting in Continual Visual Instruction Tuning Vision-Flan: Scaling Human-Labeled Tasks in Visual Instruction Tuning

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T15:46:29.656799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:46:29.656799Z digest=sha256:07a0ee68026c89d7d7461a7564d9edc461d7c382015384cb5715c09eddaea5df

Observation 035f442c-77a0-4f00-8af0-318e6cd87f1a · inbound

On Domain-Adaptive Post-Training for Multimodal Large Language Models cites this paper.

On Domain-Adaptive Post-Training for Multimodal Large Language Models Vision-Flan: Scaling Human-Labeled Tasks in Visual Instruction Tuning

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-12T05:56:46.865140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:56:46.865140Z digest=sha256:7dbd43fc85a7aaadd39dbe6c493ac33d2826df55b17b63e8c70a5a974b469944

Observation 9e5be99e-4e9a-4fdf-8ec0-67aff74c1156 · inbound

RADIOv2.5: Improved Baselines for Agglomerative Vision Foundation Models cites this paper.

RADIOv2.5: Improved Baselines for Agglomerative Vision Foundation Models Vision-Flan: Scaling Human-Labeled Tasks in Visual Instruction Tuning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T18:40:25.573763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:40:25.573763Z digest=sha256:053157eed389815bd5967a751e79f703dd54d6e0c0f3136fbf587c53fdf909bf

Observation 91456423-fd33-4212-b3bb-c03357759183 · inbound

Error-driven Data-efficient Large Multimodal Model Tuning cites this paper.

Error-driven Data-efficient Large Multimodal Model Tuning Vision-Flan: Scaling Human-Labeled Tasks in Visual Instruction Tuning

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:35.752639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:35.752639Z digest=sha256:cae00ff266d574fb907c86efb4dbec73e9751bea4b68da3c619ec9d48ec26dd8

Observation e5e26d06-e49c-4bed-9787-f7b0f29af39f · inbound

Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts cites this paper.

Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts Vision-Flan: Scaling Human-Labeled Tasks in Visual Instruction Tuning

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:12.638612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:42:12.638612Z digest=sha256:2d0cf60f57fce0bc864bd9ed71458b1af0a08ec21ffda7630ea5ba4a57e051ed

Observation 97af1f60-fa2e-47f0-9ea1-036141f5622b · inbound

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design cites this paper.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Vision-Flan: Scaling Human-Labeled Tasks in Visual Instruction Tuning

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:19.215334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:19.215334Z digest=sha256:191794c6002faae56380dc8ac17072fc8e88f5b546fc160a78ca5a976dfee08d

Observation 22adde2f-2960-441e-a9bb-fbb12cdc4497 · inbound

TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types cites this paper.

TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types Vision-Flan: Scaling Human-Labeled Tasks in Visual Instruction Tuning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T20:06:36.670967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:06:36.670967Z digest=sha256:da374483255755800b1bf0cc321f263c19bdfe90bad591bbc60449a624a453cc

Observation 4f558daa-f18f-4a5a-8bef-377f8df8026d · inbound

Seeing Beyond the Scene: Enhancing Vision-Language Models with Interactional Reasoning cites this paper.

Seeing Beyond the Scene: Enhancing Vision-Language Models with Interactional Reasoning Vision-Flan: Scaling Human-Labeled Tasks in Visual Instruction Tuning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T21:46:32.685395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:46:32.685395Z digest=sha256:c5a049c65abc813f8d4c7658906f2565b05c5d062d7d2ebad7cbb9ae097d9ffd

Observation 60240428-c3d7-42de-b785-8280a413f474 · inbound

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation cites this paper.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Vision-Flan: Scaling Human-Labeled Tasks in Visual Instruction Tuning

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:10.263657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:10.263657Z digest=sha256:103765ae85c6aaedc5f6d457ab4e8686dbfb18df33fc20f0ac75579005396977

Observation 3e7f668c-be7f-416a-affd-217cf4189c46 · inbound

FTibSuite: A Comprehensive Resource Suite for Tibetan Vision-Language Modeling cites this paper.

FTibSuite: A Comprehensive Resource Suite for Tibetan Vision-Language Modeling Vision-Flan: Scaling Human-Labeled Tasks in Visual Instruction Tuning

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:33:50.169644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-29T18:32:41.348880Z digest=sha256:50ca78311d735672eb18c2f61925852245bd876f1efdee0d51f1e49fbdebd4e2