Pith. sign in

Paper Citation Record · LEDGER

Vision-Flan: Scaling Human-Labeled Tasks in Visual Instruction Tuning

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2402.11690.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.11690 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T20:06:36.670967Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T18:33:50.167861Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f13163c8-e6d7-4c9f-a577-0302221d0768 · inbound

SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation cites this paper.

SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation Vision-Flan: Scaling Human-Labeled Tasks in Visual Instruction Tuning

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:48:36.226183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T22:48:36.010306Z digest=sha256:b556e6c13c5816f03a5d021ca68368b5742abf0a2c61cca6fac4f0640606edf4

Observation 172ad1be-4fdb-4bca-990f-31e145282e45 · inbound

LLaVA-OneVision: Easy Visual Task Transfer cites this paper.

LLaVA-OneVision: Easy Visual Task Transfer Vision-Flan: Scaling Human-Labeled Tasks in Visual Instruction Tuning

Reference 145

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:23:49.510641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:bd3b780eafa95551c34bc6fad41440ddfec5779a20c44fd7d9f6b05599ade622

Observation 22adde2f-2960-441e-a9bb-fbb12cdc4497 · inbound

TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types cites this paper.

TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types Vision-Flan: Scaling Human-Labeled Tasks in Visual Instruction Tuning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T20:06:36.670967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:06:36.670967Z digest=sha256:da752044333fa7ec0b70a69eb159064b7fc31312f7b59332dbd2388d9ed63ec7

Observation 60240428-c3d7-42de-b785-8280a413f474 · inbound

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation cites this paper.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Vision-Flan: Scaling Human-Labeled Tasks in Visual Instruction Tuning

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:10.263657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:10.263657Z digest=sha256:2ee81a76b7e1d0594783b1e036544f53d38379d722323d32f43a9f5ac2359831

Observation 3e7f668c-be7f-416a-affd-217cf4189c46 · inbound

FTibSuite: A Comprehensive Resource Suite for Tibetan Vision-Language Modeling cites this paper.

FTibSuite: A Comprehensive Resource Suite for Tibetan Vision-Language Modeling Vision-Flan: Scaling Human-Labeled Tasks in Visual Instruction Tuning

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:33:50.169644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T18:32:41.348880Z digest=sha256:113ae1da9a3a6a9ded0135b5d6326114deee10ef57ba9bc5bd48f2b9044fa9e2