Pith. sign in

Paper Citation Record · LEDGER

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation

As of 12 August 2026, this Paper Citation Record lists 67 of 67 outbound references and 2 inbound Pith citation observations for arXiv:2412.16364.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.16364 v1

Coverage vector

measured 67 of 67 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T10:43:08.380678Z

measured 69 of 69 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T15:32:48.356870Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-10T15:32:48.405652Z

Reference resolution

67 of 67 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved66
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5aacc674-890a-4e79-a0c6-81092622eb59 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.015968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.015968Z digest=sha256:1457d11967ace55732d746bc7f73b73937927a55846a70078706b59279da8107

Observation aef6de57-ada9-4950-90b9-189813c046f7 · outbound

This paper cites an unresolved cited work.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-11T10:43:09.644327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T10:43:08.022473Z digest=sha256:e5b9e9c00b15ffb91f082d6a07cad17b80ec07e901c7251c2fbd89864ec82625

Observation 7a91493e-ad36-4acd-b3f5-756b4178e5fa · outbound

This paper cites an unresolved cited work.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.028411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.028411Z digest=sha256:d75961864d92e03bb0341b51dc56880c580c5c5d4d831ca480c874d7890e8320

Observation dbad5b1a-de18-4cd4-95c4-376891cc6205 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.034935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.034935Z digest=sha256:90d9fabaea7ce63e1f4c6a3dac084e5cfe395e0da028416dc58875e35473fbd7

Observation c9aced62-d86c-499a-9f61-c0bbade12ff2 · outbound

This paper cites OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.040942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.040942Z digest=sha256:0389037584a0b6b51248a16e146f4781580598ebef90ab245dbceb47a8ba51ed

Observation 456f2e71-686d-4cbc-9525-5d55b18abacb · outbound

This paper cites Qwen Technical Report.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation Qwen Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.047462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.047462Z digest=sha256:22d30ec40b8b0f55efe144b2d06ba6ffc12c72310faa42d2c745e8652e9fc25f

Observation b1ccb274-a20f-4c6d-b211-3fd36e23abba · outbound

This paper cites an unresolved cited work.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.054234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.054234Z digest=sha256:7b67081e59717273d7c6f20fae793e62a0e64dc179606ed9cf41930c2837f007

Observation 2f1ee924-13e8-41eb-855e-e2ea59aa1595 · outbound

This paper cites an unresolved cited work.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.059750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.059750Z digest=sha256:0fe35b8388a184ecbeb6b26aa7923213ebd8a72c51fc5b6ba101473a9efe36fd

Observation 08368eb0-7df9-4183-892e-b4b7cf4c808d · outbound

This paper cites an unresolved cited work.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.065314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.065314Z digest=sha256:29d09bc24da68e31358e8780c07b123b8fa7b07c6d0bf1d7539f0f59f389c349

Observation f5819198-0513-4c32-8114-88143f42e9d3 · outbound

This paper cites an unresolved cited work.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.071267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.071267Z digest=sha256:3f5d24cebbafb08e555f51da685c20361456d794f2f0ddd952df9b363bb96926

Observation 31d5aa54-06cb-499c-a7a8-d6e5c7dbd311 · outbound

This paper cites ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.076731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.076731Z digest=sha256:d3a50818713661ada95223ed21bfd92e680a4375e8933e4f5fee6a0a9337e43b

Observation f68f39bf-0ae6-4b10-a708-822246bf7dcd · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.082289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.082289Z digest=sha256:ff30b4fc7e643016a3da41ffd93d32a6c3f66f5212c150eb34ea8545b76f6611

Observation ad00a16b-2107-40b4-9d40-b3ff799b7734 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.088107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.088107Z digest=sha256:7a27c21940589692f6f25b52747c2ba91040d4f02bdda46cdce724b203cb38b5

Observation b4953843-a2cf-4ada-84b5-63a33a282ab2 · outbound

This paper cites InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.093445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.093445Z digest=sha256:074d8268161138612eb21f527712422b8faff070f226ef8f61fbc3a449bf8e6b

Observation 03b891d6-79f4-47fe-b99a-8f1b243851b2 · outbound

This paper cites an unresolved cited work.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.099118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.099118Z digest=sha256:729067a6ccad7aba3677eabcabb830840cb489e67db948f16c84a6d324ed6bde

Observation 10e10636-9685-4993-b3c9-bb76777af97c · outbound

This paper cites InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.104693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.104693Z digest=sha256:5d17cb5594edea89e438ad1c23a503924cca2e6b0bcb3172d89c876f587b3a6b

Observation 6e901dd1-c024-4d6e-a312-6a7d435ad9e5 · outbound

This paper cites LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.111217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.111217Z digest=sha256:322083cd001242c54383e3a031402ec5c2b7c0f84700d9991f92d41f2ce39f18

Observation 1363a921-11cf-432e-9ef4-255feebfeea8 · outbound

This paper cites an unresolved cited work.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-11T10:43:09.558346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T10:43:08.116972Z digest=sha256:f2289481f0d570040ad3d87107ad6f394baf940b25fbf4bd41dd90fd4dc318da

Observation fd1c25d7-e7d1-495c-90a9-16cf49b50fc8 · outbound

This paper cites an unresolved cited work.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-11T10:43:09.541404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T10:43:08.121871Z digest=sha256:0c4169e119a616dc98afc629b1048e7b3cf2995cbd975bf24a3670e757538583

Observation 747d26dd-1952-4d61-a751-8fa681cfff86 · outbound

This paper cites FigureQA: An Annotated Figure Dataset for Visual Reasoning.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation FigureQA: An Annotated Figure Dataset for Visual Reasoning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.126593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.126593Z digest=sha256:a9a8c0f8cb1964f9fc97d92ddbe09f6a133e4380478c8e2873c3f76db5446333

Observation 4bad5c29-2f67-490c-be1d-f7bb08f05aa9 · outbound

This paper cites Building and better understanding vision-language models: insights and future directions.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation Building and better understanding vision-language models: insights and future directions

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.131404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.131404Z digest=sha256:8c0cb809bff7471f9b32a0ce3bd27f4d948fb316bd7fb0cda543d743aba0e40e

Observation d522efe4-aeba-4d33-b122-5344e7deb0a1 · outbound

This paper cites What matters when building vision-language models?.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation What matters when building vision-language models?

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.136042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.136042Z digest=sha256:28148c92cfa5b711d246007d52671269ebabe3a72b6772821f642ddfdf2cf5f6

Observation 255da361-d0a9-4079-9334-10dfc9b4a91f · outbound

This paper cites Beyond Scale: The Diversity Coefficient as a Data Quality Metric for Variability in Natural Language Data.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation Beyond Scale: The Diversity Coefficient as a Data Quality Metric for Variability in Natural Language Data

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.140756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.140756Z digest=sha256:cfaefa78c3ad73148a81d6c105d69ac42435b5375eba7759599b13624b57b975

Observation be2a0fa4-5fcd-42d2-8c89-e5a8607e13d9 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation LLaVA-OneVision: Easy Visual Task Transfer

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.145499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.145499Z digest=sha256:cd738525f855183d9ced1ba3d9ef5e993ff64612e34098ff810443d0ea6f84fd

Observation 9fc6c466-f7e4-4222-a81a-f9a6d9f3d38b · outbound

This paper cites Multimodal Foundation Models: From Specialists to General-Purpose Assistants.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation Multimodal Foundation Models: From Specialists to General-Purpose Assistants

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.150494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.150494Z digest=sha256:edca24e45bc89c53caa96740e5c089d2a4dcb87c141f6901f8443e2a725e43a5

Observation 28191912-273f-4eef-b0c1-c64b068cea95 · outbound

This paper cites an unresolved cited work.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.155150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.155150Z digest=sha256:6b3306a7f415abd21005976bbb6132abb9a565e41def35194a2d5385da8b60bb

Observation 67d16506-8c00-4edb-9dcd-180883516aa8 · outbound

This paper cites From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction Tuning.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction Tuning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.159936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.159936Z digest=sha256:7f4ce183a2f1ebf9d663011f4caf07a5f030214ea7048454c690d2d9ab01f78e

Observation 3fee9e36-0d0d-4994-9c58-de20c1335339 · outbound

This paper cites an unresolved cited work.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-11T10:43:09.513539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T10:43:08.164632Z digest=sha256:892881b70f5ffd8ae5187677113f8fd01f8a1588d750b93c53947ced6c9235cb

Observation 42e9e741-e613-40d1-912a-94582f8773d0 · outbound

This paper cites an unresolved cited work.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.169290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.169290Z digest=sha256:5eb580abd807449c0779948d014dbdbbd6a540d0fab498bf93017d6d80aa979b

Observation 8cc2fb93-52c7-4778-892d-c0421725b7e1 · outbound

This paper cites an unresolved cited work.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.174228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.174228Z digest=sha256:686c43075b5870b9ba21711e03b07ac748abd91899bb00c98d7531eee4ebb332

Observation e17d3b1d-91f7-4c8b-9d5a-d2832ddd1a74 · outbound

This paper cites an unresolved cited work.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.179262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.179262Z digest=sha256:4f9b7e731c21ac33513a537aed8cb306813ac1ca80567295b0e975ded9394d45

Observation 8b73e3aa-3a16-4350-bcc1-3e5a57987d8d · outbound

This paper cites an unresolved cited work.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.184299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.184299Z digest=sha256:0c808902f082d4e76699813edd68f92c9e0b5ce53a9f0b08978f6431d22de334

Observation 79c41576-adb8-49a9-8249-12504a124eb8 · outbound

This paper cites an unresolved cited work.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.189517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.189517Z digest=sha256:bd6c504d3f44b1a8f1f0af2073adb75c57f2dfccc3e5349200e0136a0f8753ea

Observation 823076d6-51a6-4cc2-a594-43e901f5f457 · outbound

This paper cites Visual Instruction Tuning.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation Visual Instruction Tuning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.194693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.194693Z digest=sha256:2bf86b071e275283d82af2c17ba5c450495878e8df972de7e0fbff9d98abbd6f

Observation 9d3cded9-7343-4fba-93c2-17c02b5ab6ce · outbound

This paper cites OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.200561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.200561Z digest=sha256:e95d3c565edea7fe1040f2762f3411b9a320b6861b981f382552a071e0dd2289

Observation 68958a47-46d5-4079-9441-3d430f02dc99 · outbound

This paper cites TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.206401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.206401Z digest=sha256:fe9b57f134d0228e6bc208babbd0030c19ccdd53bfbe0b7b244abb073d710e7e

Observation d7a9b546-68f6-4722-add5-1fba7b2724c1 · outbound

This paper cites Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.212053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.212053Z digest=sha256:292b00a0614a1a194e119c02af2c7a81f88b559525e84faceed799d237c2d930

Observation 637a3691-509c-46d7-9e79-adb2080d4e06 · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.217303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.217303Z digest=sha256:8650a8719b9c861ca65e91c204d8bbe8e0dea8db43aaa1e3c0fac30e1a8d3d1c

Observation f08d9e99-b392-4428-9f5f-de5d3983499d · outbound

This paper cites an unresolved cited work.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.223014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.223014Z digest=sha256:febcebdc9138e04ba5c1435563311f76fd6e34f2330da24c0fd5f5cebaec9c08

Observation ae2dc1fe-b717-4c00-97f1-1ca9741ab3a7 · outbound

This paper cites an unresolved cited work.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.227921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.227921Z digest=sha256:5bb717b6aafd2ee98a2db8ae55396eb82e75691f21170d4eecce74d8cd47e707

Observation 5511b1bc-08cb-4d2a-a873-3ee6dc34abfe · outbound

This paper cites an unresolved cited work.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation Unresolved cited work

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.233284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.233284Z digest=sha256:c47855a24934777983f7d9806004936f30d4c5c4aa5d7415e7ce5d1dfa3a3082

Observation 85c2ceb8-1290-4bf9-9aa0-94302cca30b6 · outbound

This paper cites an unresolved cited work.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation Unresolved cited work

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.238556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.238556Z digest=sha256:0a98308526297f8fac79accd2d4d253b17d517e707a446c36497ed87e78dad88

Observation 7c107ef1-2274-45ad-b018-1025e2bd7f32 · outbound

This paper cites an unresolved cited work.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation Unresolved cited work

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.243860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.243860Z digest=sha256:8e7e0626f7cf58a518a3caabe0eb450fd63723cfc30b73c53b8b1fc7dae0dbd2

Observation f2ed88f0-b8b6-4a0c-bcda-7f20f7169886 · outbound

This paper cites an unresolved cited work.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation Unresolved cited work

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.248898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.248898Z digest=sha256:0ff540f5ac08ba170c365a3257ae70e0bbd23c1bc1f8aa837f1c5180863fe8e3

Observation de5324e3-3a32-44a0-a14c-207650ee37e9 · outbound

This paper cites Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.254025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.254025Z digest=sha256:5f86553562a416736895d919ef6d1bf5ee5f107b6672b53a0dc843c959ad4039

Observation fa9e2839-d2f8-4f32-b1b4-61908a1ba1e7 · outbound

This paper cites Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.259469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.259469Z digest=sha256:c1c2538f068ce7540ef61425db9c7581b272986587c871f2437db4ebef0be56a

Observation 83e3ba0f-5a8b-44cf-98c4-82a29bed14f3 · outbound

This paper cites an unresolved cited work.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation Unresolved cited work

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.264439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.264439Z digest=sha256:db33aee6e59c636c6c4d101801ed8f060fa865191f6efcfc355d64f2cc0100db

Observation 084ea4e5-4ca3-4848-b4e4-03a678783b6c · outbound

This paper cites an unresolved cited work.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-11T10:43:09.350050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T10:43:08.269511Z digest=sha256:53f15dfae5081d16313dfa24b5597aae1474e992b446945578722f95d2bac174

Observation 6ee6884b-69f4-481a-9f67-263f29763ccb · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.274466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.274466Z digest=sha256:96743485034e924cdc4b7de61849286525f65ed8b3451ac4dee8b48ad3252409

Observation 660abd36-6da8-426d-a5df-ea81382b457c · outbound

This paper cites an unresolved cited work.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation Unresolved cited work

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.279350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.279350Z digest=sha256:19c9510cbe7e7cf766adcf4ec28f3d1f0b030fbd5c009fb5b89e2bb314aacb7c

Observation db90e8f6-2ebb-4bde-8364-563eaeb3d8d3 · outbound

This paper cites an unresolved cited work.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation Unresolved cited work

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.284180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.284180Z digest=sha256:276ef81e61c6b651ed7820c1fc93e95175b54dde61790f168856475bc9eea374

Observation 3504fa73-4fd3-4924-a1f6-f4c77c5adda4 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.289122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.289122Z digest=sha256:e03a1a03abd2b5f39f2dc00fc82edb0aa29062e34d76e2f724e83055b9132d5a

Observation 37a2c492-2125-4336-b4b9-052f6a34d4bb · outbound

This paper cites Self-Instruct: Aligning Language Models with Self-Generated Instructions.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation Self-Instruct: Aligning Language Models with Self-Generated Instructions

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.294005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.294005Z digest=sha256:dae614ff0f77a53b281c87ab32a73631f439fcc1b93ef3a38b462cedc4873eae

Observation 819bbdeb-4eee-42cb-8852-9622a2f95d9b · outbound

This paper cites an unresolved cited work.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-08-11T10:43:09.307105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T10:43:08.298833Z digest=sha256:4d5ab2d6f85f703bd319ca03803bfea07173c4c334b154a2f60693fd94672b66

Observation 1bb51277-e1c3-4935-83b6-ec35b11cdbb0 · outbound

This paper cites WizardLM: Empowering large pre-trained language models to follow complex instructions.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation WizardLM: Empowering large pre-trained language models to follow complex instructions

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.304211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.304211Z digest=sha256:fd1405d8e349c2a9d5ef942307be5179cfba13538d767fe55d0ba028414b89d7

Observation bd986c29-9567-499f-ab33-9bc948658b71 · outbound

This paper cites LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.309295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.309295Z digest=sha256:04001bdd33046fc7bd199be42da3159928187c43311e0635b2316a8b860e4cb7

Observation ef38d3f2-5405-4ae5-bee0-17e04f7452bb · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.314525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.314525Z digest=sha256:43029a42b5f99f1bab9d8c00417f1fc731413fc6b0d74813190bd086e0321852

Observation bf827e25-0aed-4782-8627-8dfd9377e165 · outbound

This paper cites UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.320223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.320223Z digest=sha256:10b555140944352d009c563f213301a3bb68b72b90efe35c105ae3ef3ae93fe7

Observation bba694e8-6c14-4073-a46e-480ac2bcdcea · outbound

This paper cites mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.330976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.330976Z digest=sha256:b85851109ef7cb4956aab16354d4000ec4fb01e91350b0394f8d0c5bcdec66c8

Observation 9e3dde9d-e7b0-4f54-b686-4a4033c772ec · outbound

This paper cites LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.336017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.336017Z digest=sha256:71ceb63f16297fc6115b63280caf7270663fead627d7cfa0b3d35b1c79ce03c5

Observation d493ed16-3acd-4a52-b132-6fa5eed4ebed · outbound

This paper cites an unresolved cited work.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation Unresolved cited work

Reference 62

Resolution
unresolved
raw_fallback, observed 2026-08-11T10:43:09.288560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T10:43:08.341568Z digest=sha256:0944df8f627016af6ba4e9ca0fbb8c214c81f8f764ce8b300c96a12a53974a42

Observation 039c9c1c-7118-4aa7-b692-2ad34cec58f7 · outbound

This paper cites an unresolved cited work.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation Unresolved cited work

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.346668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.346668Z digest=sha256:44852fd8d74651a5bd881144f551f2981ca7d5236f2e6c5c88a5281865d25072

Observation 76e70a1b-40d9-4db0-b8c3-58ef0c8b9815 · outbound

This paper cites LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.351843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.351843Z digest=sha256:6afbf972aacf2c40432b43a005140ec768040b818a9cf391debe86ebff7dadd4

Observation 9f71f3fe-2089-4064-b1c9-612f4db72381 · outbound

This paper cites Genixer: Empowering Multimodal Large Language Models as a Powerful Data Generator.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation Genixer: Empowering Multimodal Large Language Models as a Powerful Data Generator

Reference 65

Resolution
verified exact
local_arxiv, observed 2026-08-11T10:43:08.445289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T10:43:08.357447Z digest=sha256:a0b1300850563c8c497de78b629af639cdb8cb47d60c7feb24a91f9853232045

Observation f5749c3a-7a2a-4bff-86e9-52f2ca6bdcea · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.367718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.367718Z digest=sha256:6a666df3c7b0ec910bc0550f8b1a29975c4256e53a2cb0f84c275ed63642e932

Observation 746f9fb3-8af6-47cc-a585-49c169dd95e8 · outbound

This paper cites online" 'onlinestring :=.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation online" 'onlinestring :=

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.374718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.374718Z digest=sha256:ee50a99228e9461ea50179d2c26d2013932372ceee69036d288a80d4e9ccabdd

Observation 240ab25e-bb93-4ba0-9e60-afb92eeac574 · outbound

This paper cites write newline.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation write newline

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.380678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.380678Z digest=sha256:cdce294d7f5e4e7cb21afebcb4a9fe472a4cde1d131e2414226e62ced088df2f

Pith citing papers

Observation 7b2f1c54-e03b-44af-b3cf-ae100fa0aa5e · inbound

The Breeze 2 Herd of Models: Traditional Chinese LLMs Based on Llama with Vision-Aware and Function-Calling Capabilities cites this paper.

The Breeze 2 Herd of Models: Traditional Chinese LLMs Based on Llama with Vision-Aware and Function-Calling Capabilities A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-08-10T15:32:48.411909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T15:32:48.356870Z digest=sha256:80fac137a3d96c84c338e7293334d2e85201135b67b4419c3ce212bbb9b38c31

Observation 1a5d8708-20ca-457e-91b6-9079ea167a94 · inbound

A Survey on Evaluating Quality and Trustworthiness in LLM-Generated Data cites this paper.

A Survey on Evaluating Quality and Trustworthiness in LLM-Generated Data A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation

Reference 280

Resolution
unresolved
no resolver link, observed 2026-08-03T08:15:38.147584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T08:15:38.147584Z digest=sha256:2bf60a732dc98f11ad1343ffd12c0adac0d9f113b80f88443212b646f7140fb2