Pith. sign in

Paper Citation Record · LEDGER

Diffusion Instruction Tuning

As of 13 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 0 inbound Pith citation observations for arXiv:2502.06814.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.06814 v2

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T11:20:02.961462Z

measured 39 of 39 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

39 of 39 outbound references displayed

  • verified exact2
  • verified fuzzy3
  • unresolved34
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d338fb1d-b675-43e3-9579-d667c57762b9 · outbound

This paper cites Quantifying Attention Flow in Transformers.

Diffusion Instruction Tuning Quantifying Attention Flow in Transformers

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T11:20:02.827545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:20:02.827545Z digest=sha256:c0002fd465637d84c587542695484f11fa9914f7b5b2fdd4a384b79823c8394f

Observation 006e4a96-c10a-43d1-a547-73b0e193cf74 · outbound

This paper cites an unresolved cited work.

Diffusion Instruction Tuning Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:20:03.487975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-09T11:20:02.954871Z digest=sha256:8d9239e95afa99727e753fac791dba372d752c79cceb22b2c9eca34b10f625fe

Observation 753a6847-e0fb-4ca6-bff8-6e4114a0a820 · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

Diffusion Instruction Tuning Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-09T11:20:02.839886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:20:02.839886Z digest=sha256:9c91d614dcb751a96fc97efc075d00105d75d36963a986082a442a39566b9cbf

Observation d3e91e9b-6ceb-4084-807e-582139deaa96 · outbound

This paper cites Locality Alignment Improves Vision-Language Models.

Diffusion Instruction Tuning Locality Alignment Improves Vision-Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-09T11:20:02.847466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:20:02.847466Z digest=sha256:16d60f3af7ee99c405065ccdb310084aca9d9a53242b2f0fe2595dcf83361065

Observation 97079387-3cba-4818-9ce4-1de22dc43b31 · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

Diffusion Instruction Tuning InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-09T11:20:02.851199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:20:02.851199Z digest=sha256:26a2fad0d65e8bad26243a76b3d067b29ff94d6dfeea442e22648ed3cce15a4c

Observation fd0b7c91-341b-4bd7-b841-e62f05ec2dd8 · outbound

This paper cites DeepSeek-V3 Technical Report.

Diffusion Instruction Tuning DeepSeek-V3 Technical Report

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-09T11:20:02.854961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:20:02.854961Z digest=sha256:ccbd9c16a821870dd0dd896c2d4789646205c70c24421391eccf5ffadcff7fda

Observation bd5dadf8-0c1c-4964-91ca-f7ea139c9adb · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Diffusion Instruction Tuning An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T11:20:02.858643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:20:02.858643Z digest=sha256:2b407000782de52f232f95fe0a7bf5b71a0e838ec5b60a3a11ab38fd2ff6bd75

Observation 86928694-8f44-4b8e-8ac7-171e879ad086 · outbound

This paper cites The Llama 3 Herd of Models.

Diffusion Instruction Tuning The Llama 3 Herd of Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-09T11:20:02.861920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:20:02.861920Z digest=sha256:72ad93c2fd12e65f425718132c8488edcc0fc68b7e32e5cab30c284cddcb7a4a

Observation 5da7329e-dcb0-4ae5-ab1c-9af495ef7fb5 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

Diffusion Instruction Tuning MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-09T11:20:02.865354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:20:02.865354Z digest=sha256:08b8140a3086e2c3065a46fb17a08d693996d8eae37c4e36660de349ceea7b12

Observation 2ec41dfa-e504-4b23-810b-c228404d5359 · outbound

This paper cites LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model.

Diffusion Instruction Tuning LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T11:20:02.868705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:20:02.868705Z digest=sha256:0f844ccdd11aa26cb85c8388be3fd1c10267416c97a40e24e4ba9c260d096dd3

Observation d4cde6a4-3798-4820-b931-7571094cf985 · outbound

This paper cites An Image is Worth Multiple Words: Discovering Object Level Concepts using Multi-Concept Prompt Learning.

Diffusion Instruction Tuning An Image is Worth Multiple Words: Discovering Object Level Concepts using Multi-Concept Prompt Learning

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-09T11:20:03.328169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-09T11:20:02.878657Z digest=sha256:c86f521465424fa0b11f55c64e11d96bb0225259f96faad7f5f21ed441d6a9cb

Observation ce78c41b-9411-4a30-9193-69c056eb575f · outbound

This paper cites Yes I’m not able to provide a name for the person in this picture.

Diffusion Instruction Tuning Yes I’m not able to provide a name for the person in this picture

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:20:03.478271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-09T11:20:02.958257Z digest=sha256:08415dc6b00a8cab58fbc00de063592364fff3dcd9f6a4d71d729a3f802176f2

Observation 29ba56bc-ad06-4e2f-9a4f-2abd4030c8c3 · outbound

This paper cites A diagram is worth a dozen images.

Diffusion Instruction Tuning A diagram is worth a dozen images

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:20:03.513418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-09T11:20:02.885151Z digest=sha256:68b29d5e5d68ca687a2eebf799b177a1c24bbd6ea0598dac3c35952f0f499305

Observation 8cef1745-d6ec-4c29-8da3-e693cbf9c072 · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

Diffusion Instruction Tuning SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-09T11:20:02.888308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:20:02.888308Z digest=sha256:1953d52a770a4f4441d87603094bedc42e84583d2ba8b8511624c18d8a7cb5e9

Observation 8ea88340-2eda-4ee4-b965-2c0c86a1074c · outbound

This paper cites HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models.

Diffusion Instruction Tuning HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-09T11:20:02.895321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:20:02.895321Z digest=sha256:6227e81990b82068fd811ca7bb6e70a45c07e0cb91c5afb5c566ecd3cc245b97

Observation 05b1e9d4-cfdb-4221-b7d0-83d0c64e4274 · outbound

This paper cites WorldMedQA-V: a multilingual, multimodal medical examination dataset for multimodal language models evaluation.

Diffusion Instruction Tuning WorldMedQA-V: a multilingual, multimodal medical examination dataset for multimodal language models evaluation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T11:20:02.902422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:20:02.902422Z digest=sha256:e96133c41278699cd0fe63c4acc066543490db1ee60c303078b4902c613d2457

Observation 36a53c59-dd2e-4c85-a3f2-2498162596d4 · outbound

This paper cites K., and Chakraborty, A.

Diffusion Instruction Tuning K., and Chakraborty, A

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-09T11:20:02.905887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:20:02.905887Z digest=sha256:c580e366f87085a56a3aa3058ddaa7c68a01abed1bbf6889eb5f9ff502e36394

Observation dedd01ec-d94b-4eb1-82aa-2a9b642f4438 · outbound

This paper cites Null-text Inversion for Editing Real Images using Guided Diffusion Models.

Diffusion Instruction Tuning Null-text Inversion for Editing Real Images using Guided Diffusion Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-09T11:20:02.909198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:20:02.909198Z digest=sha256:9eded0da8e6ffbb2ac38f90177a7997060d842f4e3a44e2536ccb639c537b96e

Observation 533938b4-b2bb-4b3a-bd32-4cf798a355d5 · outbound

This paper cites Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders.

Diffusion Instruction Tuning Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-09T11:20:02.912656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:20:02.912656Z digest=sha256:1f92593d6c04ff451511976d8f845dd15f644c8e8c15d85b21078aef5675c85a

Observation ee30dcd7-c7ae-46b0-88c6-2244f47e3098 · outbound

This paper cites LMFusion: Adapting Pretrained Language Models for Multimodal Generation.

Diffusion Instruction Tuning LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-09T11:20:02.915960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:20:02.915960Z digest=sha256:eeda324f2c126c6cd2c4e6697061859a9d7394c8f7dc1de07af54b227a682854

Observation 98d4ebac-f61e-41de-8b99-dea333610fa5 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Diffusion Instruction Tuning Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-09T11:20:02.920209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:20:02.920209Z digest=sha256:ed1bcbc69f463ae8a5058ce8400c0634d11a6bd8ca2f070e0d3c5d26c91d8919

Observation 4eb3cb22-1e42-4136-b9b9-fa8ed3b8d626 · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

Diffusion Instruction Tuning Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-09T11:20:02.923770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:20:02.923770Z digest=sha256:76e83725f4d8dea9302370afc242bf9dd554f53ca137fc9a935acf484cd256fc

Observation 831adb08-f292-4032-9ce2-52ce6f47f6c8 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

Diffusion Instruction Tuning MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-09T11:20:02.930797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:20:02.930797Z digest=sha256:af2005353ed10439c0b8d099fc8519de857cf998f581d01590e9c3a8ccd8497c

Observation 93323bd5-84d2-4465-a886-e599e250c5b5 · outbound

This paper cites Ferret: Refer and Ground Anything Anywhere at Any Granularity.

Diffusion Instruction Tuning Ferret: Refer and Ground Anything Anywhere at Any Granularity

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-09T11:20:02.934093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:20:02.934093Z digest=sha256:0153866be9d458bb162c59234b9048185b7eaf2aec21efc0eed76ea56fd2a401

Observation fdcb8bb8-c620-4063-b910-8d2c30a36e4f · outbound

This paper cites Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model.

Diffusion Instruction Tuning Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-09T11:20:02.940685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:20:02.940685Z digest=sha256:b45fa46657e2be7b85e488665043813e943c93ccc66e5c4aad0e57f03923e40c

Observation 9b2dabd3-6c42-4bed-840f-89c5436cf7a0 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Diffusion Instruction Tuning MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-09T11:20:02.944165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:20:02.944165Z digest=sha256:00b45a2b6a330b7c0774c8cfb1470b32663324f0d7514854cd9e1d8b83fef413

Observation 282b4b10-ff46-44ee-934c-395677431439 · outbound

This paper cites MoVA: Adapting Mixture of Vision Experts to Multimodal Context.

Diffusion Instruction Tuning MoVA: Adapting Mixture of Vision Experts to Multimodal Context

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-09T11:20:02.947795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:20:02.947795Z digest=sha256:ed1442e6845cb0f37e6d47633dc09d43a4b5d4fea214dce54eacb9917a489cc0

Observation 083baf5f-7022-403b-8cd9-cfb287ed4f02 · outbound

This paper cites Lavender-Llama3.2-11B occasionally refuses to answer questions for privacy reasons, resulting in a FALSE score and reduced performance on MME as shown in Figure.

Diffusion Instruction Tuning Lavender-Llama3.2-11B occasionally refuses to answer questions for privacy reasons, resulting in a FALSE score and reduced performance on MME as shown in Figure

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:20:03.468256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-09T11:20:02.961462Z digest=sha256:07232e13e7303e6ec20578285c208347e51773d4fff7864e11c8fab003b09f6a

Observation 94208058-f678-4688-94a6-7e46de1a8519 · outbound

This paper cites Rlaif-v: Aligning mllms through open-source ai feed- back for super gpt-4v trustworthiness.

Diffusion Instruction Tuning Rlaif-v: Aligning mllms through open-source ai feed- back for super gpt-4v trustworthiness

Reference 2014

Resolution
unresolved
no resolver link, observed 2026-08-09T11:20:02.937592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:20:02.937592Z digest=sha256:c65c5995b38afa036ca95797dd9f363f0ddec5c48e3d80e45c7a6dbdadcdc832

Observation ae997530-f339-4175-a6d3-29970b2e5456 · outbound

This paper cites Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling.

Diffusion Instruction Tuning Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-09T11:20:02.843537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:20:02.843537Z digest=sha256:4f5e606fba8acbda9dd9e7bea2da13f0470290b1122b7bc75f74ed2798ab93a9

Observation 1dcd946e-6b12-49be-819c-8c930699e662 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Diffusion Instruction Tuning Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-09T11:20:02.927306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:20:02.927306Z digest=sha256:12692f0425252d1f4b4d535ee2912ff8ba72dc3aa70b6271f884c50e8b6095a2

Observation 1e2c2531-220b-4ead-ad40-fabf444c2d55 · outbound

This paper cites Generative Visual Instruction Tuning.

Diffusion Instruction Tuning Generative Visual Instruction Tuning

Reference 2018

Resolution
verified exact
local_arxiv, observed 2026-08-09T11:20:03.351075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-09T11:20:02.871867Z digest=sha256:0b742a79ac1e09eae6f080a8fc3694c7f999e73e8e19e8ee83d7b36252c5ea34

Observation 9f27da9b-4f64-418e-8d25-7da08b19e408 · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

Diffusion Instruction Tuning ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-09T11:20:02.898927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:20:02.898927Z digest=sha256:f9ec36a68af6e3ece981732fceb2799dc9ff8d8be55fbce5c4a9d89e03620faf

Observation f44f9665-ad84-49ad-a442-55c9e4ea713b · outbound

This paper cites an unresolved cited work.

Diffusion Instruction Tuning Unresolved cited work

Reference 2020

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:20:03.497726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-09T11:20:02.951042Z digest=sha256:e660a989dccf9ab2055b0e79be7c1012c142e2c564fd5c963ac40ae62f1b4020

Observation c7b5b60e-23e2-47cf-953a-30ef807b3ed0 · outbound

This paper cites From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models.

Diffusion Instruction Tuning From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-09T11:20:02.875384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:20:02.875384Z digest=sha256:e1b02075433ff4c742186be949e7175739d1c0c5726d77aa83478f1ed9cdccae

Observation 887f7ad3-bf50-4b40-a7bc-c69fb48be241 · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

Diffusion Instruction Tuning Evaluating Object Hallucination in Large Vision-Language Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-09T11:20:02.891980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:20:02.891980Z digest=sha256:9e774a8127b7dec3f52b4f0db50c9da9aa78f284c7021619db3ac7a016c99e37

Observation a4ca4871-c7c3-4e9d-bf23-56ed5f6ce608 · outbound

This paper cites Qwen Technical Report.

Diffusion Instruction Tuning Qwen Technical Report

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-09T11:20:02.836120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:20:02.836120Z digest=sha256:6d067e5332bb47ca1819d4385faec0bad278c6c48643e67d3f447972e362b784

Observation 02df6427-1317-4c79-bb21-8109ec65a7bd · outbound

This paper cites Pixtral 12B.

Diffusion Instruction Tuning Pixtral 12B

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-09T11:20:02.832217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:20:02.832217Z digest=sha256:e18d78a7bdd9489a1e434c74de51a3c1c73df5e4523069b0451c85d7349152b5

Observation dae22de9-8599-466a-b913-24ff7e4a6a9d · outbound

This paper cites Prismatic VLMs: Investigating the Design Space of Visually-Conditioned Language Models.

Diffusion Instruction Tuning Prismatic VLMs: Investigating the Design Space of Visually-Conditioned Language Models

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-09T11:20:02.881929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:20:02.881929Z digest=sha256:a3fcd4ce3d58ac92abee26f60a118ca7a58da7bb8da9443bdd6b6231b4cf7d9f

Pith citing papers

No inbound Pith citation observations are available.