Pith. sign in

Paper Citation Record · LEDGER

Instruction-augmented Multimodal Alignment for Image-Text and Element Matching

As of 18 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 0 inbound Pith citation observations for arXiv:2504.12018.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.12018 v1

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:40:26.661269Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

45 of 45 outbound references displayed

  • verified exact0
  • verified fuzzy12
  • unresolved33
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a039c636-3c55-4eaa-80e4-ffcb035b111d · outbound

This paper cites GPT-4 Technical Report.

Instruction-augmented Multimodal Alignment for Image-Text and Element Matching GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:26.393404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:26.393404Z digest=sha256:646030e94d57cf9f1996d324bcbce9beefa9123a9ba7e9dc64a165f92a97b8aa

Observation e60fa8f0-7b8c-483a-8b78-62f47d8d10ec · outbound

This paper cites Kandinsky 3.0 Technical Report.

Instruction-augmented Multimodal Alignment for Image-Text and Element Matching Kandinsky 3.0 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:26.400427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:26.400427Z digest=sha256:75d4b30edab203279a31873da4384757e5cd5c3a118715c9b5316ff67c739dd7

Observation 1da81de3-5e3b-410c-8e34-777f46280c87 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Instruction-augmented Multimodal Alignment for Image-Text and Element Matching Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:26.407193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:26.407193Z digest=sha256:082780c80215849e27ebcf613435599bfe5ef4ad9482f24764c49067e6fb20db

Observation b6942d79-9386-4c24-b0e9-c94d1b18d996 · outbound

This paper cites Qwen2.5-VL Technical Report.

Instruction-augmented Multimodal Alignment for Image-Text and Element Matching Qwen2.5-VL Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:26.413539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:26.413539Z digest=sha256:20e4f9ba252984e096ca8141fed2968ad7a4da56c12d086adf7cd7082caa8f14

Observation e51052a3-f2ff-4b79-8db3-b23f4fd2cf92 · outbound

This paper cites an unresolved cited work.

Instruction-augmented Multimodal Alignment for Image-Text and Element Matching Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-16T12:40:27.475336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:40:26.419339Z digest=sha256:80dc9773d24522d9161a6bd5f0191302d18ac153920ae7cb00e3e465555f9712

Observation 14ca790b-6790-4ff4-bdd8-0f5c54d2038e · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

Instruction-augmented Multimodal Alignment for Image-Text and Element Matching Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:26.424899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:26.424899Z digest=sha256:8dc6ce12fbf072002e334fc02a7af09e84778146bf79dee5029598757475ce0b

Observation 6d4418fb-e54f-4c40-8b10-cfdbdf92f72d · outbound

This paper cites Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks.

Instruction-augmented Multimodal Alignment for Image-Text and Element Matching Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:26.431686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:26.431686Z digest=sha256:f29186119d229f99efd704a51b7508fe687a1b1ce479535d41d0715890a5767d

Observation 82dc28c0-d4dc-4fbd-800d-7edc8d30ce76 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Instruction-augmented Multimodal Alignment for Image-Text and Element Matching An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:26.437504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:26.437504Z digest=sha256:3fe2d356d66237388d2d67e9207bbdba58cbec69bc8f3c7798f3b9585ed3acf1

Observation 44858ea2-b495-49c0-b23d-38c854cc6d35 · outbound

This paper cites Dreamina.

Instruction-augmented Multimodal Alignment for Image-Text and Element Matching Dreamina

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:40:27.441607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:40:26.444201Z digest=sha256:fee3c6f7c2766d630d4da96abffe0b0d7be31ea7c39c17e0ce97ffcbcd1430cf

Observation c5fdb2b8-f012-482b-af9b-96683ee950a4 · outbound

This paper cites Scaling recti- fied flow transformers for high-resolution image synthesis.

Instruction-augmented Multimodal Alignment for Image-Text and Element Matching Scaling recti- fied flow transformers for high-resolution image synthesis

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:26.450097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:26.450097Z digest=sha256:3c6d0b1648d77838c4cfade6828f1077b6788e0d2d8fad695ed00887d4e6d439

Observation a77dea49-59fe-459e-b1c2-2fc6a05cd78c · outbound

This paper cites Ranni: Taming Text-to-Image Diffusion for Accurate Instruction Following.

Instruction-augmented Multimodal Alignment for Image-Text and Element Matching Ranni: Taming Text-to-Image Diffusion for Accurate Instruction Following

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:26.455283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:26.455283Z digest=sha256:ce09b749b4f7ee54fd08b16682a5fed658e10de8d1e4d9fae33ff78010139227

Observation 73ddecb0-e4a0-4d62-b9b0-dec03dc27092 · outbound

This paper cites EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation.

Instruction-augmented Multimodal Alignment for Image-Text and Element Matching EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:26.460713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:26.460713Z digest=sha256:a597284074d94eb6b96c816b1f1d9caa39515b5899c30fd12bfae6c0c92d8056

Observation 5b129cfa-67e8-4506-b066-9bffb096fd3a · outbound

This paper cites NTIRE 2025 challenge on text to image generation model quality assess- ment.

Instruction-augmented Multimodal Alignment for Image-Text and Element Matching NTIRE 2025 challenge on text to image generation model quality assess- ment

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:26.466405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:26.466405Z digest=sha256:968b177a206369b7cbd3de5336515a082fc5d6c2a0030a8807c327c704217b28

Observation 0efaf066-a7b3-4984-84e1-8003036adc20 · outbound

This paper cites MARS: Mixture of Auto-Regressive Models for Fine-grained Text-to-image Synthesis.

Instruction-augmented Multimodal Alignment for Image-Text and Element Matching MARS: Mixture of Auto-Regressive Models for Fine-grained Text-to-image Synthesis

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:26.471660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:26.471660Z digest=sha256:96d0656b3140b9f192f7b48463a485cef839e8e328963fb03584014278cb66ca

Observation 0b509637-dc17-4219-ba9b-4f04119afc3d · outbound

This paper cites CLIPScore: A Reference-free Evaluation Metric for Image Captioning.

Instruction-augmented Multimodal Alignment for Image-Text and Element Matching CLIPScore: A Reference-free Evaluation Metric for Image Captioning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:26.477919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:26.477919Z digest=sha256:565ce17c398907d0180eb38826ace4487817ccd3fd121c34dbaab9f4f9f3dda5

Observation e6203e11-aa58-4904-a891-662c76648e7d · outbound

This paper cites Midjourney.

Instruction-augmented Multimodal Alignment for Image-Text and Element Matching Midjourney

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:40:27.396441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:40:26.483496Z digest=sha256:91ef856ab3a2ea08dd4c2bd7190dafe8bb68e21d9eab3e22cb9ecebbeeafe0e8

Observation 20e112cf-ee31-475f-9882-5693aa838f24 · outbound

This paper cites Lora: Low-rank adaptation of large language models.

Instruction-augmented Multimodal Alignment for Image-Text and Element Matching Lora: Low-rank adaptation of large language models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:40:27.377160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:40:26.489037Z digest=sha256:5a517b42de762af4073547b2650585a99d5b8ffb665bb79b3242dd8fa152e342

Observation 7aa9f23b-e078-4099-8f62-4668b2dcc5c2 · outbound

This paper cites Tifa: Accurate and interpretable text-to-image faithfulness evaluation with question answering.

Instruction-augmented Multimodal Alignment for Image-Text and Element Matching Tifa: Accurate and interpretable text-to-image faithfulness evaluation with question answering

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:40:27.358056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:40:26.495712Z digest=sha256:c18e21bf5545c46e7ec86f42888f031cbfe80c0bb7bf97b29d898c6d611ab116

Observation 228484db-f87f-4904-932b-b0163dae7e69 · outbound

This paper cites Pick-a-pic: An open dataset of user preferences for text-to-image generation.

Instruction-augmented Multimodal Alignment for Image-Text and Element Matching Pick-a-pic: An open dataset of user preferences for text-to-image generation

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:40:27.338434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:40:26.500934Z digest=sha256:15235ad3e322c31119bbe3c846c4e5b68db7c86a6096d0b842c223b05b32039c

Observation 41b50943-5093-4153-86d8-35780eaca9f1 · outbound

This paper cites Evaluating and improving composi- tional text-to-visual generation.

Instruction-augmented Multimodal Alignment for Image-Text and Element Matching Evaluating and improving composi- tional text-to-visual generation

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:40:27.318846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:40:26.505987Z digest=sha256:511626c80f1cf3a3044f07caff2ea6d37023b6aef3582af03518d2fc6d4a6f43

Observation 94a28a47-d2fd-4763-b475-f139e210616e · outbound

This paper cites Playground v2.5: Three Insights towards Enhancing Aesthetic Quality in Text-to-Image Generation.

Instruction-augmented Multimodal Alignment for Image-Text and Element Matching Playground v2.5: Three Insights towards Enhancing Aesthetic Quality in Text-to-Image Generation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:26.511573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:26.511573Z digest=sha256:bf977aec22a209235d73f337dbe376173fca2ece3e1492412f82ca7c580b118d

Observation e3044d9d-2f82-4d73-9e24-07b074d0c87c · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Instruction-augmented Multimodal Alignment for Image-Text and Element Matching Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:26.516854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:26.516854Z digest=sha256:a9dac0c24a99d0a95b40825918d682c00753ca40c876dd36114febae67a6d9c2

Observation 7ac6b9e6-7cf5-4f49-9e6d-cb4c928b463f · outbound

This paper cites Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding.

Instruction-augmented Multimodal Alignment for Image-Text and Element Matching Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:26.523312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:26.523312Z digest=sha256:d4473625f1b5766942ff0d004a6989a091e9bdb5318c74cdb6860bb818d27b27

Observation 700c6fc0-ddbf-46e2-ac51-a1c60f1bdd0a · outbound

This paper cites Visual instruction tuning.

Instruction-augmented Multimodal Alignment for Image-Text and Element Matching Visual instruction tuning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:26.528469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:26.528469Z digest=sha256:19b453d72984fe3bedcbabda5185669c274ea16bba6ce702b773462db24052f6

Observation f90556e0-aeb3-4d94-91a9-d8b7d5583717 · outbound

This paper cites Improved baselines with visual instruction tuning.

Instruction-augmented Multimodal Alignment for Image-Text and Element Matching Improved baselines with visual instruction tuning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:26.534548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:26.534548Z digest=sha256:0b055ddb7a303c5cf7abc49cfda4d8081769cee134cbc75d9535993cc3493997

Observation 76aa0313-43c2-454b-aa84-1bd9d64fe4b7 · outbound

This paper cites Llavanext: Improved reasoning, ocr, and world knowledge, 2024.

Instruction-augmented Multimodal Alignment for Image-Text and Element Matching Llavanext: Improved reasoning, ocr, and world knowledge, 2024

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:26.540586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:26.540586Z digest=sha256:41bbe61e6584431a8fc0287b486ccbe6f19c9ef577416930d51b8f29ec22b34c

Observation 6487a08c-b4e9-4dc1-bfc0-013f5d676c53 · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

Instruction-augmented Multimodal Alignment for Image-Text and Element Matching DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:26.545896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:26.545896Z digest=sha256:72c3404c2b5bc5ad90c18196c99155b7b85053be3240c63c894e051416347efc

Observation 4cef6bf6-12ac-4b35-b56b-21fb16ccee8a · outbound

This paper cites Ovis: Structural Embedding Alignment for Multimodal Large Language Model.

Instruction-augmented Multimodal Alignment for Image-Text and Element Matching Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:26.552528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:26.552528Z digest=sha256:b61d85fb08e576f71f4cf2a1ca167fab68e8c44c6e923055c0bfe2776334c834

Observation 191beef7-3212-43e2-93c6-bb505d01d53f · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

Instruction-augmented Multimodal Alignment for Image-Text and Element Matching Learn- ing transferable visual models from natural language super- vision

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:26.559152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:26.559152Z digest=sha256:77bda0708fef910bbc85019f8a13fc356a2777f5818e58fc48ab4d110888eb27

Observation 0cb6d91b-0910-43c7-9e2a-97883407daf7 · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

Instruction-augmented Multimodal Alignment for Image-Text and Element Matching Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:26.565081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:26.565081Z digest=sha256:9d1e3d5db1961a22343dce414881124a32d3c64f601a985eb9da8b3f0858dece

Observation 24505366-057f-4bdc-974c-7527ec0aa346 · outbound

This paper cites High-resolution image syn- thesis with latent diffusion models.

Instruction-augmented Multimodal Alignment for Image-Text and Element Matching High-resolution image syn- thesis with latent diffusion models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:26.572519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:26.572519Z digest=sha256:a2f1c581edcfaac6f8030e75dd77d71256a2c871f5f265ccefaaff47c3d47a29

Observation 677aa1d1-1f91-4a58-b64d-8021cd5b4c79 · outbound

This paper cites Adversarial diffusion distillation.

Instruction-augmented Multimodal Alignment for Image-Text and Element Matching Adversarial diffusion distillation

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:40:27.226754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:40:26.577850Z digest=sha256:5b7d7779332389e21b21d5c25a381e0e0cb7d557e0242ab2ac0f58341d930dc7

Observation cc23c617-bf0b-4ca4-9609-ba611575f1ed · outbound

This paper cites Hugginggpt: Solving ai tasks with chatgpt and its friends in huggingface.

Instruction-augmented Multimodal Alignment for Image-Text and Element Matching Hugginggpt: Solving ai tasks with chatgpt and its friends in huggingface

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:40:27.208423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:40:26.583786Z digest=sha256:d525e456b278c0910fd7908036c77ca64be25cbaeda79bffe7fab41fa3bc384a

Observation 83dae9d7-bcc8-4359-a459-81d9aa2ac634 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Instruction-augmented Multimodal Alignment for Image-Text and Element Matching Gemini: A Family of Highly Capable Multimodal Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:26.589564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:26.589564Z digest=sha256:972559400e7fd9fa89088f4a5287aae2795bf4ab7fefdfce60ad6b6023519174

Observation d12f04cc-a6a6-4592-ad0a-8fa2761096de · outbound

This paper cites Kolors: Effective training of diffusion model for photorealistic text-to-image synthesis.

Instruction-augmented Multimodal Alignment for Image-Text and Element Matching Kolors: Effective training of diffusion model for photorealistic text-to-image synthesis

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:26.595625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:26.595625Z digest=sha256:e74ce76c71bc56e363b13bb166fdd8fd7c7bf4f12221ddd30f2c8455c4dda467

Observation d3eac9ab-a837-4762-bb2a-efdca2f4f1c4 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large lan- guage models.

Instruction-augmented Multimodal Alignment for Image-Text and Element Matching Chain-of-thought prompting elicits reasoning in large lan- guage models

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:40:27.177195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:40:26.602565Z digest=sha256:663377100dacc839bc3828b7ea425d3abd109fac955206166380de9e05904bb9

Observation aa4a7cf8-23cb-4208-9fa9-1ba2a7e0dada · outbound

This paper cites Revisiting Text-to-Image Evaluation with Gecko: On Metrics, Prompts, and Human Ratings.

Instruction-augmented Multimodal Alignment for Image-Text and Element Matching Revisiting Text-to-Image Evaluation with Gecko: On Metrics, Prompts, and Human Ratings

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:26.609996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:26.609996Z digest=sha256:ff45f20f976c3b0b0daaacba7e8f47e02929887e048d86db7dbb14c21b76a3f0

Observation 4557b27a-5f78-45e7-9426-e61b90c5d129 · outbound

This paper cites Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models.

Instruction-augmented Multimodal Alignment for Image-Text and Element Matching Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:26.617119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:26.617119Z digest=sha256:b26e9f712280b625f28ed356fb1ed97bc4974d94ebc643f40f3e46e2049c2f99

Observation 40a63b09-6cd1-4e1e-b70f-df790981cee3 · outbound

This paper cites Q-align: Teaching lmms for visual scoring via discrete text-defined levels.

Instruction-augmented Multimodal Alignment for Image-Text and Element Matching Q-align: Teaching lmms for visual scoring via discrete text-defined levels

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:40:27.157242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:40:26.624217Z digest=sha256:3c2798b4a5c79629ad26acb2ddc4434fcd7238c8102e325c7c71c89292a52045

Observation eac17490-2252-4540-9922-a54025055502 · outbound

This paper cites Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis.

Instruction-augmented Multimodal Alignment for Image-Text and Element Matching Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:26.629918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:26.629918Z digest=sha256:8a61269f1e667df7ee8aaec39e233901d777426d49a59fee7a3f544d615e059e

Observation 309f6589-57dd-473e-895c-105382268df6 · outbound

This paper cites Imagere- ward: Learning and evaluating human preferences for text- to-image generation.

Instruction-augmented Multimodal Alignment for Image-Text and Element Matching Imagere- ward: Learning and evaluating human preferences for text- to-image generation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:26.635879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:26.635879Z digest=sha256:d1bffc6d3627193440f7a4877aad0974a8e0cf330e0254b3591fd20dd3f94083

Observation 6eada856-94fd-4d4a-b983-3d8988269077 · outbound

This paper cites What you see is what you read? improving text- image alignment evaluation.

Instruction-augmented Multimodal Alignment for Image-Text and Element Matching What you see is what you read? improving text- image alignment evaluation

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:40:27.123600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:40:26.641902Z digest=sha256:a40d3d7bfa7360ab66804fc71b7242e8c2a6a3edbb29ffb54ad2d0e00c440d0d

Observation 6d3eace7-8ac9-421e-8c84-af0da819edef · outbound

This paper cites mplug- owl3: Towards long image-sequence understanding in multi- modal large language models.

Instruction-augmented Multimodal Alignment for Image-Text and Element Matching mplug- owl3: Towards long image-sequence understanding in multi- modal large language models

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:40:27.102486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:40:26.648908Z digest=sha256:6b13b23eb026a56b260b6bd3a84fe804022ea46aa50b93234df905dc3496db3b

Observation de44943b-669d-4577-a7f3-33a215fc197d · outbound

This paper cites Swift:a scal- able lightweight infrastructure for fine-tuning, 2024.

Instruction-augmented Multimodal Alignment for Image-Text and Element Matching Swift:a scal- able lightweight infrastructure for fine-tuning, 2024

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:26.655485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:26.655485Z digest=sha256:7647813b009fee1dad6ba762a42d49eb73f9c669ffefc6db9d27c2904545df51

Observation 83488eb1-ff86-4677-a0fe-16a21ea732ff · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Instruction-augmented Multimodal Alignment for Image-Text and Element Matching MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:26.661269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:26.661269Z digest=sha256:fb0b835d21ef336df4a0f6eff26aebb28485578e8b8610daf7fe5cce9d73c5b1

Pith citing papers

No inbound Pith citation observations are available.