Pith. sign in

Paper Citation Record · LEDGER

Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models

As of 11 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2501.07086.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.07086 v1

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T20:53:27.913006Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact1
  • verified fuzzy20
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5c09b0f4-655d-4712-b47e-ace5d8872c24 · outbound

This paper cites GLIDE: towards photorealistic image generation and editing with text-guided diffusion models,.

Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models GLIDE: towards photorealistic image generation and editing with text-guided diffusion models,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:28.561574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:53:27.669201Z digest=sha256:12a4eedd833f52a30f68782110fff3b00c2d90a6235e59cfbd1709c4825bcb8b

Observation bc7188f2-a1ad-4184-9c6b-93f933c60f96 · outbound

This paper cites Zero-shot text-to-image generation,.

Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models Zero-shot text-to-image generation,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:28.546171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:53:27.676905Z digest=sha256:fc38a88c5da20b573d3581c6662bc4f94e2e84f4dd7925062bae051c32138fbc

Observation ff87c1b7-825f-46b8-9328-962e2d6ade18 · outbound

This paper cites Photorealistic text-to-image diffusion models with deep language understanding,.

Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models Photorealistic text-to-image diffusion models with deep language understanding,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:28.528242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:53:27.706976Z digest=sha256:802bdf6b7fe657d31bda5d5b05b9882a50a40372fb4020fd7a3b49ae3aebaa69

Observation e843a7ca-0d43-46e0-83f7-cdae742cdf4c · outbound

This paper cites The revolution of multimodal large language models: A survey,.

Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models The revolution of multimodal large language models: A survey,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:28.512292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:53:27.757052Z digest=sha256:e7641cb63a537fc1328885910085cd2858e52f164144a95a237438aff311f116

Observation 3b34aee4-f94c-43ab-aeb4-6d37ff67b7a5 · outbound

This paper cites Design guidelines for prompt engineering text- to-image generative models,.

Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models Design guidelines for prompt engineering text- to-image generative models,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:28.496749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:53:27.789964Z digest=sha256:d4d07b12b8e9e633b9db1b7c03562a7b1b6806efaadb10c0cfc76d7b29d10677

Observation 64dee35c-00a8-4223-9320-0abecbadc0b8 · outbound

This paper cites Generative Multimodal Models are In-Context Learners.

Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models Generative Multimodal Models are In-Context Learners

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:27.795138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:53:27.795138Z digest=sha256:bca23f55983d2766593a42396e5cd726b67deaabd44b7f2a3bd263ffcc586fdd

Observation 6c2327d7-e453-4c93-b25e-6cd6ba745609 · outbound

This paper cites Promptcot: Align prompt distribution via adapted chain- of-thought,.

Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models Promptcot: Align prompt distribution via adapted chain- of-thought,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:28.480015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:53:27.804754Z digest=sha256:45d6ea553d102ca28791dfa788a3037123604b765a700bccd72c4d44446febda

Observation d45e8b67-f011-4d74-b167-4f73e178c5f5 · outbound

This paper cites Not all languages are created equal in llms: Improving multilingual capability by cross-lingual-thought prompting,.

Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models Not all languages are created equal in llms: Improving multilingual capability by cross-lingual-thought prompting,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:28.464282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:53:27.816471Z digest=sha256:42f9d7002bf791c86edbe18a99be4b85be8c78726dc237aa19cdf30353dd42f9

Observation 5d79981d-5e90-4e23-8910-71bee16d0996 · outbound

This paper cites Cross-lingual prompt- ing: Improving zero-shot chain-of-thought reasoning across languages,.

Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models Cross-lingual prompt- ing: Improving zero-shot chain-of-thought reasoning across languages,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:28.432734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:53:27.825982Z digest=sha256:93d94bd5216ab70abbe6029fe0e47dc7fa0cfd6f2b9085e46ee1e538c75e805b

Observation da7a3b5b-2483-43b3-a1c7-525a0fb0cc1a · outbound

This paper cites PLUG: leveraging pivot language in cross-lingual instruction tuning,.

Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models PLUG: leveraging pivot language in cross-lingual instruction tuning,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:28.413772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:53:27.830870Z digest=sha256:243d32198ade4c4ae888e160704546a00733b59d46eca2da51aee258e08c0f3b

Observation cf8ad285-c9ec-4bbf-b056-67851978658c · outbound

This paper cites Revealing the Parallel Multilingual Learning within Large Language Models.

Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models Revealing the Parallel Multilingual Learning within Large Language Models

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-10T20:53:28.055231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:53:27.835654Z digest=sha256:c33e131ba1648ec8fae3a76e36a67d47f2b50f15bae093599b7ad24972822b02

Observation ae032c7f-a368-466f-bcff-c2fdbf153b55 · outbound

This paper cites LAION-5B: an open large-scale dataset for training next generation image-text models,.

Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models LAION-5B: an open large-scale dataset for training next generation image-text models,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:28.397020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:53:27.840560Z digest=sha256:ea14a047f6a537366af3166c01dd75a1a5e27f1429b6aa94530bf48730327229

Observation d6bf1b94-9ff5-47d2-805e-a1f6d0ee3f85 · outbound

This paper cites Denoising diffusion probabilistic models,.

Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models Denoising diffusion probabilistic models,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:28.378402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:53:27.845353Z digest=sha256:1ce5aaa553710a8008921eea8dad58133394f91d90a60e1d8013cb0425253e81

Observation 245c3e8f-6ab6-461c-88c1-83c2602f244a · outbound

This paper cites Taming transformers for high- resolution image synthesis,.

Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models Taming transformers for high- resolution image synthesis,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:28.362226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:53:27.850132Z digest=sha256:ce2dcbff55d07576a72f98456aebd8d427227aa5459194d5df5064c296c9fdae

Observation b85d43a4-c851-4411-bc88-d6b2da11086b · outbound

This paper cites Clipscore: A reference-free evaluation metric for image captioning,.

Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models Clipscore: A reference-free evaluation metric for image captioning,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:28.344341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:53:27.855077Z digest=sha256:8079229688470b82663e268c47d6cdd174694eae8525f633d7dfe53547813bbc

Observation eeed5c60-77dd-40a2-ac61-3415df90a96f · outbound

This paper cites Microsoft COCO: common objects in context,.

Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models Microsoft COCO: common objects in context,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:27.859826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:53:27.859826Z digest=sha256:d7b59f7cf292aff9f0f788bd72e1ac9822957997a9c7bc7683fcd9ca9fd83284

Observation 8e30d65f-3b00-49c4-bca0-7f73ba2739b0 · outbound

This paper cites Improving image generation with better captions,.

Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models Improving image generation with better captions,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:27.864672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:53:27.864672Z digest=sha256:0d7b6e4ad70b860517ba2f4a60fde9be1fbb9b6ccb063b233564a87c716728e1

Observation 2cc3403d-8551-4c7a-b4ff-b67900ccc57b · outbound

This paper cites T2i-compbench: A comprehensive benchmark for open-world compositional text-to-image generation,.

Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models T2i-compbench: A comprehensive benchmark for open-world compositional text-to-image generation,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:28.302428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:53:27.874475Z digest=sha256:a74a8adc533adabacbd09071eeab4fab40c7c5a4a1ad32b97695b11adc0b13c3

Observation aeb945c6-ee58-432a-b11f-023d886ea8e1 · outbound

This paper cites Magicbrush: A manually annotated dataset for instruction-guided image editing,.

Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models Magicbrush: A manually annotated dataset for instruction-guided image editing,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:28.281489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:53:27.879790Z digest=sha256:838c3946bf9eea747a07b3757ff7ddfe09df3d84ea6d134b1d2978ad933800ee

Observation 76e9cbcf-bcb4-4703-b250-ecfdefdfaa1c · outbound

This paper cites BLIP: bootstrapping language- image pre-training for unified vision-language understanding and genera- tion,.

Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models BLIP: bootstrapping language- image pre-training for unified vision-language understanding and genera- tion,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:28.262795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:53:27.884483Z digest=sha256:66280afe42e2585adc7bc286e47107a22a4531ef20b7193022f405af9fbf5be6

Observation 49515561-7c1d-4a3c-9bde-93083d2c8c80 · outbound

This paper cites Imagereward: Learning and evaluating human preferences for text- to-image generation,.

Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models Imagereward: Learning and evaluating human preferences for text- to-image generation,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:28.235684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:53:27.889016Z digest=sha256:211a75beff539cff0697ab2fffb970ed3aa15c9681a59a208c5d45ab3239e015

Observation d4fd8d51-e318-42f3-adb0-a86d990013b0 · outbound

This paper cites Lumina-Next: Making Lumina-T2X Stronger and Faster with Next-DiT.

Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models Lumina-Next: Making Lumina-T2X Stronger and Faster with Next-DiT

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:27.893502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:53:27.893502Z digest=sha256:3e736ed59b77ce9e593756aa507cf9db1e41b91d490980ef890c9ae79254c1ac

Observation d40e715a-7169-44ab-b756-e860f0a55c8e · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models Gemma: Open Models Based on Gemini Research and Technology

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:27.898679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:53:27.898679Z digest=sha256:7f49dd309d963b04a8f7d202ed734b71716096c15f04df8d8fbfbfb27a92f77b

Observation aa3b67f8-2e80-4e97-ac35-e0e4c0e4dc06 · outbound

This paper cites Optimizing prompts for text- to-image generation,.

Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models Optimizing prompts for text- to-image generation,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:28.183564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:53:27.903719Z digest=sha256:e6d9edff6c7f77b5972e28b5cb13f22f31ea1bfdd7b334300d6fbbdf87acc101

Observation f22d0cdc-9799-45d9-af56-b3e61064f226 · outbound

This paper cites Aligner: Efficient Alignment by Learning to Correct.

Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models Aligner: Efficient Alignment by Learning to Correct

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:27.908141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:53:27.908141Z digest=sha256:e0089a060a997f852107c630e6db48a6026dd24d6727657221c1540503aaead3

Observation 17196ef3-5d1f-4214-b4b3-cdf3c1cb54e7 · outbound

This paper cites Weak-to-strong generalization: Eliciting strong capabilities with weak supervision,.

Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models Weak-to-strong generalization: Eliciting strong capabilities with weak supervision,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:28.114770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:53:27.913006Z digest=sha256:6c43f8baba4a3f93fc8de268562e7313a9c5f057e7556fc69654a3abf07424a1

Observation 1a4156a8-c63e-4c90-ba54-adc4cab0af09 · outbound

This paper cites 12 365– 12 394.

Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models 12 365– 12 394

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:28.449198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:53:27.821191Z digest=sha256:8803c5910301201d7baf0fc7dc97d3a8d22ceaf9270ab5661107d508f1a7d993

Pith citing papers

No inbound Pith citation observations are available.