Pith. sign in

Paper Citation Record · LEDGER

CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder

As of 20 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 4 inbound Pith citation observations for arXiv:2412.17225.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.17225 v1

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T05:45:54.168593Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T05:51:12.656627Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T21:28:58.246299Z

Reference resolution

39 of 39 outbound references displayed

  • verified exact1
  • verified fuzzy21
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 32dcee00-c070-47fe-a50a-c395288bfe3c · outbound

This paper cites Improving Image Gen- eration with Better Captions.

CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder Improving Image Gen- eration with Better Captions

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:45:54.632585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T05:45:54.027737Z digest=sha256:0d6db3cfbdbc132110f45e7f37bb5b295928c78044a4692095c768995ef920d0

Observation d644f7f7-218a-4241-8348-8e52a477c2d7 · outbound

This paper cites Freeman, Michael Rubinstein, Yuanzhen Li, and Dilip Krishnan.

CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder Freeman, Michael Rubinstein, Yuanzhen Li, and Dilip Krishnan

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:45:54.620509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T05:45:54.032151Z digest=sha256:c43e6c59bedba585f0d8b685a296acca874cfcb9d4dbfc49770193bf54b536f2

Observation aa59fca9-478b-42bf-ad77-66a533438e79 · outbound

This paper cites Diffute: Universal text editing diffusion model.

CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder Diffute: Universal text editing diffusion model

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:45:54.610514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T05:45:54.037798Z digest=sha256:1a4f5d29f6212a5099cf20f4a78e577ee7a77eb564a49d61b9f0c6194abfb22f

Observation af6c4dbe-363e-4cd0-a785-3fd33ceaf91e · outbound

This paper cites TextDiffuser-2: Unleashing the Power of Language Models for Text Rendering.

CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder TextDiffuser-2: Unleashing the Power of Language Models for Text Rendering

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T05:45:54.042101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:45:54.042101Z digest=sha256:093bea3e554aaf957d5f3d0b73390c17e86e88e932d2f64ad8609a4de08700ed

Observation 5be04580-7bc3-41bf-a3b5-123cfd3eeea4 · outbound

This paper cites Textdiffuser: Diffusion models as text painters.

CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder Textdiffuser: Diffusion models as text painters

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:45:54.600198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T05:45:54.046269Z digest=sha256:2c6fa802d7456591c3831031100594386f2d53b4a265d3ae668c5bf55b9425ec

Observation 0f6a3fba-73d8-4b6f-9915-8b934c5132e1 · outbound

This paper cites Diffusion Models Beat GANs on Image Synthesis.

CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder Diffusion Models Beat GANs on Image Synthesis

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T05:45:54.049939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:45:54.049939Z digest=sha256:4e6ad502546a07b61de87a8532b52fe0c45d36ffc866340ba098641088465138

Observation a105cabf-43bc-426b-b935-3f12b69ce34c · outbound

This paper cites Odm: A text-image further alignment pre-training approach for scene text detection and spotting.

CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder Odm: A text-image further alignment pre-training approach for scene text detection and spotting

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:45:54.589550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T05:45:54.054305Z digest=sha256:983f65abd885492a5b42d0358d997ee7885f5474761497ec379809d3b17a60e5

Observation a33befb2-80d5-47d5-98fd-700c80dbde74 · outbound

This paper cites Duguangocr.

CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder Duguangocr

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:45:54.578511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T05:45:54.057863Z digest=sha256:7a2ab1e7dd3739180644a539aa5161dfedc1751ad36d3a119f1df64a96a7cc59

Observation 4fdb2fe9-d5c1-46df-a71d-66b3a84b97ba · outbound

This paper cites Scaling recti- fied flow transformers for high-resolution image synthesis.

CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder Scaling recti- fied flow transformers for high-resolution image synthesis

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T05:45:54.061352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:45:54.061352Z digest=sha256:d38d198731a6231b4d1e8e7da9c06690a9a8750e7d2730153e4c57c886216cbf

Observation f6730221-9216-40c4-bccd-9301fa904acb · outbound

This paper cites Wukong: 100 million large-scale chinese cross-modal pre-training dataset and a foundation frame- work, 2022.

CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder Wukong: 100 million large-scale chinese cross-modal pre-training dataset and a foundation frame- work, 2022

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:45:54.558944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T05:45:54.064964Z digest=sha256:5fb1adfd1941fbba8a7acdd0219399dcc737185924c3381472705e0b23d58b06

Observation 78774e74-34d5-4000-bd64-99259f3d4ee9 · outbound

This paper cites Denoising Dif- fusion Probabilistic Models, 2020.

CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder Denoising Dif- fusion Probabilistic Models, 2020

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:45:54.546182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T05:45:54.068410Z digest=sha256:6a5118a1837188affe74f8591d9310731569d1cf3d99b3110c37c5f2cedc62db

Observation 87be2020-c36d-43de-ad5c-6df12c40f426 · outbound

This paper cites Auto-encoding varia- tional bayes, 2022.

CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder Auto-encoding varia- tional bayes, 2022

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:45:54.533173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T05:45:54.072712Z digest=sha256:36b473c3b3e7e76c4ac494c583f6c5e49fd89e0c4e2951d3f871f513707abe4e

Observation 11d9325a-7517-4838-9095-db005ac98a92 · outbound

This paper cites Pp-ocrv3: More attempts for the improvement of ultra lightweight ocr system,.

CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder Pp-ocrv3: More attempts for the improvement of ultra lightweight ocr system,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:45:54.521099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T05:45:54.076401Z digest=sha256:a4cfc658c33a85bb693a5d2f539fa80a2dd1a5830b2791dc61683282840400db

Observation 4a50ce3e-f731-474e-8afe-a05b6511c937 · outbound

This paper cites Character-aware models improve visual text rendering, 2023.

CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder Character-aware models improve visual text rendering, 2023

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:45:54.509116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T05:45:54.079905Z digest=sha256:a05dcb21074952c0f1ff57ec5008d329dfb030c5695cc1a25349d02174881776

Observation ba0c7ce2-e223-4f68-b8c1-413f868882c1 · outbound

This paper cites Character-Aware Models Improve Visual Text Rendering.

CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder Character-Aware Models Improve Visual Text Rendering

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T05:45:54.083336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:45:54.083336Z digest=sha256:26dcd113d7adc0420d787ed13636a4b904dad19c18b05c05cfd0d7fa7c74cee3

Observation 5708ea8d-5278-4cc3-8bc1-46d4e43f1593 · outbound

This paper cites Glyph-ByT5: A Cus- tomized Text Encoder for Accurate Visual Text Rendering,.

CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder Glyph-ByT5: A Cus- tomized Text Encoder for Accurate Visual Text Rendering,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:45:54.495477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T05:45:54.086895Z digest=sha256:c975c6a825951e4018d8da69abcad28dd42af691494c34ac18e7e5f9e0a8b40d

Observation 0c45c0d0-fc72-4ab9-b27b-a285b2cdd64f · outbound

This paper cites Glyph-ByT5-v2: A Strong Aesthetic Baseline for Accurate Multilingual Visual Text Rendering.

CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder Glyph-ByT5-v2: A Strong Aesthetic Baseline for Accurate Multilingual Visual Text Rendering

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T05:45:54.094448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:45:54.094448Z digest=sha256:4333856c032b0ce1c7af2c90baa5e1bad0600b2f625e9b45f41c7ca93ac21d56

Observation 2ecb5dfd-4802-4cc8-b0dc-60425446d0bd · outbound

This paper cites GlyphDraw: Seamlessly Rendering Text with Intricate Spatial Structures in Text-to-Image Generation.

CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder GlyphDraw: Seamlessly Rendering Text with Intricate Spatial Structures in Text-to-Image Generation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T05:45:54.097886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:45:54.097886Z digest=sha256:d1673b67fc691a5e74bbefbf678d97dd8982c87dc9de92ca7e2b2997bd8d6178

Observation ad03fc6a-a382-4c15-9bc4-a7aba992e08e · outbound

This paper cites Glyphdraw2: Automatic generation of complex glyph posters with diffusion models and large language models,.

CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder Glyphdraw2: Automatic generation of complex glyph posters with diffusion models and large language models,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:45:54.482727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T05:45:54.101959Z digest=sha256:7efd1e8fbf9628f3f8478bacc2c65443d5d2951ed5ba6e3a6a46f20d750e6f37

Observation 5e344d80-5f47-487b-9baf-fec339db194f · outbound

This paper cites Midjourney.

CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder Midjourney

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:45:54.468583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T05:45:54.105504Z digest=sha256:8ddb76ab1354391c7195e3276a07b4cde3626b4ee3e411318cc32d51d77f26db

Observation 37c395c1-3718-43a0-91ce-d99a458aecaf · outbound

This paper cites Improved denoising dif- fusion probabilistic models, 2021.

CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder Improved denoising dif- fusion probabilistic models, 2021

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:45:54.455745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T05:45:54.108873Z digest=sha256:67a5ff9f4b5e1b8f05fe7ade4322d5e037f1b300b29b74a5102aa16833d50dd9

Observation fe117036-b7ff-4752-85c2-25ea340eb4ee · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder Learning transferable visual models from natural language supervi- sion

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T05:45:54.112155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:45:54.112155Z digest=sha256:936ed9f164e5c1b390d60cfe15622de78b8ddfd333b71fd7cfe550d5a1c48fa2

Observation 098b4e94-2f61-4228-a1b5-6d8025581694 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T05:45:54.115472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:45:54.115472Z digest=sha256:85a2408854de96d1385d8ff524e2d51d0a45685028bd174d071d5bee4a1079c4

Observation 59b35edd-5e0c-46be-8542-5ed4add82c13 · outbound

This paper cites Zero-shot text-to-image generation, 2021.

CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder Zero-shot text-to-image generation, 2021

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T05:45:54.118609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:45:54.118609Z digest=sha256:341d62d625a3bcdd2c97589b42882369d1962f7b51b09c0b88d0d1c623786861

Observation f74ce017-246f-456e-85b6-1d9d3748c3ed · outbound

This paper cites Hierarchical text-conditional image gener- ation with clip latents, 2022.

CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder Hierarchical text-conditional image gener- ation with clip latents, 2022

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T05:45:54.122161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:45:54.122161Z digest=sha256:b2e4545683a6484052ba918a97844575f523ecf043f0e3ba484ad5ea7d84c608

Observation 8f16315b-4474-4658-a26c-796c1c0801fe · outbound

This paper cites Ocr-vqgan: Taming text- within-image generation.

CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder Ocr-vqgan: Taming text- within-image generation

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:45:54.410844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T05:45:54.126126Z digest=sha256:25da8c24ebf653508f1f816a9249368599d7e9db5a1b1e735897fdd899ad11be

Observation 18160161-c409-440f-88f1-a2b150a5e58c · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder High-resolution image synthesis with latent diffusion models

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:45:54.400593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T05:45:54.129746Z digest=sha256:5213c325180a46ee543d59b35e95f99fba3a57e23a3b7ec29cf92fcb4b2b6b7d

Observation 2b65121d-7aa4-46a7-8d28-d0f51b1b9144 · outbound

This paper cites Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding.

CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T05:45:54.134047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:45:54.134047Z digest=sha256:590446cef45a112e5ad193db146ce82fb58b17f3e02917eaa308c4149d6b24ed

Observation a1522277-adbe-44d6-8b8e-e58ff6b58c12 · outbound

This paper cites LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs.

CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T05:45:54.137766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:45:54.137766Z digest=sha256:98c0234f283ce9c5c2e116b54c8af2899c86cee9080fc1a844efd6dc6989998a

Observation 3dfd3bf8-f8d6-4cdb-a245-7165c583f528 · outbound

This paper cites Kingma, Ab- hishek Kumar, Stefano Ermon, and Ben Poole.

CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder Kingma, Ab- hishek Kumar, Stefano Ermon, and Ben Poole

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T05:45:54.141208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:45:54.141208Z digest=sha256:1f316f94b8895ecbc77dfbff5a226937937293a096323e72d25d803c0cbd34e7

Observation 0d02f40a-f4d3-4215-9539-a03046be485f · outbound

This paper cites AnyText: Multilingual Visual Text Generation And Editing.

CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder AnyText: Multilingual Visual Text Generation And Editing

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T05:45:54.144468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:45:54.144468Z digest=sha256:fbf21255030aec23e4cb675bfae62eb2f5447c726d2c2e09465a4560d8d5f576

Observation 5c17193d-3251-4244-af5a-c4db9ddaa9a3 · outbound

This paper cites Glyphcontrol: Glyph conditional control for visual text generation.

CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder Glyphcontrol: Glyph conditional control for visual text generation

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:45:54.379824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T05:45:54.147962Z digest=sha256:9dc90d3402cd10071e730f416dd96254aa62d7a12499586e8bb76f373ab963f9

Observation 2d8c04b4-8d3e-46c3-a1c3-eee0ee9c063d · outbound

This paper cites Long-clip: Unlocking the long-text capability of clip, 2024.

CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder Long-clip: Unlocking the long-text capability of clip, 2024

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:45:54.367440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T05:45:54.151114Z digest=sha256:818f34d79d444c38f10ca22fabc75fbe1bde018fda5f0541c1de48770513323d

Observation f03aad97-a15e-49a2-ad58-e27632021fe1 · outbound

This paper cites Adding conditional control to text-to-image diffusion models.

CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder Adding conditional control to text-to-image diffusion models

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:45:54.354679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T05:45:54.154138Z digest=sha256:27d9c894fd2178795a50230fd245156098b28b1258c26cb8b5b27339b09697a3

Observation c02d6877-585f-4d82-ab71-d409ae8d0b2b · outbound

This paper cites Adding Conditional Control to Text-to-Image Diffusion Models,.

CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder Adding Conditional Control to Text-to-Image Diffusion Models,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T05:45:54.157738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:45:54.157738Z digest=sha256:9e3894811581b50907c01f5d770da0ea65ab023bfc79774189eae9c617aa6a4b

Observation 40ef395d-b24b-4269-8b3e-ad783d5a255e · outbound

This paper cites Brush your text: Synthesize any scene text on im- ages via diffusion model.

CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder Brush your text: Synthesize any scene text on im- ages via diffusion model

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:45:54.335900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T05:45:54.165257Z digest=sha256:95a4632551418eff1e242df2efc6141b72fe2f9bf4bd8071b576bf76da737c3f

Observation 13b981e2-4f9b-4bc2-9c69-9a6572dca6a3 · outbound

This paper cites UDiffText: A Unified Framework for High-quality Text Synthesis in Arbitrary Images via Character-aware Diffusion Models.

CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder UDiffText: A Unified Framework for High-quality Text Synthesis in Arbitrary Images via Character-aware Diffusion Models

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-08-11T05:45:54.207012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T05:45:54.168593Z digest=sha256:454b3742cdbb86979a0fafe4f2ab3da4c4def7ab7f056b3dc6e39ac82c8e3ddb

Observation bfd22a16-6f0b-4425-9200-70da3b6ed929 · outbound

This paper cites Adding Conditional Control to Text-to-Image Diffusion Models.

CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder Adding Conditional Control to Text-to-Image Diffusion Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-11T05:45:54.161125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:45:54.161125Z digest=sha256:a78e5c8880702511951e79970b116baea8d430765bcb72d7e8a2ce44a598659c

Observation fa5c5579-92f7-4b53-9ea1-4408977906f0 · outbound

This paper cites Glyph-ByT5: A Customized Text Encoder for Accurate Visual Text Rendering.

CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder Glyph-ByT5: A Customized Text Encoder for Accurate Visual Text Rendering

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-11T05:45:54.090483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:45:54.090483Z digest=sha256:7349cbe0e117ce8e5ed37c1634a4955cb4a992604c2fbffecb433426dbe994f7

Pith citing papers

Observation 800ea31d-774f-4249-835e-68c4aa3970ec · inbound

Holding the FP8 Quality Ceiling at 8-Bit Weights and Activations: INT8 and GGUF Post-Training Quantization of Ideogram 4.0 for Consumer GPUs cites this paper.

Holding the FP8 Quality Ceiling at 8-Bit Weights and Activations: INT8 and GGUF Post-Training Quantization of Ideogram 4.0 for Consumer GPUs CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:17:58.047058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T10:05:40.235610Z digest=sha256:128af26627c26ec6cbc47b9ffb6115203cb778df8b367d97f2f1c6c15bc66e02

Observation 8daa24de-0235-4b42-a24e-0ddd934bb10e · inbound

GMO-E$^2$DIT: Grounded Multi-Operation Editing for E-Commerce Images cites this paper.

GMO-E$^2$DIT: Grounded Multi-Operation Editing for E-Commerce Images CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-02T13:56:58.996534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-02T13:56:43.671622Z digest=sha256:4f7e43763ad1c307fbd3934e607b361000d901a960c06c7bd323e8beb0c6b25a

Observation f9fc143e-5a29-4ec9-8905-b2e272ea7998 · inbound

GMO-E$^2$DIT: Grounded Multi-Operation Editing for E-Commerce Images cites this paper.

GMO-E$^2$DIT: Grounded Multi-Operation Editing for E-Commerce Images CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-03T21:28:58.247930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-03T21:23:41.271521Z digest=sha256:a72ecd098c3cff954827ff522dff2fbcb49079209a8246de393b9f5f124fd18c

Observation df547916-2f60-4f9e-a17d-0164114a8eb5 · inbound

InnoText: A Unified Model for Visual Text Generation and Editing cites this paper.

InnoText: A Unified Model for Visual Text Generation and Editing CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T05:51:12.656627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T05:51:12.656627Z digest=sha256:31d3a15840ea1a069ef56aa76f56e22b5096e280b1e6ed46c8ab8c1052ba35d3