Pith. sign in

Paper Citation Record · LEDGER

StyleTextGen: Style-Conditioned Multilingual Scene Text Generation

As of 22 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 1 inbound Pith citation observation for arXiv:2605.14708.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.14708 v1

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-15T04:54:52.729257Z

measured 62 of 62 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T07:57:26.635757Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

61 of 61 outbound references displayed

  • verified exact17
  • verified fuzzy43
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 24f6cd5f-7c37-47b7-9b9c-b53d144ffb50 · outbound

This paper cites In- structpix2pix: Learning to follow image editing instructions.

StyleTextGen: Style-Conditioned Multilingual Scene Text Generation In- structpix2pix: Learning to follow image editing instructions

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T04:55:04.241800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T04:54:52.729257Z digest=sha256:22da1aa7989c194090d3de04fa607992fef9d42dbc5960930921f895aa98a70b

Observation 60d920da-3ac4-4aa7-9f69-33cbecf20565 · outbound

This paper cites The Devil is in Fine-tuning and Long-tailed Problems:A New Benchmark for Scene Text Detection.

StyleTextGen: Style-Conditioned Multilingual Scene Text Generation The Devil is in Fine-tuning and Long-tailed Problems:A New Benchmark for Scene Text Detection

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-15T04:55:03.039023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T04:54:52.729257Z digest=sha256:a421267e521dd3c7afc3a53be947b6f86a455cbef88e97bde2e5f6e196157bb2

Observation c5b95678-6ddf-42f5-ba33-1d092b835608 · outbound

This paper cites Posta: A go-to framework for customized artistic poster gen- eration.

StyleTextGen: Style-Conditioned Multilingual Scene Text Generation Posta: A go-to framework for customized artistic poster gen- eration

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T04:55:04.252601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T04:54:52.729257Z digest=sha256:a1d498b5668b9d576d16e300b5b354fb4add5f3f9b1fa8b0bdefb69d1f186e22

Observation 7fc71fe7-3ca7-4aa9-84c1-6cc260d9ea84 · outbound

This paper cites TextDiffuser-2: Unleashing the Power of Language Models for Text Rendering.

StyleTextGen: Style-Conditioned Multilingual Scene Text Generation TextDiffuser-2: Unleashing the Power of Language Models for Text Rendering

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-15T04:55:03.030702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T04:54:52.729257Z digest=sha256:92ac4596f21970887a89909e85fdf7c3c3ccd7dad27257f66999fbba2b02b8cc

Observation caa5316c-fe76-4e77-bf7f-122c122bc808 · outbound

This paper cites TextDiffuser: Diffusion Models as Text Painters.

StyleTextGen: Style-Conditioned Multilingual Scene Text Generation TextDiffuser: Diffusion Models as Text Painters

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-15T04:55:03.023699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T04:54:52.729257Z digest=sha256:03813749c2956228e01e6134448de7e25cc9cb6f1587778a1c35ca93d968e622

Observation 6c96dadd-75a6-4ddf-aaed-ae4b57bbe04f · outbound

This paper cites Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks.

StyleTextGen: Style-Conditioned Multilingual Scene Text Generation Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T04:55:04.236672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T04:54:52.729257Z digest=sha256:a676ee2d89d24ba0c90c44f9d353b052f69807f99fc21ebe8866a868f4775589

Observation 902ca4c2-9e11-423e-afab-27f34339040c · outbound

This paper cites Context per- ception parallel decoder for scene text recognition.IEEE Trans.

StyleTextGen: Style-Conditioned Multilingual Scene Text Generation Context per- ception parallel decoder for scene text recognition.IEEE Trans

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T04:55:04.231255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T04:54:52.729257Z digest=sha256:44682b3467013488a910a0159dad6b5577c4917be2098a296598a230dcbc92bc

Observation 4495f058-4850-4ca9-bbd5-b86a6b55c5bb · outbound

This paper cites Instruction-guided scene text recognition.IEEE Trans.

StyleTextGen: Style-Conditioned Multilingual Scene Text Generation Instruction-guided scene text recognition.IEEE Trans

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T04:55:04.247565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T04:54:52.729257Z digest=sha256:430bf1498fea5b2d7148da3928b1c4936533d824dd9568a17c17cbdf4ab3d5ec

Observation f6c43d30-243e-4612-9f66-e719d814ee67 · outbound

This paper cites Recognition-synergistic scene text editing.

StyleTextGen: Style-Conditioned Multilingual Scene Text Generation Recognition-synergistic scene text editing

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T07:20:12.908645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T04:54:52.729257Z digest=sha256:1c2dd51d80e579bece1ec5022c4f7e431449a00ef652a1d40d1c258d48ae933b

Observation 7836c6b6-8dbf-4860-a29e-ec44c115caf7 · outbound

This paper cites Im- age style transfer using convolutional neural networks.

StyleTextGen: Style-Conditioned Multilingual Scene Text Generation Im- age style transfer using convolutional neural networks

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T07:20:12.852573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T04:54:52.729257Z digest=sha256:33bbdf5817318de2cd848e4135ae7dac21e34d11527074bbd2c86ed1497d95b3

Observation 3a06ee62-7ad7-44a4-b8b4-d0b1bf4a9e40 · outbound

This paper cites A Token-level Text Image Foundation Model for Document Understanding.

StyleTextGen: Style-Conditioned Multilingual Scene Text Generation A Token-level Text Image Foundation Model for Document Understanding

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-15T04:55:03.129477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T04:54:52.729257Z digest=sha256:0a275fa40900a1875c4190e0d34048555bc1285170bb1c1093061773b3ed23f4

Observation b966bc17-091a-4388-b39a-6e999f1c6252 · outbound

This paper cites Arbitrary style transfer in real-time with adaptive instance normalization.

StyleTextGen: Style-Conditioned Multilingual Scene Text Generation Arbitrary style transfer in real-time with adaptive instance normalization

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T07:20:12.899534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T04:54:52.729257Z digest=sha256:fc6f002fb50d1e124923bde122c0dbfc959991501cc336ed0b567d66ee9acb2e

Observation 626bf490-34ba-4912-b3b6-9850b353cad5 · outbound

This paper cites Improving diffusion models for scene text editing with dual encoders.

StyleTextGen: Style-Conditioned Multilingual Scene Text Generation Improving diffusion models for scene text editing with dual encoders

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T04:55:04.257564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T04:54:52.729257Z digest=sha256:fb86e0c90f2c556bfeb4b4bac46d958c082b87d894507067045e5df08b10e647

Observation 6b84a317-e4ba-47ee-ab65-e1ba3c68fc8d · outbound

This paper cites Perceptual losses for real-time style transfer and super-resolution.

StyleTextGen: Style-Conditioned Multilingual Scene Text Generation Perceptual losses for real-time style transfer and super-resolution

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T07:20:12.894610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T04:54:52.729257Z digest=sha256:24bb370e9a42057496187853e794ca644990915074dceac782b87cf4fdd6c7a6

Observation 430c5956-0bea-4121-a516-8f13a58eab1d · outbound

This paper cites Brushnet: A plug-and-play image inpainting model with decomposed dual-branch diffusion.

StyleTextGen: Style-Conditioned Multilingual Scene Text Generation Brushnet: A plug-and-play image inpainting model with decomposed dual-branch diffusion

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T04:55:04.268962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T04:54:52.729257Z digest=sha256:6f126bdecd4b117f902ad74acaa32620a92fc1e5875e2b30cb1466627f905ed8

Observation e85e3cba-5395-4590-a7bd-7ca14e1815fb · outbound

This paper cites A style-based generator architecture for generative adversarial networks.

StyleTextGen: Style-Conditioned Multilingual Scene Text Generation A style-based generator architecture for generative adversarial networks

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T04:55:04.263430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T04:54:52.729257Z digest=sha256:bf24805d68039b33b960087f529e9327bfab552d43d2402c13cbc193f2590183

Observation 94293e39-064e-493c-ada6-754504b2350c · outbound

This paper cites Analyzing and improv- ing the image quality of stylegan.

StyleTextGen: Style-Conditioned Multilingual Scene Text Generation Analyzing and improv- ing the image quality of stylegan

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T07:20:12.904178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T04:54:52.729257Z digest=sha256:d250f0f5ff696d10908a613a7e7a389c3e4c2e6915ae11195067a4b199b142b1

Observation bb89238a-b708-460e-80da-5483f466d16e · outbound

This paper cites Textstylebrush: transfer of text aes- thetics from a single example.IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(7):9122–9134.

StyleTextGen: Style-Conditioned Multilingual Scene Text Generation Textstylebrush: transfer of text aes- thetics from a single example.IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(7):9122–9134

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T07:20:12.948917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T04:54:52.729257Z digest=sha256:dd9869f085211bf580badf6d65a9f159f79788fe829cf0f333012ed372738842

Observation 2aa1d89c-70a4-4d74-9641-47d488e542f3 · outbound

This paper cites Clipstyler: Image style transfer with a single text condition.

StyleTextGen: Style-Conditioned Multilingual Scene Text Generation Clipstyler: Image style transfer with a single text condition

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T04:55:04.273860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T04:54:52.729257Z digest=sha256:bbadc8add98ce2da1118158adbed68456c249524206313c1ff734c6a208ed6a6

Observation 91bffa3e-894f-47ab-be99-c0d089e70064 · outbound

This paper cites Flux.https://github.com/ black-forest-labs/flux.

StyleTextGen: Style-Conditioned Multilingual Scene Text Generation Flux.https://github.com/ black-forest-labs/flux

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T04:55:04.284998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T04:54:52.729257Z digest=sha256:1a6e1ba263093b99e42515bed4ad89d265b322842d409193961ac7caf769dc80

Observation b07b3d65-a598-4afd-8651-f0379c47d390 · outbound

This paper cites Stylestudio: Text-driven style transfer with selective control of style elements.

StyleTextGen: Style-Conditioned Multilingual Scene Text Generation Stylestudio: Text-driven style transfer with selective control of style elements

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T07:20:12.887919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T04:54:52.729257Z digest=sha256:d030eddb8ec30f6e5d0a6981747b2a89644c1159d06bba3ceeff328b16ae4621

Observation 7db49b0b-05ee-4345-bb82-7178b900b9ec · outbound

This paper cites PP-OCRv3: More Attempts for the Improvement of Ultra Lightweight OCR System.

StyleTextGen: Style-Conditioned Multilingual Scene Text Generation PP-OCRv3: More Attempts for the Improvement of Ultra Lightweight OCR System

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-15T04:55:03.123084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T04:54:52.729257Z digest=sha256:8f024c50d45212188586d4887e1908932e6b7e36f2c74cc87e93b3d7bdb38d17

Observation c671ba7f-0827-4476-acb7-10d6bfc44c0a · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

StyleTextGen: Style-Conditioned Multilingual Scene Text Generation BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-15T04:55:03.054185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T04:54:52.729257Z digest=sha256:7982e6118bb6141643944c383996b02cde103e066919a5b5540d34063c905d27

Observation aebb15c8-c436-4541-aa27-3842e4ad3a1b · outbound

This paper cites First Creating Backgrounds Then Rendering Texts: A New Paradigm for Visual Text Blending.

StyleTextGen: Style-Conditioned Multilingual Scene Text Generation First Creating Backgrounds Then Rendering Texts: A New Paradigm for Visual Text Blending

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-15T04:55:03.136540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T04:54:52.729257Z digest=sha256:c791e8f5ec018843edd64ac4a21b562cb94c8e05903f8259bcedcfe4781cfeb8

Observation f75e0ea4-693f-4f07-9d7d-1379600444a1 · outbound

This paper cites an unresolved cited work.

StyleTextGen: Style-Conditioned Multilingual Scene Text Generation Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-05-15T04:55:04.289742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T04:54:52.729257Z digest=sha256:4ab9ef74210a89e0b1309fc5cdcbee2045dad549621e63e4ff53b5e494272723

Observation f55839a9-78de-4a8f-bace-6c146e728901 · outbound

This paper cites Glyph-ByT5: A Customized Text Encoder for Accurate Visual Text Rendering.

StyleTextGen: Style-Conditioned Multilingual Scene Text Generation Glyph-ByT5: A Customized Text Encoder for Accurate Visual Text Rendering

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-15T04:55:03.113705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T04:54:52.729257Z digest=sha256:19a3cf82fcd0fa8bd8d13261bc534bf6ecfcb7cd103aa1ffb3784f7ceb8a9863

Observation 1313bef7-0887-4496-bbd8-82504feb98a1 · outbound

This paper cites Decoupled weight decay regularization.

StyleTextGen: Style-Conditioned Multilingual Scene Text Generation Decoupled weight decay regularization

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T07:20:12.862509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T04:54:52.729257Z digest=sha256:91db6ef89d934af0a049b55cf7c4f05e03a05f6486f919614fea7e50cd52ec6a

Observation b4c5beea-9355-4e28-a22b-751207d41fe8 · outbound

This paper cites Arbitrary reading order scene text spotter with local semantics guidance.

StyleTextGen: Style-Conditioned Multilingual Scene Text Generation Arbitrary reading order scene text spotter with local semantics guidance

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T04:55:04.279901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T04:54:52.729257Z digest=sha256:3d170c8a8b585d75c9b2958e9cfbb8b959f6a5c13914afef67dd31df459326d3

Observation 8b7e3217-047a-4b1a-bfdb-8a739093ae27 · outbound

This paper cites Glyphdraw2: Automatic generation of complex glyph posters with diffusion models and large language models.

StyleTextGen: Style-Conditioned Multilingual Scene Text Generation Glyphdraw2: Automatic generation of complex glyph posters with diffusion models and large language models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T07:20:12.873182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T04:54:52.729257Z digest=sha256:24a8dfd6e3a5a9fd9943a2d70d80c80970215f6e5ebb38ae77efb400df3fff18

Observation ef27a535-1264-49cd-9e1b-458f20a6f053 · outbound

This paper cites Calligrapher: Freestyle Text Image Customization.

StyleTextGen: Style-Conditioned Multilingual Scene Text Generation Calligrapher: Freestyle Text Image Customization

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-15T04:55:03.061571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T04:54:52.729257Z digest=sha256:9fe1efd8cdd374ff1192d7db764d1aa68e4032122f0de7862d21dea411a61544

Observation e6c1a51e-59f4-4da4-917b-bc9918444490 · outbound

This paper cites Dall·e3.https://openai.com/index/ dall-e-3/.

StyleTextGen: Style-Conditioned Multilingual Scene Text Generation Dall·e3.https://openai.com/index/ dall-e-3/

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T07:20:12.857502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T04:54:52.729257Z digest=sha256:7fb0c5bdfcc2ba61cbaadac18bc912ad5aa46daf5a994c251a31bcc8e2cd2202

Observation 1484b868-44d6-4e82-987b-ceaa28739b7a · outbound

This paper cites SDXL: improving latent diffusion models for high-resolution image synthesis.

StyleTextGen: Style-Conditioned Multilingual Scene Text Generation SDXL: improving latent diffusion models for high-resolution image synthesis

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T07:20:12.883191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T04:54:52.729257Z digest=sha256:be8d26b7bf962db2fd29321992bc2e552b56b5d095be320c4ed13fe8ea29f9c1

Observation c97e3310-7959-42d3-a712-b1f7bb44cc42 · outbound

This paper cites Exploring stroke-level mod- ifications for scene text editing.

StyleTextGen: Style-Conditioned Multilingual Scene Text Generation Exploring stroke-level mod- ifications for scene text editing

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T07:20:12.867547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T04:54:52.729257Z digest=sha256:f69c34c7b2fe55723883c2c8e9924768147fd0b10e92424e5f09e0bace34fcd7

Observation 93d38291-dcaf-4a7a-8345-8faa7ce923ae · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

StyleTextGen: Style-Conditioned Multilingual Scene Text Generation Learning transferable visual models from natural language supervi- sion

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T07:20:12.846566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T04:54:52.729257Z digest=sha256:8dc532697e3dfb5990c0493b01be53718653a13fd826cbb54fe1d3fda541272b

Observation 51a25496-5d3d-4c3a-a4ac-e6e3ffeaf186 · outbound

This paper cites Zero-shot text-to-image generation.

StyleTextGen: Style-Conditioned Multilingual Scene Text Generation Zero-shot text-to-image generation

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T04:55:04.362173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T04:54:52.729257Z digest=sha256:3528408dcb9755226dbbac602bf3946c48886b85ddd578e030b011ee007be525

Observation ce0c8142-189b-4f20-af77-995a7ea217b1 · outbound

This paper cites High-resolution image syn- thesis with latent diffusion models.

StyleTextGen: Style-Conditioned Multilingual Scene Text Generation High-resolution image syn- thesis with latent diffusion models

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T04:55:04.366525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T04:54:52.729257Z digest=sha256:72230febb191fe6f5d26e89912c479c6f704a35e97513b4a3e804a7b5e8bd369

Observation 5c8fe24e-7e37-418b-9f71-5c50edacb4fa · outbound

This paper cites Semantic Image Inversion and Editing using Rectified Stochastic Differential Equations.

StyleTextGen: Style-Conditioned Multilingual Scene Text Generation Semantic Image Inversion and Editing using Rectified Stochastic Differential Equations

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-15T04:55:03.075170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T04:54:52.729257Z digest=sha256:571a29b394e8e5a597a7b19489c584ade0f5dc36c26a38992c3c524db852c799

Observation 75818bf6-e02d-4129-ab7d-955819eb9968 · outbound

This paper cites Stefann: scene text editor using font adaptive neural network.

StyleTextGen: Style-Conditioned Multilingual Scene Text Generation Stefann: scene text editor using font adaptive neural network

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T04:55:04.357753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T04:54:52.729257Z digest=sha256:a2e4fea743a4fcf8fc0863d6f22fe9049910f7e91bbcbfdfd9dc421b75e45c2f

Observation 29259fcd-a411-4a31-90cb-799c3e89b229 · outbound

This paper cites Denton, Seyed Kamyar Seyed Ghasemipour, Raphael Gontijo Lopes, Burcu Karagol Ayan, Tim Salimans, Jonathan Ho, David J.

StyleTextGen: Style-Conditioned Multilingual Scene Text Generation Denton, Seyed Kamyar Seyed Ghasemipour, Raphael Gontijo Lopes, Burcu Karagol Ayan, Tim Salimans, Jonathan Ho, David J

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T04:55:04.348122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T04:54:52.729257Z digest=sha256:caeefcc493f444ef1212d2bd625c0e9b1bfec750315ec8310d7a7fa3d5535e8f

Observation 78d21797-800a-4469-94e0-84037d7696e9 · outbound

This paper cites When semantics mislead vision: Mitigating large mul- timodal models hallucinations in scene text spotting and un- derstanding.

StyleTextGen: Style-Conditioned Multilingual Scene Text Generation When semantics mislead vision: Mitigating large mul- timodal models hallucinations in scene text spotting and un- derstanding

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-15T04:55:03.045948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T04:54:52.729257Z digest=sha256:b89ce2c0063d97b0f5a341e821f46ac177558e54166a787a5ec04d0517a86c61

Observation 93729c8d-bf0a-4826-8b46-49988fb29bdd · outbound

This paper cites Visual Text Processing: A Comprehensive Review and Unified Evaluation.

StyleTextGen: Style-Conditioned Multilingual Scene Text Generation Visual Text Processing: A Comprehensive Review and Unified Evaluation

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-15T04:55:03.094426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T04:54:52.729257Z digest=sha256:8555e5a2761d980078d080efb3df6e0d5eb09290694f9ff3f089b24244f7371e

Observation 0957c762-b8ee-4911-9656-f2377559ce32 · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

StyleTextGen: Style-Conditioned Multilingual Scene Text Generation SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-05-15T04:55:03.102437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T04:54:52.729257Z digest=sha256:577bbbfcfbf93da40a2786298bce598843b228dd1c7b6ba01bf0ff4a1559f3b5

Observation f1b26d75-779d-43cd-b31d-7fd72d309db3 · outbound

This paper cites Anytext: Multilingual visual text gener- ation and editing.arXiv.

StyleTextGen: Style-Conditioned Multilingual Scene Text Generation Anytext: Multilingual visual text gener- ation and editing.arXiv

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T04:55:04.343764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T04:54:52.729257Z digest=sha256:f966cabc41b89f72d4e6ce10bb785217a04746d011964171c585f089002257e8

Observation 300d6b35-b263-4480-b82a-c444e2d1f2b6 · outbound

This paper cites Anytext2: Vi- sual text generation and editing with customizable attributes.

StyleTextGen: Style-Conditioned Multilingual Scene Text Generation Anytext2: Vi- sual text generation and editing with customizable attributes

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T04:55:04.352651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T04:54:52.729257Z digest=sha256:7400208a2dae4a83a6cd83d0a26fc5827d3b94193a1ca74c628bc455421cf811

Observation d974ff6a-40aa-4cf8-99fd-5c1d75e2ee38 · outbound

This paper cites Texture Networks: Feed-forward Synthesis of Textures and Stylized Images.

StyleTextGen: Style-Conditioned Multilingual Scene Text Generation Texture Networks: Feed-forward Synthesis of Textures and Stylized Images

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-05-15T04:55:03.081180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T04:54:52.729257Z digest=sha256:72de7d129d5d21fc373a72859be34b353071f0641f70e5fda94e635de689d039

Observation 0fd9fdae-9e72-475d-874b-a5b19a406da7 · outbound

This paper cites Rectified diffusion: Straightness is not your need in rectified flow.

StyleTextGen: Style-Conditioned Multilingual Scene Text Generation Rectified diffusion: Straightness is not your need in rectified flow

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T07:20:12.943190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T04:54:52.729257Z digest=sha256:71969c37e4a8c69d3976181802866583ba0dcd14b3808ea5475fa718ebb61c10

Observation 2912d3be-6bd9-4df4-9717-dc30a8d8ba65 · outbound

This paper cites Glyphmastero: A glyph encoder for high-fidelity scene text editing.

StyleTextGen: Style-Conditioned Multilingual Scene Text Generation Glyphmastero: A glyph encoder for high-fidelity scene text editing

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T04:55:04.333921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T04:54:52.729257Z digest=sha256:e624e04720a6785fc2b6db79853f2ff799ddf7e973d674de6d8f19dfc6f47cbe

Observation 4f6efb53-3845-466f-8d3b-692671bcd8de · outbound

This paper cites Editing text in the wild.

StyleTextGen: Style-Conditioned Multilingual Scene Text Generation Editing text in the wild

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T04:55:04.323337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T04:54:52.729257Z digest=sha256:f0d22af1f21523a562ad27adf9fe655211a4524ed961a1e7ac809c24c36819a2

Observation d1f1c944-eb21-48db-9ec8-cb092acebc66 · outbound

This paper cites Textflux: An ocr-free dit model for high-fidelity multilingual scene text synthesis.arXiv preprint arXiv:2505.17778.

StyleTextGen: Style-Conditioned Multilingual Scene Text Generation Textflux: An ocr-free dit model for high-fidelity multilingual scene text synthesis.arXiv preprint arXiv:2505.17778

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-15T04:55:03.143340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T04:54:52.729257Z digest=sha256:6c4bd623e0a3761455bba4ce890285b691e5cd6c7c9cdec7222e96eba176fb22

Observation c08c2294-5336-4bb9-8889-a38a1b573721 · outbound

This paper cites Swaptext: Image based texts transfer in scenes.

StyleTextGen: Style-Conditioned Multilingual Scene Text Generation Swaptext: Image based texts transfer in scenes

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T04:55:04.328606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T04:54:52.729257Z digest=sha256:13d1a88bca97f01f1a43fca90a53c1a9b7ffaa4fd8fe8569a9b93fc6a240b69b

Observation 2554edca-9fa7-43c1-b0cc-2158b819da69 · outbound

This paper cites Ipad: iterative, parallel, and diffusion-based network for scene text recog- nition.International Journal of Computer Vision, 133(8): 5589–5609.

StyleTextGen: Style-Conditioned Multilingual Scene Text Generation Ipad: iterative, parallel, and diffusion-based network for scene text recog- nition.International Journal of Computer Vision, 133(8): 5589–5609

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T04:55:04.339039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T04:54:52.729257Z digest=sha256:42cacd23f58b8d299733d9b7777c20714cb8a895055dc49c9b011850856c79c7

Observation 121a68de-a528-4e89-89f5-29744d22b450 · outbound

This paper cites GlyphControl: Glyph Conditional Control for Visual Text Generation.

StyleTextGen: Style-Conditioned Multilingual Scene Text Generation GlyphControl: Glyph Conditional Control for Visual Text Generation

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-15T04:55:03.087731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T04:54:52.729257Z digest=sha256:fada6304796b5a93dd1ce869d7eeeecfe195d445da88a4b4aa2f7d3af202a2f4

Observation e8f17d2c-d57b-4922-8662-abc9ace99d1d · outbound

This paper cites arXiv preprint arXiv:2505.22810 , year=.

StyleTextGen: Style-Conditioned Multilingual Scene Text Generation arXiv preprint arXiv:2505.22810 , year=

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-15T04:55:03.068402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T04:54:52.729257Z digest=sha256:636523f71a5473ce01456d3b5ee0bed902849604e80800510b07ad4bb2670893

Observation cfe462b8-b7b8-4990-9ee4-f27c391810af · outbound

This paper cites Ip- adapter: Text compatible image prompt adapter for text-to- image diffusion models.

StyleTextGen: Style-Conditioned Multilingual Scene Text Generation Ip- adapter: Text compatible image prompt adapter for text-to- image diffusion models

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T07:20:12.936906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T04:54:52.729257Z digest=sha256:49e842e9b5604c0f0e3cc8f4b4082223ddab72d2ad6c670bea19eff4b732a6bd

Observation 7b7f2c66-4699-4540-b582-ee87ceb0a248 · outbound

This paper cites Hi-sam: Marrying segment anything model for hierarchical text segmentation.

StyleTextGen: Style-Conditioned Multilingual Scene Text Generation Hi-sam: Marrying segment anything model for hierarchical text segmentation

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T07:20:12.878553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T04:54:52.729257Z digest=sha256:7e975371189fced63e6ba0b28afac37711846d162f902ff22bf08822cff54666

Observation 87570db8-7e2f-42b8-b61a-9e9e0312c53f · outbound

This paper cites Textctrl: Diffusion-based scene text editing with prior guidance control.Advances in Neural Information Pro- cessing Systems, 37:138569–138594.

StyleTextGen: Style-Conditioned Multilingual Scene Text Generation Textctrl: Diffusion-based scene text editing with prior guidance control.Advances in Neural Information Pro- cessing Systems, 37:138569–138594

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T04:55:04.304966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T04:54:52.729257Z digest=sha256:6c4b59e4df2a60b139c4cf63966a13348e48d982d150b6cf324098fd519f6227

Observation e7495cd9-e347-423b-9e39-f8e65ea7fbfb · outbound

This paper cites Inversion-based style transfer with diffusion models.

StyleTextGen: Style-Conditioned Multilingual Scene Text Generation Inversion-based style transfer with diffusion models

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T04:55:04.310106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T04:54:52.729257Z digest=sha256:022899307aa80a6473c8346e85e822416fa3448134f8f0947fbb2d129f550a7a

Observation e4ae7549-31c1-4e7b-a760-2329c5722787 · outbound

This paper cites Metaxas, and Praveen Krishnan.

StyleTextGen: Style-Conditioned Multilingual Scene Text Generation Metaxas, and Praveen Krishnan

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T04:55:04.314497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T04:54:52.729257Z digest=sha256:ffc0b640727c45352137724864a3e9414ec5fb8e148bfd5d021062c4954641e5

Observation 8e81627d-c3bd-491c-80cf-8240363d5df2 · outbound

This paper cites Udifftext: A unified frame- work for high-quality text synthesis in arbitrary images via character-aware diffusion models.

StyleTextGen: Style-Conditioned Multilingual Scene Text Generation Udifftext: A unified frame- work for high-quality text synthesis in arbitrary images via character-aware diffusion models

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T04:55:04.294612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T04:54:52.729257Z digest=sha256:e6254c9540496b73a97ab6eba87e874fe1d57235f1a7597a63f232117df7549a

Observation c752ae6c-d36c-49fd-ad23-ddda37fc2d31 · outbound

This paper cites Cdistnet: Perceiving multi-domain character distance for robust text recognition.IJCV, 132(2): 300–318.

StyleTextGen: Style-Conditioned Multilingual Scene Text Generation Cdistnet: Perceiving multi-domain character distance for robust text recognition.IJCV, 132(2): 300–318

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T04:55:04.300088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T04:54:52.729257Z digest=sha256:26ed11c9149f8c4e53fecfefc6a7b2a587f317d7b131afc004684bea9f3ba905

Observation 524f0385-de17-4e79-a476-d7699576e3d2 · outbound

This paper cites Explicitly-decoupled text transfer with the minimized background reconstruction for scene text editing.

StyleTextGen: Style-Conditioned Multilingual Scene Text Generation Explicitly-decoupled text transfer with the minimized background reconstruction for scene text editing

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T04:55:04.319121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T04:54:52.729257Z digest=sha256:0757d82cda04f754474affafddd4914e006258a5c5a07532d1d634427284b07e

Pith citing papers

Observation d2e714ed-0934-4fec-b6a4-7827e5626993 · inbound

SlerpFlow: Spherical Trajectory Correction for Rectified Flow Inversion cites this paper.

SlerpFlow: Spherical Trajectory Correction for Rectified Flow Inversion StyleTextGen: Style-Conditioned Multilingual Scene Text Generation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-01T07:57:26.635757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:57:26.635757Z digest=sha256:48eeea14cfad8051a26ae787ec2d9ee4acc5bbe1d1c86c672cf804d7ad5f9a01