Pith. sign in

Paper Citation Record · LEDGER

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

As of 8 August 2026, this Paper Citation Record lists 87 of 87 outbound references and 34 inbound Pith citation observations for arXiv:2508.09987.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.09987 v1

Coverage vector

measured 87 of 87 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T20:44:18.490290Z

measured 121 of 121 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 34 of 34 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T06:47:10.728416Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T21:10:09.706550Z

Reference resolution

87 of 87 outbound references displayed

  • verified exact0
  • verified fuzzy15
  • unresolved72
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8617fb45-85f5-4cac-b6d2-2f2376b89421 · outbound

This paper cites Lawrence Zitnick, Devi Parikh, and Dhruv Batra.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Lawrence Zitnick, Devi Parikh, and Dhruv Batra

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:10.478127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:10.478127Z digest=sha256:c8c6d419592ffdd821ac5a4e388522fc41d12250c4e10c11cb7d3eee5c6cd2df

Observation 971a41a5-30bb-4562-b464-0237a213c23b · outbound

This paper cites Sd3-medium.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Sd3-medium

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:10.536355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:10.536355Z digest=sha256:a2342600de66cf4b8b19733c9c6d3f87a2ab81bd9fdcce69c2deb828ed4b9cfd

Observation 97555354-04a7-49f3-9c76-1e9d20533c84 · outbound

This paper cites Qwen Technical Report.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Qwen Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:10.622883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:10.622883Z digest=sha256:dfcbd5b23d3955222d5f60dcbf519b04dbc0308501c4c5b8698a7aee9f909f2c

Observation b99b6dda-b6a7-4f10-985a-3a24a173b7f9 · outbound

This paper cites Qwen2.5-VL Technical Report.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Qwen2.5-VL Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:10.688300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:10.688300Z digest=sha256:0b97a751a227b58446edf794c86996bbf835a58dbf9b6b51d5ea6abdcce03a0b

Observation f773e167-5a88-44cd-be4e-e7ebc581fa0d · outbound

This paper cites Instructpix2pix: Learning to follow image editing instructions.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Instructpix2pix: Learning to follow image editing instructions

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:10.745556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:10.745556Z digest=sha256:d795db5d4c2d7d9e163ccd66b85765df9d13501df53072c226ebe58386dfe7d5

Observation 34d256f9-ab71-4f3d-a38d-cc6df7905906 · outbound

This paper cites Allava: Harnessing gpt4v-synthesized data for a lite vision-language model, 2024 a.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Allava: Harnessing gpt4v-synthesized data for a lite vision-language model, 2024 a

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:10.827910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:10.827910Z digest=sha256:b0cb7a5d0e07d853a71d9984ecf42883ddb47c5c3c46e2fe9ddf3594360a493c

Observation fb06e333-9da3-4435-ba91-0b2c8208054e · outbound

This paper cites BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:10.890407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:10.890407Z digest=sha256:6565e67d7f14b9b7336586074af6eb7f92f00b7fc53d39dacfc2cd9de84e748f

Observation 0a9f6bb9-d2f5-455e-8ebc-1323b3f2d242 · outbound

This paper cites PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:10.965681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:10.965681Z digest=sha256:e815061d20abf30e37cce2b8e3f95d80555a0fcb54e1c19dd1cb332a594e1a9c

Observation 97a4601e-976b-4aee-a89a-3b446283c17c · outbound

This paper cites ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:11.031053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:11.031053Z digest=sha256:f7e2f1011d7fbe63a0bbae454cddf2a4dcaf2ef3301b51d51f93a1b9549192b1

Observation cc1c951d-86d4-43dc-88e6-8323404fe7e1 · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:11.140230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:11.140230Z digest=sha256:6c166f86b99725eb2479ba3721447b9dc15558d988f0fc7f2956bcaff1630ce5

Observation a9b7a0c5-af64-46f3-9b24-bc9e6904c3cb · outbound

This paper cites Sharegpt4video: Improving video understanding and generation with better captions.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Sharegpt4video: Improving video understanding and generation with better captions

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:11.209115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:11.209115Z digest=sha256:176eaad29b8670a374a9c27b24d5a3a0ea2951785199572c4d3d18d048018e3a

Observation 0413e382-52db-46f9-8719-9efc0395d285 · outbound

This paper cites Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:11.307662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:11.307662Z digest=sha256:6324b5c10c2bdc10aedd46b29e59f1ba6d2e89fa96e0b54e0d4b532aaf5c6a7d

Observation 669b36a1-b70e-4040-8632-c5031034c7f0 · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Gonzalez, Ion Stoica, and Eric P

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:11.389528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:11.389528Z digest=sha256:0b108ef790198c22ac77c62948dcadf6cf0c70465ddfc23e11ef86c6888af2f5

Observation a496aa47-6186-4223-bf92-5645788ca429 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Training Verifiers to Solve Math Word Problems

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:11.474169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:11.474169Z digest=sha256:c32c5764a08dd9a1e77b981c5e4deb19824c9e0ec8d46087b891effce3186497

Observation e08b543a-3a34-488d-a6e8-45e8dd3a0fbf · outbound

This paper cites Emerging Properties in Unified Multimodal Pretraining.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Emerging Properties in Unified Multimodal Pretraining

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:11.546425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:11.546425Z digest=sha256:890053096c3289e228ce4896ae13a4a8da6baf575a638996ed89bad510c7ec94

Observation 2495bef3-5bb7-47dd-8736-453fa1f105bd · outbound

This paper cites Autoregressive Video Generation without Vector Quantization.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Autoregressive Video Generation without Vector Quantization

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:11.606456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:11.606456Z digest=sha256:d88abfb7dcf6218f3bf8145957de3494bfafe45720af15bd7a08cefd426df88b

Observation 21cde564-0d55-4a8e-a3ae-3cdf3dcd7b23 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:11.699910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:11.699910Z digest=sha256:0399c2465015dbbc12b43391ffb289c97fd283196c2d76147e172ab26229af1b

Observation 46ce3066-ded4-48b6-a4fc-a590c1e153e8 · outbound

This paper cites Scaling rectified flow transformers for high-resolution image synthesis.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Scaling rectified flow transformers for high-resolution image synthesis

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:11.819994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:11.819994Z digest=sha256:4a25a8113dd18346098fefe076c2291daabe3abf31d264afa14467a1787b903e

Observation 013bb6d9-fef6-4856-a0fd-0524bf938b8f · outbound

This paper cites GenEval: An Object-Focused Framework for Evaluating Text-to-Image Alignment.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation GenEval: An Object-Focused Framework for Evaluating Text-to-Image Alignment

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:11.883278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:11.883278Z digest=sha256:4bdd2ad7e8c87cff95356e5c197094c8915580fa9c0d55baf8706bd395417c7f

Observation 0c5d3d43-a58b-42ee-8789-60210e68f368 · outbound

This paper cites Geneval: An object-focused framework for evaluating text-to-image alignment.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Geneval: An object-focused framework for evaluating text-to-image alignment

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:44:23.523614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T20:44:11.977252Z digest=sha256:6ff30351f9f1420990fd8000c4c5b3217235a5e576dbdc853e425f704be3ab66

Observation 6f1fba65-6838-45c5-9b27-3fb226a9fba5 · outbound

This paper cites Geneval: An object-focused framework for evaluating text-to-image alignment.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Geneval: An object-focused framework for evaluating text-to-image alignment

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:12.088623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:12.088623Z digest=sha256:9d57e62fcd4d4de30f30aa71e4eb3d9fcc021372784655ebf73075f333479e02

Observation c8cae32e-ceaa-4ea0-aeb3-6c90f0bf8e42 · outbound

This paper cites Gemini 2.0 flash.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Gemini 2.0 flash

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:44:23.361409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T20:44:12.163832Z digest=sha256:e3b1e9569ab646548e14e3ff0ede0c7027750bedaad4ff564567dc4405d8b8b8

Observation 9254bbaa-5f05-4d51-bb8e-8dda23a2fa65 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:12.258440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:12.258440Z digest=sha256:3b2b509bca1deadef4071ffabf8466005649b72f35632e4ea58489e6888a8ba3

Observation 3cb084c7-d4b8-4bcc-b561-a4bbe46415cd · outbound

This paper cites PiSA: A Self-Augmented Data Engine and Training Strategy for 3D Understanding with Large Models.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation PiSA: A Self-Augmented Data Engine and Training Strategy for 3D Understanding with Large Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:12.369370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:12.369370Z digest=sha256:9966b2577f455e099a95cddebae0f53a4393135064c95258c3578d981aff1f9a

Observation 1c41b22a-8423-4dd8-9ff9-b14b3a0b11d8 · outbound

This paper cites Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:12.432203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:12.432203Z digest=sha256:a912913996f43e84b2252d9bea8a2e0a6ee28cacf58423828418c742dfb81b5f

Observation 5ac129a1-469f-4d0e-8853-84325d4ba9fd · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Measuring Mathematical Problem Solving With the MATH Dataset

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:12.504055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:12.504055Z digest=sha256:c3c4f69751e0f8a62efba0f9edef0ce831d478785527fde8cdb404a06ea31a61

Observation 97dabc6a-db8a-40ef-beeb-1ff6075eb843 · outbound

This paper cites ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:12.564771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:12.564771Z digest=sha256:7ffbd95905bdc58a0b9aaea6c62e23e5621562aa59068eb9d614bb77037004f2

Observation 7a515664-b7a0-4767-a102-1e81076cb6f6 · outbound

This paper cites Tifa: Accurate and interpretable text-to-image faithfulness evaluation with question answering.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Tifa: Accurate and interpretable text-to-image faithfulness evaluation with question answering

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:44:23.131160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T20:44:12.627331Z digest=sha256:62de5514d44af83962b6d7495bc9896cf9c33fb1187ae05975a99be9252a51dc

Observation 536242bd-79f5-4edf-b5cc-89f572b67e5e · outbound

This paper cites T2I-CompBench++: An Enhanced and Comprehensive Benchmark for Compositional Text-to-image Generation.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation T2I-CompBench++: An Enhanced and Comprehensive Benchmark for Compositional Text-to-image Generation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:12.680706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:12.680706Z digest=sha256:dbd96eb95bed0c91c048f0267fd20c928c3967b38dc76a8835a3d1d0f0baf9c2

Observation c7d45430-4670-46bf-abdf-01758b3f0373 · outbound

This paper cites CoMat: Aligning Text-to-Image Diffusion Model with Image-to-Text Concept Matching.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation CoMat: Aligning Text-to-Image Diffusion Model with Image-to-Text Concept Matching

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:12.755158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:12.755158Z digest=sha256:7f9d38d444669451db1bdc28e1bf300fe65a134f1074d12b01209e65ab009899

Observation b065b8fc-0f30-42c5-adfa-00e05b1bc5f6 · outbound

This paper cites MMSearch: Benchmarking the Potential of Large Models as Multi-modal Search Engines.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation MMSearch: Benchmarking the Potential of Large Models as Multi-modal Search Engines

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:12.841584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:12.841584Z digest=sha256:fc092db350d7f64f4b608b900bb39033a1a33338c8f28f40991f731f7b8b9ba3

Observation ac2ac4a8-37f7-4a10-9308-141a84c5d824 · outbound

This paper cites T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:12.948621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:12.948621Z digest=sha256:90ee6f9c4b6974bf660fef36d3124ad5d79455e41ed9bfddd1d85536edb9f3c5

Observation 537b8e27-1d48-45fe-9349-229f5710d9d2 · outbound

This paper cites MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:13.035019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:13.035019Z digest=sha256:09dd800c5a96f1c2304779e7a37a13270fc5dd2a918d6cf84f30f17880e5f23e

Observation eeee2d81-7577-4268-aec4-0d0e8cb0bc80 · outbound

This paper cites Auto-Encoding Variational Bayes.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Auto-Encoding Variational Bayes

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:13.131965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:13.131965Z digest=sha256:91cf2fcbbb1666703ce184d0a3e8924ea33d7fe67b8b66f1a95a034c0a39631c

Observation 9f0e7380-5565-4469-8545-0d46e55d4818 · outbound

This paper cites Viescore: Towards explainable metrics for conditional image synthesis evaluation, 2023.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Viescore: Towards explainable metrics for conditional image synthesis evaluation, 2023

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:13.239307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:13.239307Z digest=sha256:a911f5351a12f6db778094e50ba1715ee7bc07201f749f5f1ec8b9e9702c21b8

Observation d76b024a-e6ab-4322-a1b3-20081f7708fd · outbound

This paper cites an unresolved cited work.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:13.299556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:13.299556Z digest=sha256:41b6f997d89531fe8a58a74072cff2f3d019b6fc4067912dedb6a952bf70e5db

Observation d33fa69b-7a28-4909-9114-fbebdc2da3be · outbound

This paper cites GenAI-Bench: Evaluating and Improving Compositional Text-to-Visual Generation.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation GenAI-Bench: Evaluating and Improving Compositional Text-to-Visual Generation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:13.367107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:13.367107Z digest=sha256:08f46d18b125a8ac6a32f9fce7a0f5e4881089c44f2347e31f612db53b371f32

Observation e8fecf9f-ad82-47a4-839e-4c0e0c6197b2 · outbound

This paper cites CrossViewDiff: A Cross-View Diffusion Model for Satellite-to-Street View Synthesis.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation CrossViewDiff: A Cross-View Diffusion Model for Satellite-to-Street View Synthesis

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:13.455328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:13.455328Z digest=sha256:158ec981cb32d2e635e8ac9de92583a4f83d4083d8819b2813d9a29ec0d92f4a

Observation e7854012-5878-4145-a2f3-f0b127f40fdd · outbound

This paper cites Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:13.554546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:13.554546Z digest=sha256:dff9d707856cde7edb359a5fce70c06088f0e1b2679a740dc45ddc4ad70c7df1

Observation 1edcc029-c7a2-4424-9eaa-778cfe197606 · outbound

This paper cites UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:13.670112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:13.670112Z digest=sha256:43fc0ba33a5c1f917bbd57419d4890c8f6625d8e6b35b5db68762de0f12a8e52

Observation 96bb787e-a117-4b9c-b4d2-0490c681795f · outbound

This paper cites Microsoft coco: Common objects in context.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Microsoft coco: Common objects in context

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:13.852362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:13.852362Z digest=sha256:c54fa021b9a820079d80918b40488999d9bf2ea74ff5a8a4f44ab4f64727ed33

Observation 15f05b24-c517-4cf2-a712-3970452c8e49 · outbound

This paper cites Evaluating Text-to-Visual Generation with Image-to-Text Generation.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Evaluating Text-to-Visual Generation with Image-to-Text Generation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:13.931759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:13.931759Z digest=sha256:526b26a82a09bf40a4a4420dd75b93eb98cdd1a1946d469c0b563fa47ff9a365

Observation ef264d48-ff5c-4941-9c5c-0fde57f60077 · outbound

This paper cites Visual instruction tuning.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Visual instruction tuning

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:44:22.745660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T20:44:14.038559Z digest=sha256:0193a53c64efe34ad652cb2c4880754c4e643bfec58569dca701002a2dcd45ea

Observation b3b2deda-1d67-44f3-ae94-ec315b1903d6 · outbound

This paper cites Visual instruction tuning.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Visual instruction tuning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:14.104263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:14.104263Z digest=sha256:f9d6e6585d4ac4e7b6bfb9b1674d7cd970ff18c3e1c49835dcdf9a4144cf8b09

Observation 29b277f6-3478-4afb-b12d-ca87597a64cc · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation MMBench: Is Your Multi-modal Model an All-around Player?

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:14.255400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:14.255400Z digest=sha256:d7823ff5ccb46c084468d0596b79c53bf803a57071721e33b9a1a75130f9f177

Observation b74d1110-41e6-40c1-9fd0-3fde59ad31c3 · outbound

This paper cites The Llama 3 Herd of Models.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation The Llama 3 Herd of Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:14.390201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:14.390201Z digest=sha256:0819caeb1e69e94547932536ace83e89242b4718a5b1190cabb7b88c319d8ccb

Observation e491591a-c5fb-4ea4-9fb1-c8ef4e85a808 · outbound

This paper cites an unresolved cited work.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Unresolved cited work

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:14.465050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:14.465050Z digest=sha256:ed7e4b55120c50054815b924c9931f87d22ebd2f53ffe9b4704cb2988b0bd60b

Observation 8b0ca092-b47e-4b6a-b559-4a2cc6f71a60 · outbound

This paper cites GPT-4 Technical Report.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation GPT-4 Technical Report

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:14.544936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:14.544936Z digest=sha256:4ace57c391ff848ddf92eb0ed266cd58a198d6bc7d2d46995465b2c2eec394c5

Observation 6b7be0b3-184d-4b69-8eb5-b3ca8f3a3970 · outbound

This paper cites GPT-4V(ision) system card, 2023 c.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation GPT-4V(ision) system card, 2023 c

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:44:22.415811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T20:44:14.654402Z digest=sha256:66202b58b728426dd480d40a4abca77438d8632066552033dce9292b8f397f74

Observation 700e44bb-11b0-4977-9a4b-b32311c363a4 · outbound

This paper cites Dall·e 3.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Dall·e 3

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:14.813143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:14.813143Z digest=sha256:daadba4b8ae75ca57662ddfec35a19d45e4a1018daad1f18c724d09cd646abe8

Observation 786b6bb7-1fe6-4836-9248-a1b61e0b8f3a · outbound

This paper cites an unresolved cited work.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-05T20:44:22.066520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T20:44:14.924248Z digest=sha256:f08d9f6aa67bf3d5b9348caf77a2b4111f7a48bcc008f2ebabf926c1e87d4963

Observation ce6d28a4-dd5a-4a0d-b3b8-3b3878bf56ec · outbound

This paper cites an unresolved cited work.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-05T20:44:21.653553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T20:44:15.018814Z digest=sha256:7a66fd5376e05504771c1b1a630e7b75a7955ea5ca49f13b05f19cfef9f2378e

Observation bc58b3d3-28a5-44a2-9759-bd83c6352742 · outbound

This paper cites Transfer between Modalities with MetaQueries.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Transfer between Modalities with MetaQueries

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:15.111083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:15.111083Z digest=sha256:6f0c90f54a03dfc6e1ae9ac0b00912c0564e8e5be8f1cd09b663dba153116a44

Observation 1268b0aa-837c-497b-9fbd-7c962c5667d5 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:15.208196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:15.208196Z digest=sha256:a43dce44f13117d2a71dbeae1f480b73e325457cca89938ba10f948163254d8d

Observation d626d2a7-e5f8-4924-b316-49e5c9a29f2b · outbound

This paper cites Tokenflow: Unified image tokenizer for multimodal understanding and generation.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Tokenflow: Unified image tokenizer for multimodal understanding and generation

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:44:21.328873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T20:44:15.342084Z digest=sha256:24aa441a23f8f3e0a5db110b9a677141532306335422604e7003b104b7f9d7bd

Observation d4ce61b1-0977-44b7-943a-a93b7694132a · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:15.466929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:15.466929Z digest=sha256:a0ec858e467e4885b667091a9fa1cee4f4a356daae18be93849a42137766bf00

Observation 1565b050-13d7-44d7-89b9-1d9e7d97a88b · outbound

This paper cites Gpqa: A graduate-level google-proof q&a benchmark.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Gpqa: A graduate-level google-proof q&a benchmark

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:15.560365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:15.560365Z digest=sha256:2fe0bde72b3b92ebf65ce1c25214ba200ae3a213274febcab1a2ba595f439785

Observation e043e2f4-65c6-4329-a8f4-45e308aabdc9 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation High-resolution image synthesis with latent diffusion models

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:44:21.096417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T20:44:15.634551Z digest=sha256:d35a40c18044bf948a419a4a5f2995065337c22e74ba62959354b8027fddf769

Observation 7e1b5794-57e4-43d6-ae34-5f736ffe59bc · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation High-resolution image synthesis with latent diffusion models

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:44:20.943984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T20:44:15.688147Z digest=sha256:c28b38a8b18f2c0e9a97a74dce44d771e01a2f0c67d88740e87c7ea1bfa4965e

Observation 86ca084c-fab6-4139-a63d-6d3ac42082a7 · outbound

This paper cites Journeydb: A benchmark for generative image understanding.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Journeydb: A benchmark for generative image understanding

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:15.762368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:15.762368Z digest=sha256:4aefa789416164c44092e8730944c593fbf3a9391d699ebe9428fe82a1ad3f41

Observation 0751e862-9eb7-42f8-90f8-edd4ff0e57a6 · outbound

This paper cites Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:15.870519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:15.870519Z digest=sha256:0360373a98af270eda950b324925b6e0adedfeba4e77c9a2ecd3fc6e35e57b3f

Observation eec3108d-0d5d-42b9-8fe0-56002d0fe716 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation LLaMA: Open and Efficient Foundation Language Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:16.057133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:16.057133Z digest=sha256:b51fcfc07d36b3089d7211b7cac6dd783e9a6b58a165895f587a2367c75e6227

Observation 2adcf704-cc20-496a-9476-b17d332f7b92 · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Emu3: Next-Token Prediction is All You Need

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:16.154776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:16.154776Z digest=sha256:56d88e0dd3529064e41662d0f7aea5c4581f03a98e078976a82a446230c1d7b2

Observation ccef5db5-15eb-4134-9cfd-857a8cfa91a8 · outbound

This paper cites GPT-IMAGE-EDIT-1.5M: A Million-Scale, GPT-Generated Image Dataset.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation GPT-IMAGE-EDIT-1.5M: A Million-Scale, GPT-Generated Image Dataset

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:16.254186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:16.254186Z digest=sha256:8c22c79b8eecdb123887f2b29a4d359b5cc12fa5df8072088ff433a01b245aaa

Observation 31c843ec-e643-4c38-9c8d-3749a0c7960a · outbound

This paper cites TIIF-Bench: How Does Your T2I Model Follow Your Instructions?.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation TIIF-Bench: How Does Your T2I Model Follow Your Instructions?

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:16.332457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:16.332457Z digest=sha256:33db794a14fdf9eb1fdf50ca76bce0046c0ac702208ed40efa95a56e3dc3b8ff

Observation b202c619-3cc5-410c-a751-e1392b7e1a01 · outbound

This paper cites Janus: Decoupling visual encoding for unified multimodal understanding and generation.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Janus: Decoupling visual encoding for unified multimodal understanding and generation

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:16.378281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:16.378281Z digest=sha256:99d26dbc58d953f0644674c23579669b075d375472989e5f43000322b6fb2e0c

Observation 276f6e35-9182-4dfc-a0df-9f322ea10bf2 · outbound

This paper cites OmniGen2: Towards Instruction-Aligned Multimodal Generation.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:16.475472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:16.475472Z digest=sha256:b30828d1575cd3db52d205939eb4229d2cfb51a8d7cb51b805d5e819a160e2a0

Observation 6770393b-c88b-4663-8b23-dc95edadc091 · outbound

This paper cites Less-to-More Generalization: Unlocking More Controllability by In-Context Generation.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Less-to-More Generalization: Unlocking More Controllability by In-Context Generation

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:16.671224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:16.671224Z digest=sha256:8d4753bfd7bd97a7805a41970d57ded096c0e9f34432c297b9d8155dd761aca3

Observation eeea55b8-05d8-4c65-b7ac-5545717ab100 · outbound

This paper cites Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:16.792997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:16.792997Z digest=sha256:b12b92681b8b557d104055f8eaf2bfc6707bcb415947528b50beabf6f1ed8d66

Observation 22868f17-5beb-4301-b17d-3d16e94d28f9 · outbound

This paper cites Omnigen: Unified image generation.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Omnigen: Unified image generation

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:44:20.728214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T20:44:16.871964Z digest=sha256:f23f684d666b63dc2be3ab67a48e2ed87bf9088dafb6656671c5813f331b5563

Observation d3fbf88e-5d6a-48df-a184-85ccf5454910 · outbound

This paper cites Show-o: One Single Transformer to Unify Multimodal Understanding and Generation.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:16.966007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:16.966007Z digest=sha256:c7634f215326beb91066a746069b69bd9b41eb41f9088b85e6ecbfee7d04f645

Observation 33200aa2-a6a0-47a6-a36b-2cecc935da35 · outbound

This paper cites VisionReward: Fine-Grained Multi-Dimensional Human Preference Learning for Image and Video Generation.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation VisionReward: Fine-Grained Multi-Dimensional Human Preference Learning for Image and Video Generation

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:17.042652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:17.042652Z digest=sha256:2b6dc689ab218702b900f1e7207a18f7341abc176dafa45cf8cc0845df5d6671

Observation 6a496748-550e-4b94-9a5a-8e115ba8c9a7 · outbound

This paper cites PointLLM: Empowering Large Language Models to Understand Point Clouds.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation PointLLM: Empowering Large Language Models to Understand Point Clouds

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:17.158713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:17.158713Z digest=sha256:9e4634449463c67210c1c0914ee84dcda26ea7c2a79658ccad98881c4b1c4a22

Observation c0c1295e-055b-4fd9-a112-8a432002e628 · outbound

This paper cites GPT-ImgEval: A Comprehensive Benchmark for Diagnosing GPT4o in Image Generation.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation GPT-ImgEval: A Comprehensive Benchmark for Diagnosing GPT4o in Image Generation

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:17.270406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:17.270406Z digest=sha256:5903254532ebb3060c1ba6d36f2abd4c83139ca4fc3fcaf1b9aa4dbcd2ca24fb

Observation 8ab861fb-602b-4843-91ac-80f88e1e7445 · outbound

This paper cites The fabrication of reality and fantasy: Scene generation with llm-assisted prompt interpretation.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation The fabrication of reality and fantasy: Scene generation with llm-assisted prompt interpretation

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:44:20.513604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T20:44:17.333380Z digest=sha256:05d721c16568a2820c59b95d33cd120ab88f702278e0cb5c93be1f752c7d5b28

Observation 152a51c7-8452-4b4f-a891-a2d42955a03b · outbound

This paper cites Skydiffusion: Street-to-satellite image synthesis with diffusion models and bev paradigm.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Skydiffusion: Street-to-satellite image synthesis with diffusion models and bev paradigm

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:44:20.319888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T20:44:17.442397Z digest=sha256:fccf142d26e780aa6902759aa659c03c85aa76d5a9922338d256b85c5042e20d

Observation 48755bd2-c57a-48c5-aec5-8dd61a2ccc70 · outbound

This paper cites LOKI: A Comprehensive Synthetic Data Detection Benchmark using Large Multimodal Models.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation LOKI: A Comprehensive Synthetic Data Detection Benchmark using Large Multimodal Models

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:17.585671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:17.585671Z digest=sha256:fd044dc7c18107752093d159142eb626540f280ecc93d70323230c938d3288fe

Observation 8b0dddf8-44ef-465b-a070-e08ecac09a8b · outbound

This paper cites Sigmoid loss for language image pre-training.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Sigmoid loss for language image pre-training

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:44:20.123840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T20:44:17.696061Z digest=sha256:6c4ce81aeccdf26ec84339d5cc62f5e2ed3751233b1fe41c2eefe2acda792e51

Observation 71c09a01-7924-40c1-98d8-da3eda218ec8 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:17.843196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:17.843196Z digest=sha256:0fed1aa48e6ac3a066da1c23857e773d951d30029719d46fe7fce9c1a1e5d2fd

Observation 368ac5c2-1d7a-43de-aef5-1e26b2a2556e · outbound

This paper cites Adding conditional control to text-to-image diffusion models.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Adding conditional control to text-to-image diffusion models

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:44:19.903316Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T20:44:17.937041Z digest=sha256:3d7b8dba6eb369f718288b090812aeaa65c485e11fb10b33fd87ed8790cdd646

Observation 1fbe3f38-a77e-41fb-a23d-3a9c4e3cf118 · outbound

This paper cites Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? In European Conference on Computer Vision, pages 169--186.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? In European Conference on Computer Vision, pages 169--186

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:44:19.701308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T20:44:18.003309Z digest=sha256:14cff8477bc28bccec5a9bc1ec6b1c56c84a3f6999b60b50fbeb98c23e0446d2

Observation 1c842d2f-a386-4325-bfd3-8d7b4aa80f9d · outbound

This paper cites MAVIS: Mathematical Visual Instruction Tuning with an Automatic Data Engine.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation MAVIS: Mathematical Visual Instruction Tuning with an Automatic Data Engine

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:18.084915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:18.084915Z digest=sha256:00b97d2aa2a1a97d8587a4994eda95060a85c012a42feb30a808b152ddf3bdc0

Observation 57992e02-9dd8-4f1c-b753-d67c0251ccf2 · outbound

This paper cites Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:18.171068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:18.171068Z digest=sha256:a0cfcdc37c99fecdb4830723fc0286819d27fffb9aac8e8e92a06f84b195a8bc

Observation 5fe73f18-f381-49f7-b573-9eb9369a696d · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:18.224864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:18.224864Z digest=sha256:b7cf76b33284928bab77fae198522476b815781d8c4230503e49e6f99bd87180

Observation 6210cb88-0136-4d27-ac1b-3b0f2b9bb64d · outbound

This paper cites Lumina-next: Making lumina-t2x stronger and faster with next-dit.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Lumina-next: Making lumina-t2x stronger and faster with next-dit

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:44:19.536676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T20:44:18.313183Z digest=sha256:0070da6aac547d7728b67c355a561d1b32d681c24b74c8533ee8f64d802686af

Observation 598ab3a7-4c0c-4cc4-a8a6-16ea8eceeea2 · outbound

This paper cites EasyRef: Omni-Generalized Group Image Reference for Diffusion Models via Multimodal LLM.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation EasyRef: Omni-Generalized Group Image Reference for Diffusion Models via Multimodal LLM

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:18.395972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:18.395972Z digest=sha256:637935d6b51096c689cdcbda1b86e13d8b6a0edd74fcccd2b6b2616c58fe16a7

Observation 0ff93221-48a4-4c12-9888-aff39271040f · outbound

This paper cites MoVA: Adapting Mixture of Vision Experts to Multimodal Context.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation MoVA: Adapting Mixture of Vision Experts to Multimodal Context

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:18.490290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:18.490290Z digest=sha256:7d1cb197788cd656c7ecf500e81bbe975343ec376b28f3847c9debfb01566755

Pith citing papers

Observation 930cac71-df1c-4ade-85a8-a73f533250b2 · inbound

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation cites this paper.

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:10.728416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:10.728416Z digest=sha256:7c87e9ed54f3b670692924a7676bfce1bdb5ce08cbbc49e5ffa4f0d10cda3c08

Observation 9077d993-71cf-4d75-a590-62ba585e9563 · inbound

MICo-150K: A Comprehensive Dataset Advancing Multi-Image Composition cites this paper.

MICo-150K: A Comprehensive Dataset Advancing Multi-Image Composition Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 95

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:21:23.408189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T00:20:58.483350Z digest=sha256:4679234a70066383c6eb730efffd18a2f0067a55e01d83d21370068c1d469833

Observation ee11affc-1452-4348-b2b3-13d6e7e85949 · inbound

InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMs cites this paper.

InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMs Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:10:45.467647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T08:09:19.209759Z digest=sha256:07410d51ff4466e3040ea60c2ce0ed63dd8db7da216dbe20669abe424b46b39a

Observation 3727f093-580b-42bd-b965-57a7b0d45c8e · inbound

Prompt Reinjection: Alleviating Prompt Forgetting in Multimodal Diffusion Transformers for Text-to-Image Generation cites this paper.

Prompt Reinjection: Alleviating Prompt Forgetting in Multimodal Diffusion Transformers for Text-to-Image Generation Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:04:10.562099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T13:00:22.875471Z digest=sha256:5af31dbcb144cb259d3cfe60f2425d5ac8c87734a19908040a7ec914e9f7582b

Observation 4b0d842f-3bb5-43bd-b628-7734f017fe59 · inbound

Prompt Reinjection: Alleviating Prompt Forgetting in Multimodal Diffusion Transformers for Text-to-Image Generation cites this paper.

Prompt Reinjection: Alleviating Prompt Forgetting in Multimodal Diffusion Transformers for Text-to-Image Generation Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T03:50:46.874348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:50:46.874348Z digest=sha256:f9382d1f204a57822aa06e4ceccec1509aadaf6c634d28cd602c846cb93ee8d7

Observation 431cceef-dc42-4ef3-a2e3-06fdd4459c29 · inbound

WeEdit: A Dataset, Benchmark and Glyph-Guided Framework for Text-centric Image Editing cites this paper.

WeEdit: A Dataset, Benchmark and Glyph-Guided Framework for Text-centric Image Editing Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-02T18:25:58.290824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:25:58.290824Z digest=sha256:8eca86727ef2bb3ef566fd81f31e2154806f64c5fac0309fedf6af38617088c0

Observation 87d8668a-d32e-40c4-b5ad-1c8834f1d5ef · inbound

LIVE: Leveraging Image Manipulation Priors for Instruction-based Video Editing cites this paper.

LIVE: Leveraging Image Manipulation Priors for Instruction-based Video Editing Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 53

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T07:21:55.099717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T07:21:41.483427Z digest=sha256:d638c2997f8a4ba4b311c3ec91fbf87d7c49e2e5d555803af416c66038de0f9c

Observation 578ef19e-c334-4815-b475-25e8afa0b1a0 · inbound

IncreFA: Breaking the Static Wall of Generative Model Attribution cites this paper.

IncreFA: Breaking the Static Wall of Generative Model Attribution Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:36:01.504198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T05:35:46.381465Z digest=sha256:7853fd09b2bf5bfdd0be05e7aa2dd490232cf6b1c7b61e0b5a35598b7d6440a5

Observation 3d432ef3-d8fb-47df-b223-a5d3b74ead1e · inbound

Extending One-Step Image Generation from Class Labels to Text via Discriminative Text Representation cites this paper.

Extending One-Step Image Generation from Class Labels to Text via Discriminative Text Representation Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:28:39.636695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T05:15:22.907880Z digest=sha256:e5d3cba93c9f05f391bcf7475b26dc4e3556b72a3ce917bed9123a9bc92bd7eb

Observation 3eec311c-890e-447f-aa40-f6e2275e829d · inbound

LLaDA2.0-Uni: Unifying Multimodal Understanding and Generation with Diffusion Large Language Model cites this paper.

LLaDA2.0-Uni: Unifying Multimodal Understanding and Generation with Diffusion Large Language Model Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-10T00:49:48.357764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T00:49:38.156237Z digest=sha256:773388dbc849ea680b1938b19d5e8ca880ec3975debb97f084b41d30a0fd65e6

Observation 04c45296-5970-461d-b729-da28578ca476 · inbound

Automated In-the-Wild Data Collection for Continual AI Generated Image Detection cites this paper.

Automated In-the-Wild Data Collection for Continual AI Generated Image Detection Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:05:35.272848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T19:00:45.628715Z digest=sha256:6f137017926c457d985e84aad4b45be8fff5eb916b4bc0abd9e5de1d157fdda7

Observation a69c4dbc-3e41-46e2-8883-0988202182e8 · inbound

DiffCap-Bench: A Comprehensive, Challenging, Robust Benchmark for Image Difference Captioning cites this paper.

DiffCap-Bench: A Comprehensive, Challenging, Robust Benchmark for Image Difference Captioning Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:55:44.257603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T17:56:47.621050Z digest=sha256:4e8a3e919a6f0821ea697f586819ec5b5b2e56b8bffa6ae625f89720aef1dfc1

Observation e6f3e836-e3b6-44dd-b520-5a428ec346ac · inbound

HiDream-O1-Image: A Natively Unified Image Generative Foundation Model with Pixel-level Unified Transformer cites this paper.

HiDream-O1-Image: A Natively Unified Image Generative Foundation Model with Pixel-level Unified Transformer Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 59

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T07:32:29.604208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T07:30:53.939221Z digest=sha256:849578a04e1b1243b2a99d7263fe756c912cbc06086c900a81ae67486b313ace

Observation 5a634008-15f1-4d99-9bdb-a99ff3c88d2d · inbound

UniPath: Adaptive Coordination of Understanding and Generation for Unified Multimodal Reasoning cites this paper.

UniPath: Adaptive Coordination of Understanding and Generation for Unified Multimodal Reasoning Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:47:04.888423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T01:29:23.774196Z digest=sha256:3ee2cc249e4b1f949b9c365d8be198e2c9de7140322ec0e2b2f028fa70724ef9

Observation 9cb5b6af-48b2-49d6-8066-3d516d37c56f · inbound

UniCustom: Unified Visual Conditioning for Multi-Reference Image Generation cites this paper.

UniCustom: Unified Visual Conditioning for Multi-Reference Image Generation Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:02:23.959544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T05:53:21.851578Z digest=sha256:5ae13f52aa2106609632bc019dfb2d8bca4fc3297d5483e4c1aae4a9f07fce00

Observation 8b878cdd-2cd2-48db-8273-c4b711f3440d · inbound

UniCustom: Unified Visual Conditioning for Multi-Reference Image Generation cites this paper.

UniCustom: Unified Visual Conditioning for Multi-Reference Image Generation Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:03:03.420134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-14T22:00:01.349754Z digest=sha256:5154856cae0fd78c9758e4f88a491c69d540ec41497e55fb30567ae1d9c6b542

Observation 9f65ce3f-0c10-407a-ab3a-167a1e28387d · inbound

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation cites this paper.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:52:22.717435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:9c5ed018dee57361b233e6cf838ac51d99b06378b6e8aa622b3a0c27e3e3e88d

Observation d80e7656-d23a-42e3-8b1d-56b62a9a65ae · inbound

InsightTok: Improving Text and Face Fidelity in Discrete Tokenization for Autoregressive Image Generation cites this paper.

InsightTok: Improving Text and Face Fidelity in Discrete Tokenization for Autoregressive Image Generation Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:13:30.544636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T02:10:18.004724Z digest=sha256:06704b9b07b0c4ba1624bec66104062b59d8c7ceca7532f458d1dbea9633eea4

Observation 7b84827b-75db-4d6c-856e-0c62ab14ec03 · inbound

Unlocking Complex Visual Generation via Closed-Loop Verified Reasoning cites this paper.

Unlocking Complex Visual Generation via Closed-Loop Verified Reasoning Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-19T16:37:39.705633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T16:35:43.166697Z digest=sha256:67a1cbd9a924a06dce1f8395b48c1accd17d43b53072f249c56978bef8c16256

Observation 22edd390-9aec-43c3-90e6-85f988829d6e · inbound

TextSculptor: Training and Benchmarking Scene Text Editing cites this paper.

TextSculptor: Training and Benchmarking Scene Text Editing Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-21T05:19:39.346678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T05:16:43.756525Z digest=sha256:1cab156172f7a1c215a46d1ecaa4a8acbc6f3950026fea16bc434eaa58b94f6b

Observation 820dacfe-2939-4c1e-835f-8d8cdf1b3416 · inbound

VINS-120K: Ultra High-Resolution Image Editing with A Large-Scale Dataset cites this paper.

VINS-120K: Ultra High-Resolution Image Editing with A Large-Scale Dataset Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-25T04:25:19.730880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-25T04:21:05.811662Z digest=sha256:0486c861ed8b3c3d0bd0f88ff3f9768d67ac3e78bbae50a2fe91c553324a2eae

Observation c0059742-8e6d-4142-a723-dbbc63bd864c · inbound

FakeVLM-R1: Internalizing Physical Laws via CoT for Synthetic Image Detection cites this paper.

FakeVLM-R1: Internalizing Physical Laws via CoT for Synthetic Image Detection Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:03:14.105435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T07:59:44.436264Z digest=sha256:304a91328ffd99f35f3e4258cf5b40c781192ea0082728dbfb68eae408c23453

Observation d90f1426-8151-49e6-9048-e1ba002e6484 · inbound

GenClaw: Code-Driven Agentic Image Generation cites this paper.

GenClaw: Code-Driven Agentic Image Generation Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-06-29T07:43:13.991999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T07:36:54.292448Z digest=sha256:ac175056ef5c1d2b0e0da45ad8870fb8e7ac30f817db1b17c854c24f8721d08f

Observation c4e0c9e5-a423-4714-88cd-6dd7da9438c5 · inbound

MemoGen: Can Past Experience Improve Future Text-to-Image Generation? cites this paper.

MemoGen: Can Past Experience Improve Future Text-to-Image Generation? Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:36:27.150642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T10:48:18.630777Z digest=sha256:6b6c0e464803d09a5896215f6b34a90496089183be1da8467f0ec18e6fd9bcd1

Observation abf57ffa-361c-4e17-9e30-b478241da339 · inbound

IV-CoT: Implicit Visual Chain-of-Thought for Structure-Aware Text-to-Image Generation cites this paper.

IV-CoT: Implicit Visual Chain-of-Thought for Structure-Aware Text-to-Image Generation Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 57

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T16:39:58.135045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-26T00:19:49.071495Z digest=sha256:d61b49077b51116f09facf4b82899a651ecb57e260c6b15a8f3136d6d77c8245

Observation 560ea876-8511-4e1e-867e-76970259df4b · inbound

DomainShuttle: Freeform Open Domain Subject-driven Text-to-video Generation cites this paper.

DomainShuttle: Freeform Open Domain Subject-driven Text-to-video Generation Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-07-04T21:10:09.708346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-25T19:00:23.260939Z digest=sha256:8324dccf6bb7019c19e12354776299e74f307388f95c5baea8c42c37f31841a3

Observation f3feccb5-939d-4652-a186-3fb74814f15f · inbound

Qwen-Image-Agent: Bridging the Context Gap in Real-World Image Generation cites this paper.

Qwen-Image-Agent: Bridging the Context Gap in Real-World Image Generation Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T13:19:50.702108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T05:19:21.433049Z digest=sha256:3cbc8c58d2e33afc2946346f9c9e1662923d4ed01b20d3d72afd245607e63290

Observation f665122f-5ab2-4958-a065-9e5e97710d82 · inbound

Qwen-Image-Agent: Bridging the Context Gap in Real-World Image Generation cites this paper.

Qwen-Image-Agent: Bridging the Context Gap in Real-World Image Generation Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T19:03:52.050125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T04:58:40.065005Z digest=sha256:48507fea9c2d319b46d159a6e9b845ca0422fd26abc0bb1c9e8399a81cea8dab

Observation 4bc666a7-0f71-40ba-82a3-fcba477715e7 · inbound

SciIR: A Large-scale Training Dataset and Benchmark for Scientific Image Reasoning Generation cites this paper.

SciIR: A Large-scale Training Dataset and Benchmark for Scientific Image Reasoning Generation Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 46

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T06:54:20.742715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T06:48:22.594002Z digest=sha256:39b070db9b06875c06c3961eb4d1e1c233036fcccb9397356fda5bdec5756ed2

Observation d10be50f-ff23-4c43-9fd5-9c5543cfb045 · inbound

Bridging Video Understanding and Generation in a Unified Framework cites this paper.

Bridging Video Understanding and Generation in a Unified Framework Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-07-01T10:05:40.288631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-01T05:57:54.653504Z digest=sha256:e82902d83b68a969d6b1734bab95064cd3ba9f4730eba13884d235f21d8af4cd

Observation 6d962e50-6fa4-47a8-9724-bb4ce6867b2d · inbound

StructGen: Disambiguating Multi-Reference Image Generation via Structured Context Modeling cites this paper.

StructGen: Disambiguating Multi-Reference Image Generation via Structured Context Modeling Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T22:47:42.650721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:47:42.650721Z digest=sha256:1ee4f07df838474bc8561bc6dfb030f39b324251423898d04297f9ec5536fc0d

Observation a22c9aec-68da-4707-8639-adf7a373f25f · inbound

ExpertVerse: A General-Purpose Benchmark for Expert-Level Reasoning in Knowledge-Intensive Visual Synthesis cites this paper.

ExpertVerse: A General-Purpose Benchmark for Expert-Level Reasoning in Knowledge-Intensive Visual Synthesis Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T12:46:26.424119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:46:26.424119Z digest=sha256:cc8bac347d921fe9b3b85970143ce2e635a1299092d15a584f0cce0061dfd624

Observation e6e8cc14-ba48-4b0f-8e93-6b5832611c76 · inbound

Argus-Unified: Towards A Compact and Economical Unified Model for Image Understanding and Generation cites this paper.

Argus-Unified: Towards A Compact and Economical Unified Model for Image Understanding and Generation Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-01T02:14:09.772449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:14:09.772449Z digest=sha256:49d6007296fc0b60f734bf72e76e78ee352ae1159e56e2b7c4c551c5bf75e8b0

Observation e2f9fab1-341b-4ec4-b735-2dbbf7c15ceb · inbound

Amortized Moment Matching for Visual Generation cites this paper.

Amortized Moment Matching for Visual Generation Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 151

Resolution
unresolved
no resolver link, observed 2026-07-30T18:58:28.496196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T18:58:28.496196Z digest=sha256:93793ca1b1fbda00b0e58790910bec4181ecf75ff1acf45e7b066882966c993c