Pith. sign in

Paper Citation Record · LEDGER

ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation

As of 19 August 2026, this Paper Citation Record lists 69 of 69 outbound references and 34 inbound Pith citation observations for arXiv:2506.18095.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.18095 v1

Coverage vector

measured 69 of 69 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:28:09.873139Z

measured 103 of 103 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 34 of 34 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T13:05:39.665891Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

69 of 69 outbound references displayed

  • verified exact0
  • verified fuzzy12
  • unresolved56
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 69c5562d-0f13-4a78-8d26-0031208ab9d0 · outbound

This paper cites GPT-4 Technical Report.

ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T23:28:02.069987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:28:02.069987Z digest=sha256:a8265a81eafb263f08aa060709d84dd1044c9c7ee5c401af6b421454ef39d49b

Observation 5b691a77-58c4-4ab2-90ab-953e6a7ba643 · outbound

This paper cites Qwen2.5-VL Technical Report.

ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Qwen2.5-VL Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T23:28:02.150573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:28:02.150573Z digest=sha256:6862cd6c425f38011705a6a4d896a7f5699279cebb4d83a11a64312c7a8023bb

Observation 6c933785-3e08-43d6-b7d8-d0b435ead2e0 · outbound

This paper cites Improving image generation with better captions.

ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Improving image generation with better captions

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T23:28:02.300477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:28:02.300477Z digest=sha256:005b77aa193dc2a3edfb52d6c05a2dc1ec3042c2d97275fc3b357de4c4959054

Observation 70ff78fa-e686-41a1-8739-eadb3ae22e54 · outbound

This paper cites an unresolved cited work.

ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:28:13.805385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T23:28:02.436973Z digest=sha256:e706188c3751633fc9e52ea982dad1eae3c476daeee24a0f7f146d8aaf16033d

Observation ef533ac3-2a70-43fc-b9e2-bd3923428b92 · outbound

This paper cites HiDream-I1: A High-Efficient Image Generative Foundation Model with Sparse Diffusion Transformer.

ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation HiDream-I1: A High-Efficient Image Generative Foundation Model with Sparse Diffusion Transformer

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T23:28:02.515024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:28:02.515024Z digest=sha256:07448baa4e0ddca84bdeb81639862886045b50fdecc9232de7028f3617866654

Observation cbff37d4-5075-466a-a5d1-40fc08ab1094 · outbound

This paper cites Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts.

ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T23:28:02.647600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:28:02.647600Z digest=sha256:b25b4766f9c68109e90ddbdc730257fa61b675b1616ac511cfa81d31b4741e09

Observation eac4d61d-d032-4427-a682-7161337aadda · outbound

This paper cites ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models.

ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T23:28:02.771576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:28:02.771576Z digest=sha256:431c73637b225e78313dab7cba1e075c6c79afec528451753c641985797b02ff

Observation c30384a4-4ee4-4ed4-a4e7-a5cea2b3ee23 · outbound

This paper cites Textdiffuser: Diffusion models as text painters.

ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Textdiffuser: Diffusion models as text painters

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:28:13.571786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T23:28:02.860703Z digest=sha256:ba575479b3b23e3bd1dcdaca2f16d3b02e61a175ec18cf89fc664c0c1521493a

Observation a7146d27-7331-49b9-aca6-3cc16e2da07e · outbound

This paper cites BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset.

ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T23:28:02.977351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:28:02.977351Z digest=sha256:75cd463a55ffc5fb6b557a556a4cf5758f8d4995109ff524e22f142d19cc490d

Observation b8da046a-4135-48c8-be05-d26a0e3795d7 · outbound

This paper cites PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis.

ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T23:28:03.158995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:28:03.158995Z digest=sha256:6e02e9f650d1cc01cd641d22cf9c61ca3728b5cbb21d401c201986f74af533de

Observation 85473dbf-336f-4058-a13a-706dc983d539 · outbound

This paper cites PixArt-\Sigma: Weak-to-Strong Training of Diffusion Transformer for 4K Text-to-Image Generation.

ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation PixArt-\Sigma: Weak-to-Strong Training of Diffusion Transformer for 4K Text-to-Image Generation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T23:28:03.229878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:28:03.229878Z digest=sha256:69464b9604918a5a969d7682825b469132f60606850ee75d5052801302630ecd

Observation 2aaaca21-4a0c-4264-82cf-4fb231c583d6 · outbound

This paper cites Towards injecting medical visual knowledge into multimodal llms at scale.

ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Towards injecting medical visual knowledge into multimodal llms at scale

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:28:13.550252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T23:28:03.351626Z digest=sha256:adccebff615805dc0a51a6e824c7d8eced94de07e53259513c7382d3d05d5ef8

Observation 957fd8ea-0870-43da-8c62-b3105c881098 · outbound

This paper cites An Empirical Study of GPT-4o Image Generation Capabilities.

ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation An Empirical Study of GPT-4o Image Generation Capabilities

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T23:28:03.456631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:28:03.456631Z digest=sha256:0ad26d9646913e8969221c712aaa157e3acf4a668a3fd96bcca7579a459e6c17

Observation d08889de-d609-4182-bf83-dc48d97d6027 · outbound

This paper cites Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling.

ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T23:28:03.554887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:28:03.554887Z digest=sha256:28a8059519f17dda57f9fa7a62c7e44858d972ae50a209ca25510ade300adda7

Observation 15442cd7-75d2-4569-9597-4fe2acd1e963 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:28:13.528954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T23:28:03.658952Z digest=sha256:9f23ed8ffda78df45340f27957362a1fca6f584e1c2ecdd0377bde2ecf8d487e

Observation 40d29074-f65e-410f-a2d0-d123bed6d857 · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Imagenet: A large-scale hierarchical image database

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T23:28:03.736199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:28:03.736199Z digest=sha256:d81a28c4d29ffa9b050a4eeddd8d68118148e2832817888dcbe47495bbf9f8d1

Observation bb9e5f82-cff9-40ff-96da-a3b628019a21 · outbound

This paper cites Taming transformers for high-resolution image synthesis.

ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Taming transformers for high-resolution image synthesis

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T23:28:03.942387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:28:03.942387Z digest=sha256:869378509002dba0789ce38671813dce6e75f963944f66ba00021d6079d9ee92

Observation e4c5993d-a8bc-4bfc-8321-bf08e951258a · outbound

This paper cites Scaling Rectified Flow Transformers for High-Resolution Image Synthesis.

ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Scaling Rectified Flow Transformers for High-Resolution Image Synthesis

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T23:28:04.066433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:28:04.066433Z digest=sha256:aeb110e121dab6b116fe8ea96d171b9c90360a7ea68ebc2b1db29f4eb6cf571e

Observation 6433a298-c677-49e5-a1fd-37950c61de70 · outbound

This paper cites Vitron: A unified pixel-level vision llm for understanding, generating, segmenting, editing.

ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Vitron: A unified pixel-level vision llm for understanding, generating, segmenting, editing

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:28:13.389494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T23:28:04.224741Z digest=sha256:fae8be82d825e5c98261094709cbb659ca1ead678d9d85b6c772e3c9e08aaabc

Observation 87b0722d-66c5-4169-8e41-392d3cf70911 · outbound

This paper cites SEED-Data-Edit Technical Report: A Hybrid Dataset for Instructional Image Editing.

ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation SEED-Data-Edit Technical Report: A Hybrid Dataset for Instructional Image Editing

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T23:28:04.382575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:28:04.382575Z digest=sha256:19a048bfe51fe26512a95e939ba65d1577bc267bc0d0df1348fe7f4936100e11

Observation 6f1804a4-f939-48ba-b28c-95fd060943be · outbound

This paper cites SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation.

ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T23:28:04.533542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:28:04.533542Z digest=sha256:30b319d5a4f8709e78daf66074d1a12576e61d4c658e0093a0778bb13a479bb6

Observation b44d41eb-f8f1-48f9-9326-50c82c9e473c · outbound

This paper cites Geneval: An object-focused framework for evaluating text-to-image alignment.

ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Geneval: An object-focused framework for evaluating text-to-image alignment

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T23:28:04.640332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:28:04.640332Z digest=sha256:f5ad6f43c84134d1a0f480ad6305b8f69c443db4c8a7004f0127850bc7704c3d

Observation 6024e104-6b00-4937-8047-8d428796140e · outbound

This paper cites Scaling Laws for Autoregressive Generative Modeling.

ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Scaling Laws for Autoregressive Generative Modeling

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T23:28:04.820031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:28:04.820031Z digest=sha256:d9e19f3aba1929e4097122d96438fdc9b4585db8fab1f88b4ad999a02804a365

Observation f246f47c-5b94-48ad-9136-c22984658118 · outbound

This paper cites Classifier-Free Diffusion Guidance.

ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Classifier-Free Diffusion Guidance

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T23:28:05.045258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:28:05.045258Z digest=sha256:d28b722b3f7dfc29f37ffac7c3e87aa55ee27d72c46aebb1cc7aa2fa8333765a

Observation 38cfac70-e405-4248-a732-bbf13ee34dc9 · outbound

This paper cites Denoising diffusion probabilistic models.

ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Denoising diffusion probabilistic models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T23:28:05.174734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:28:05.174734Z digest=sha256:e8ed7def9b5c6455900fc1c6a606225a8b7284e402a381e3cbdf36b33cc7bd68

Observation e4e2ec45-017a-4238-a3ab-4b06efe22c14 · outbound

This paper cites ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment.

ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T23:28:05.319362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:28:05.319362Z digest=sha256:1a1efd8469f0813e61753cf75dac9e626e41e4be956a5590dde7dd561877120d

Observation 06002829-b5c3-45a0-afbb-b23882b30f20 · outbound

This paper cites HQ-Edit: A High-Quality Dataset for Instruction-based Image Editing.

ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation HQ-Edit: A High-Quality Dataset for Instruction-based Image Editing

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T23:28:05.426710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:28:05.426710Z digest=sha256:45e6f001e4f4ed5ab979b3a0f2c09582691196c2878092c0d29e0ac32b892765

Observation 7b25588c-f228-4a12-ba65-2cef87527286 · outbound

This paper cites Playground v2.5: Three Insights towards Enhancing Aesthetic Quality in Text-to-Image Generation.

ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Playground v2.5: Three Insights towards Enhancing Aesthetic Quality in Text-to-Image Generation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T23:28:05.513912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:28:05.513912Z digest=sha256:77c11a26f124e85feb7fd3c27f085aabfab55ee34562555334bc8a0e8f2f4baf

Observation 7ba124ae-8510-4cb4-890d-62a2b53796e0 · outbound

This paper cites Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding.

ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T23:28:05.645405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:28:05.645405Z digest=sha256:2e930cf01f5d7b11ad3c1db430c5fb55a4434009151204d9921408714d2a7f0d

Observation e562b3fb-f63d-4cc7-b413-3d1c74b0aa45 · outbound

This paper cites Microsoft coco: Common objects in context.

ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Microsoft coco: Common objects in context

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T23:28:05.789504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:28:05.789504Z digest=sha256:9d8eaf993aaf457eabaa0fddb7187bf3b85043ff42d4c07a4d6665cbea96a806

Observation c300218b-9319-49f5-a65d-ccb32c38d7c8 · outbound

This paper cites World Model on Million-Length Video And Language With Blockwise RingAttention.

ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation World Model on Million-Length Video And Language With Blockwise RingAttention

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T23:28:05.906884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:28:05.906884Z digest=sha256:d5b3c1aa3aabd6451ddd3161bfd0e2478d1f25cd701e4a86ecfb0fcbadfa187d

Observation cff1c216-f4b1-4c03-b46f-4de3610ba0e9 · outbound

This paper cites Visual instruction tuning.

ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Visual instruction tuning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T23:28:06.009347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:28:06.009347Z digest=sha256:51c6d025dc944cdea96a30f53ddef09ea4d5bf029f53826cda86331c9e6ba53b

Observation fc9bb83e-0a06-4809-bc1f-7394afe78b75 · outbound

This paper cites Step1X-Edit: A Practical Framework for General Image Editing.

ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Step1X-Edit: A Practical Framework for General Image Editing

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T23:28:06.131216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:28:06.131216Z digest=sha256:3cdf42b9006b8a45e9dce76e231c22dd300c2264deac95700087c5c2a13ce06d

Observation 3c72b30d-b76f-46e3-904b-5b40ce4cc75e · outbound

This paper cites Unified-io 2: Scaling autoregressive multimodal models with vision language audio and action.

ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Unified-io 2: Scaling autoregressive multimodal models with vision language audio and action

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:28:13.179146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T23:28:06.214740Z digest=sha256:c25155c325d9cf6fbf9234dd3a5406bf9c120aa8ffab10833168e1c63ccfadd0

Observation 274d2a4c-60dc-47e0-afae-701416afd088 · outbound

This paper cites Introducing 4o image generation.

ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Introducing 4o image generation

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:28:13.154825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T23:28:06.334827Z digest=sha256:0aee4dd1edf8f0692f127bdb3c368e9697f5978a3022ffc7511dfc0246567e0a

Observation cb7f369c-0901-4b24-9a8d-7f61448e8fda · outbound

This paper cites Hello gpt-4o, 2024.

ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Hello gpt-4o, 2024

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:28:13.118415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T23:28:06.509959Z digest=sha256:3a1c475c051754eb4b467605eec1886cafb7d9e4bd6409c2c57bdb9dd751359d

Observation 1f9c8120-20d2-4887-b0f5-f3a4ccebf895 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T23:28:06.646624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:28:06.646624Z digest=sha256:758bd820b4dcf5839c439cf1a60202c8cbf0698fd8ca5c337a931ea8f4eec2b9

Observation 833463f7-e50d-4d93-a8f3-1c9cb36c4c79 · outbound

This paper cites Zero-shot text-to-image generation.

ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Zero-shot text-to-image generation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T23:28:06.689684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:28:06.689684Z digest=sha256:4176e8f093bf25de70de8a815260f5646a660518018e47f54b87b11641368a8e

Observation 72909ad8-c2ee-4c84-bf18-a35670526051 · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T23:28:06.766108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:28:06.766108Z digest=sha256:18b562dd421ecafcd18cd3d476a35cc9c873a69046dee789c73c474be5aba49f

Observation 6c462b7d-8688-40c6-a78f-b0e484c67518 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation High-resolution image synthesis with latent diffusion models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T23:28:06.914738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:28:06.914738Z digest=sha256:d9a24127c6c64ae6b93b3e2962f5c19eb39fca120c693e0b5915b2b022b05de0

Observation 12a8bce6-abcc-4ab9-9525-50ad057c2fc2 · outbound

This paper cites Laion-5b: An open large-scale dataset for training next generation image-text models.

ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Laion-5b: An open large-scale dataset for training next generation image-text models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T23:28:06.979295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:28:06.979295Z digest=sha256:cc7afb004a3534472ebb81072a17cf7a3400b158cbb2d2ee9231e89dc4704be4

Observation 222979f4-59dd-4405-ae40-8aaa9832f1ab · outbound

This paper cites Textcaps: a dataset for image captioning with reading comprehension.

ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Textcaps: a dataset for image captioning with reading comprehension

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:28:12.844257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T23:28:07.083550Z digest=sha256:c8fe5e91b0fcafdd34dc5c6f16cf96057b8a7dafc6958b6ed6de1ac3ce4b29ac

Observation 51f3cb0c-7788-42fd-addc-18d22233d511 · outbound

This paper cites Generative modeling by estimating gradients of the data distribution.

ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Generative modeling by estimating gradients of the data distribution

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T23:28:07.165955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:28:07.165955Z digest=sha256:8668f1e78df66e83b15fe9b8eccdbab8b07c0cac4f4c75a3384393d3e0f917d0

Observation 3edbe10a-0c9e-48bd-8f73-267640380da4 · outbound

This paper cites Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation.

ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T23:28:07.246822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:28:07.246822Z digest=sha256:d9872d41037e3037b61bbe9733ba15b3615f7dd5eda32d44dfa231b7a817fa54

Observation f3c691b7-ca7e-43f3-9127-bcf2e679a695 · outbound

This paper cites Emu: Generative Pretraining in Multimodality.

ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Emu: Generative Pretraining in Multimodality

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T23:28:07.325623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:28:07.325623Z digest=sha256:ccbf0da732c74113a57f0eafba883de868bc30b0fd64db14ac72de95fc27540b

Observation 2dc2524e-c15d-4508-9692-c276ee0a6476 · outbound

This paper cites Multimodal Latent Language Modeling with Next-Token Diffusion.

ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Multimodal Latent Language Modeling with Next-Token Diffusion

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T23:28:07.466483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:28:07.466483Z digest=sha256:6de7dc6fc02809d84c92b4bba5867ca2d1be273f68619fc95221b90e8a6eb005

Observation fdd28907-eb26-4f3d-9ab0-5626dd9ad9f3 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T23:28:07.564744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:28:07.564744Z digest=sha256:5c1b9c283a6414cbf1ca4023dc148c29f6f9cb7e5c55e79e495c5db518452097

Observation 2c17f4fa-c2bd-4259-9522-c95de7b73cfc · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Gemini: A Family of Highly Capable Multimodal Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T23:28:07.666376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:28:07.666376Z digest=sha256:2e7901bad54d48debea474423869d7db0e6f5923107661978ab43bc4d795dad5

Observation 6c6b19f1-d144-4ac2-baae-41fc252e9ec4 · outbound

This paper cites Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction.

ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T23:28:07.776955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:28:07.776955Z digest=sha256:eb9aca44951fb080d6210a8b9976fd87138789f6eb8421bc504aa7c6ba570e65

Observation c624ce04-be73-430f-ac7c-a92fdeb31ee0 · outbound

This paper cites AnyText: Multilingual Visual Text Generation And Editing.

ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation AnyText: Multilingual Visual Text Generation And Editing

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T23:28:07.866027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:28:07.866027Z digest=sha256:58a37692b15bbc35f47ad9879c26430113fe763278ee981293d4ff0fc57423e6

Observation 50751418-9b2e-4298-80b4-0b39d30c5a8c · outbound

This paper cites Textatlas5m: A large-scale dataset for dense text image generation.

ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Textatlas5m: A large-scale dataset for dense text image generation

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T23:28:07.957576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:28:07.957576Z digest=sha256:5f84334ab1c178eba70daf9a5d20376d2fb15c283c243f9b958d19529ccbc0e3

Observation bf8346c1-e658-4964-a6ab-e8e635871ba8 · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Emu3: Next-Token Prediction is All You Need

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T23:28:08.063683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:28:08.063683Z digest=sha256:42b8f25916fd718b0c80265963ba9b12f86ab573bfce2043c088a22930fce1b1

Observation dfdb897c-4045-4ef7-a3e9-6674083af233 · outbound

This paper cites Janus: Decoupling visual encoding for unified multimodal understanding and generation.

ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Janus: Decoupling visual encoding for unified multimodal understanding and generation

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:28:12.599256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T23:28:08.191964Z digest=sha256:d108caa1b5c9f771dabf38f5cdde29eb6fd1bd2c896ca8c97908f3480d696c7a

Observation 495838a0-7edd-4c79-a3d1-63b7286c119d · outbound

This paper cites Next-gpt: Any-to-any multimodal llm.

ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Next-gpt: Any-to-any multimodal llm

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:28:12.344746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T23:28:08.271954Z digest=sha256:7c0e97495bdc1e5a9cd21931b6c3be50966e96c5ae8efec886ad479ccb6fa58a

Observation ddc754d1-d7c9-41eb-957b-146175a044f5 · outbound

This paper cites Show-o: One Single Transformer to Unify Multimodal Understanding and Generation.

ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T23:28:08.396644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:28:08.396644Z digest=sha256:ebf1d9a78c879b4272e08b852d12e29e8042de739301ea34d9408645deedc3d6

Observation 8c02d195-e9af-4ad9-9409-7113542331ad · outbound

This paper cites GPT-ImgEval: A Comprehensive Benchmark for Diagnosing GPT4o in Image Generation.

ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation GPT-ImgEval: A Comprehensive Benchmark for Diagnosing GPT4o in Image Generation

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T23:28:08.514846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:28:08.514846Z digest=sha256:8eb5420579cb1a1436074dfc22d882d2dd65ed101706cc099d996cacec7c45d2

Observation 0742d445-1c92-44cc-84ee-7aa31c60195b · outbound

This paper cites ImgEdit: A Unified Image Editing Dataset and Benchmark.

ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation ImgEdit: A Unified Image Editing Dataset and Benchmark

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T23:28:08.633139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:28:08.633139Z digest=sha256:032ab47813698b3399a0ec3b149f95258fbb5d6630546af71ec9f4bad38a360d

Observation 78361dec-7c30-48e3-bd4c-a52ff3363c04 · outbound

This paper cites Scaling Autoregressive Models for Content-Rich Text-to-Image Generation.

ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T23:28:08.742835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:28:08.742835Z digest=sha256:2398d16438ae83b0215fbab1c033aa1c70dc86550af5eba50d003462850c1df7

Observation 5a485763-c51a-4191-a88b-d9b7a0856edf · outbound

This paper cites AnyEdit: Mastering Unified High-Quality Image Editing for Any Idea.

ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation AnyEdit: Mastering Unified High-Quality Image Editing for Any Idea

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T23:28:08.784667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:28:08.784667Z digest=sha256:8e44e25025475eb22cf406ca25e6e8ef04b77211e640eb6ff7264b0e89ea81db

Observation 8a224390-44a0-45a7-8793-be16ff421cfa · outbound

This paper cites AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling.

ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T23:28:08.892356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:28:08.892356Z digest=sha256:f0e1793f36e880c8e8cc5ff0371520d978d385ecc85d03e28038d88c8098532a

Observation 2d8274e0-8a91-441b-a62f-d36fd59cca34 · outbound

This paper cites Nexus-Gen: Unified Image Understanding, Generation, and Editing via Prefilled Autoregression in Shared Embedding Space.

ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Nexus-Gen: Unified Image Understanding, Generation, and Editing via Prefilled Autoregression in Shared Embedding Space

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T23:28:09.017945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:28:09.017945Z digest=sha256:dc900e727abb0607c827c017bbc2c912d0fb6d282421ff0a2345f049df86d077

Observation 1116264d-5982-416e-b4cb-21b2edcf599a · outbound

This paper cites Magicbrush: A manually annotated dataset for instruction-guided image editing.

ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Magicbrush: A manually annotated dataset for instruction-guided image editing

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:28:11.964652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T23:28:09.124231Z digest=sha256:993133f71571adf246ca1df46198675edcbd5856bada4fb626361ea28ff65df9

Observation 6a1d66f3-ab3f-4fee-89fb-2bce5ce3dbf8 · outbound

This paper cites Ultraedit: Instruction-based fine-grained image editing at scale.

ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Ultraedit: Instruction-based fine-grained image editing at scale

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:28:11.572391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T23:28:09.246658Z digest=sha256:970c9b0333b5bc342b924ca48c2a536a5b7315f674055a109fccd665b58f2d3e

Observation bfb47594-c57c-4467-b740-9912eabe5eb3 · outbound

This paper cites Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model.

ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T23:28:09.349741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:28:09.349741Z digest=sha256:7cff7ccace785620285ada3c875c476f7537df10ab642d8894bb3d100bc4dab1

Observation 6e8b3b05-e52d-4294-b1d1-4cad0a13b504 · outbound

This paper cites Lumina-Next: Making Lumina-T2X Stronger and Faster with Next-DiT.

ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Lumina-Next: Making Lumina-T2X Stronger and Faster with Next-DiT

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T23:28:09.448639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:28:09.448639Z digest=sha256:08e277ad99a6c6a7c1d4e3108423c152d63e27c3d6fc3b6c8b39306f7dce7703

Observation bd4dc842-09ed-4014-b8b8-caa0e77bb21a · outbound

This paper cites write newline.

ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation write newline

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T23:28:09.580228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:28:09.580228Z digest=sha256:85971b04ab28af26bb67fd2ce9b588797814957f4980f166fa4c5655ae3a3da8

Observation b1018978-483a-44b2-8dad-4850576be72d · outbound

This paper cites @esa (Ref.

ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation @esa (Ref

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T23:28:09.673093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:28:09.673093Z digest=sha256:87a07edb2f50f91d6cc6d56d466e6d39791d608e0a62d06f8e18577fe8f7f4cc

Observation e9f828fd-6031-4f3a-8220-d88fd0f0b425 · outbound

This paper cites an unresolved cited work.

ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Unresolved cited work

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T23:28:09.791138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:28:09.791138Z digest=sha256:7e5f5782eaad2f3f54c5f292e6db7d5df09682ba31c541f22562818852d52262

Observation da56fdde-9320-4ccc-9888-19350a643d65 · outbound

This paper cites v6kL8 | c6 +zP? z ۷ 7?Y| -w?iiiEޔCT֐ t1gΜjsss + B M ]UVڡth<3 E NUݺu WM .\(P& @ס^^^= ;+[i4o ԗ/_FuPg- >ƹs ݴ_4.

ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation v6kL8 | c6 +zP? z ۷ 7?Y| -w?iiiEޔCT֐ t1gΜjsss + B M ]UVڡth<3 E NUݺu WM .\(P& @ס^^^= ;+[i4o ԗ/_FuPg- >ƹs ݴ_4

Reference 69

Resolution
malformed identifier
raw_fallback, observed 2026-08-06T23:28:11.254750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T23:28:09.873139Z digest=sha256:1cc879c5d15e79379788f34cfc4b76801df03792187bd4c34cb678b743ba96c5

Pith citing papers

Observation 309f22f9-9241-4f0d-b7d2-f8eccc15f32e · inbound

GPT-IMAGE-EDIT-1.5M: A Million-Scale, GPT-Generated Image Dataset cites this paper.

GPT-IMAGE-EDIT-1.5M: A Million-Scale, GPT-Generated Image Dataset ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T13:05:39.665891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:05:39.665891Z digest=sha256:1b1c13208827dfb4e1441d383e7bf7b6a300e57cf44e0b273153711cc619f56b

Observation 97a4601e-976b-4aee-a89a-3b446283c17c · inbound

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation cites this paper.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:11.031053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:11.031053Z digest=sha256:0ade37c711e075af44592a9754b92e52aee57530daab6405c93a5f18e0cadc8a

Observation 5975655a-e5c2-4bd4-8859-fb51ab03caff · inbound

Draw-In-Mind: Rebalancing Designer-Painter Roles in Unified Multimodal Models Benefits Image Editing cites this paper.

Draw-In-Mind: Rebalancing Designer-Painter Roles in Unified Multimodal Models Benefits Image Editing ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-18T20:06:50.141079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T20:04:00.773443Z digest=sha256:d7c691aad57589ddd3b9e231f25c568e03587692903f08e703e8363a1eef7dc4

Observation 9761af27-5252-44df-9dd5-1a00c12bafb7 · inbound

Reconstruction Alignment Improves Unified Multimodal Models cites this paper.

Reconstruction Alignment Improves Unified Multimodal Models ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T22:36:07.886211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T22:36:07.886211Z digest=sha256:41fc776bd576c68b5de505f6a90e6ec29a84bea247833beaf8ad0e72a4ce24ca

Observation fc41ad36-ad34-4c40-83e6-3745a1c768ae · inbound

Lavida-O: Elastic Large Masked Diffusion Models for Unified Multimodal Understanding and Generation cites this paper.

Lavida-O: Elastic Large Masked Diffusion Models for Unified Multimodal Understanding and Generation ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T15:39:34.264959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T15:39:34.264959Z digest=sha256:8cdefd86229ef44d7d9f872dd62d1632f69e9dbd511fcc0d8032e23c6fd7c34d

Observation 5aa384f7-4c39-4008-bbee-35e3365c1c77 · inbound

Emu3.5: Native Multimodal Models are World Learners cites this paper.

Emu3.5: Native Multimodal Models are World Learners ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:12:13.530814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:13039e8a8e2fc84fcbe36ea08c97f8b8d00a5cdd472d0b4252db222f6305a3ab

Observation 8a09de47-3500-4297-b64b-a42980b220fa · inbound

Distribution Matching Distillation Meets Reinforcement Learning cites this paper.

Distribution Matching Distillation Meets Reinforcement Learning ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T21:47:19.795407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:47:19.795407Z digest=sha256:c42d8fd7cd931dcb872a4128379b970cbbf9fcfd3c3167a61665d4d236eab9c4

Observation 754a545d-74ad-4438-a302-90c6c1f46be8 · inbound

AIA: Rethinking Architecture Decoupling Strategy In Unified Multimodal Model cites this paper.

AIA: Rethinking Architecture Decoupling Strategy In Unified Multimodal Model ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-17T04:19:00.627838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-17T04:17:07.534291Z digest=sha256:6514dd0a2ac6c06e9621f5775e8444708bf6788d62f38e6ae73d67b6c65a1900

Observation cb519f03-a043-4030-8e65-b652ae1f883a · inbound

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models cites this paper.

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T16:21:31.813465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:21:31.813465Z digest=sha256:e0f6b312ff00e2559ed172aff64de37044c1912f930dc9a7c34cd9c2b46fb191

Observation fa3362e7-21a0-402f-b312-f19d70ea8568 · inbound

PlanViz: Evaluating Planning-Oriented Image Generation and Editing for Computer-Use Tasks cites this paper.

PlanViz: Evaluating Planning-Oriented Image Generation and Editing for Computer-Use Tasks ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:00:42.763568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-16T07:00:18.807922Z digest=sha256:f79f2c9023b7112507b2ad1bca9a5dd3aa2f20a14c03c7123e7c318bdf18af49

Observation e2391c0d-6086-4f42-b788-8402b6c26b5d · inbound

WeEdit: A Dataset, Benchmark and Glyph-Guided Framework for Text-centric Image Editing cites this paper.

WeEdit: A Dataset, Benchmark and Glyph-Guided Framework for Text-centric Image Editing ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T18:25:52.501547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:25:52.501547Z digest=sha256:bd7c854fa1a8cd5e0fabe2b06f06523d4e22180a705be1668c6d11ef505c8fe5

Observation f36e7ec0-3247-4d13-aab6-fb00b16848f4 · inbound

Automating Crash Diagram Generation Using Vision-Language Models: A Case Study on Multi-Lane Roundabouts cites this paper.

Automating Crash Diagram Generation Using Vision-Language Models: A Case Study on Multi-Lane Roundabouts ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:50:04.432062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T14:46:41.877762Z digest=sha256:0ad697740996b073c810da933b7f2a7055716fb579dfb02733ddc9f882b714c7

Observation 4837438a-6459-479e-a0b3-ecbd2e941b85 · inbound

IncreFA: Breaking the Static Wall of Generative Model Attribution cites this paper.

IncreFA: Breaking the Static Wall of Generative Model Attribution ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:36:01.509277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T05:35:46.381465Z digest=sha256:673473a72d1aa9f373564f4fe156e2cd90e0f05442a54e33e04bef6ad7824f75

Observation 09230663-92aa-4ccf-a58c-26893973f84b · inbound

Extending One-Step Image Generation from Class Labels to Text via Discriminative Text Representation cites this paper.

Extending One-Step Image Generation from Class Labels to Text via Discriminative Text Representation ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:28:39.596721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T05:15:22.907880Z digest=sha256:fe9bed05216089de419e73caa5d1a67e820c82700224f007b99dad45f3eb48f2

Observation d9d9b23f-5dde-4f85-93ee-e9369d2e2083 · inbound

Guiding Distribution Matching Distillation with Gradient-Based Reinforcement Learning cites this paper.

Guiding Distribution Matching Distillation with Gradient-Based Reinforcement Learning ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T02:48:27.456349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T02:45:35.600729Z digest=sha256:5e34287b4d48bd2f9b62b323ae892b21efd2b421fecb05fd0edec8454043e17c

Observation 036cb635-679c-45df-b969-4de48f914a80 · inbound

MUSE: Resolving Manifold Misalignment in Visual Tokenization via Topological Orthogonality cites this paper.

MUSE: Resolving Manifold Misalignment in Visual Tokenization via Topological Orthogonality ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation

Reference 91

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:36:08.006096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-08T15:04:41.518195Z digest=sha256:e42ca19b4a99dd95a535a7f4f622bce1d59f8e3538aa665d6aa3beb5aee43d04

Observation bb07873a-e4b5-4054-b07a-f76fb25daa70 · inbound

Auto-Rubric as Reward: From Implicit Preferences to Explicit Multimodal Generative Criteria cites this paper.

Auto-Rubric as Reward: From Implicit Preferences to Explicit Multimodal Generative Criteria ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:36:24.278633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-12T00:58:53.626310Z digest=sha256:afddba12ccfa248a2d331238009a198c18927aae03be04acc92785b74f61baae

Observation 316fed8e-229c-45ae-93e4-189f4ff73509 · inbound

FlashAR: Efficient Post-Training Acceleration for Autoregressive Image Generation cites this paper.

FlashAR: Efficient Post-Training Acceleration for Autoregressive Image Generation ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-12T02:21:16.540851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-12T02:20:25.071428Z digest=sha256:2353b8a3b7bcdbd48acfc445e557b90e504e338f17da7eca62cde5ddb5aff7cc

Observation 0c7faca6-810c-47fd-a7bb-81d967daa41a · inbound

FlashAR: Efficient Post-Training Acceleration for Autoregressive Image Generation cites this paper.

FlashAR: Efficient Post-Training Acceleration for Autoregressive Image Generation ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:02:23.630811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-13T05:54:24.248910Z digest=sha256:da5e38cb2d7d8afa9c668e8cb49a4be457cae66ca029eed9cae741c7db70fbdd

Observation d6a7db49-d335-4b4a-810a-1172fdb05aef · inbound

Inline Critic Steers Image Editing cites this paper.

Inline Critic Steers Image Editing ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:52:57.810329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-14T20:52:10.428866Z digest=sha256:7dbdce20e20eb15831db55da2b82301b95767d8d4c357ca687233923dd83cb6b

Observation 1766dc85-7e6e-424f-9f0d-7b44974e8ed6 · inbound

NanoCP: Request-Level Dynamic Context Parallelism for Data-Expert Parallel Decoding cites this paper.

NanoCP: Request-Level Dynamic Context Parallelism for Data-Expert Parallel Decoding ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-21T01:53:54.532915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-21T01:51:38.671365Z digest=sha256:0baa09841bb6d035e0d62587284fa187c152418d71083e5a9d7333cbc1d6dadf

Observation e6560117-548a-4d13-be28-0661929cf818 · inbound

Bernini: Latent Semantic Planning for Video Diffusion cites this paper.

Bernini: Latent Semantic Planning for Video Diffusion ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-22T06:41:10.556131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T06:39:47.124605Z digest=sha256:2533f74f11f8e5ef34ce557485b3cbf7918b1928ba89888074f7634ea90cdbcb

Observation 85c82c7e-a6da-4d5c-8a6c-2801a4224f68 · inbound

Reinforcing Few-step Generators via Reward-Tilted Distribution Matching cites this paper.

Reinforcing Few-step Generators via Reward-Tilted Distribution Matching ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:24:00.240761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-29T22:21:06.005125Z digest=sha256:8ef2874edd48bac6afcda15a855953e9839ce0ef4ee65346d7829e02d2fe45bf

Observation b5f23702-f0d4-4c15-b8c3-f3327b84f720 · inbound

Imagine Before You Draw: Visual Prompt Engineering for Image Generation cites this paper.

Imagine Before You Draw: Visual Prompt Engineering for Image Generation ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation

Reference 50

Resolution
malformed identifier
arxiv_id, observed 2026-07-02T07:06:43.968385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T07:12:24.297839Z digest=sha256:517ce5531c08034771896a4aefcd9090af6fc30c977c52517a8e2fdcc08bfdbb

Observation c6f74ddb-5aa5-4c90-aa36-b29908259095 · inbound

Is This Edit Correct? A Multi-Dimensional Benchmark for Reasoning-Aware Image Editing cites this paper.

Is This Edit Correct? A Multi-Dimensional Benchmark for Reasoning-Aware Image Editing ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-12T20:04:25.886330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T20:04:25.886330Z digest=sha256:cfd60275ca0e8eb104e5bc957b3a74c9eaf8ad72a94aae1e43af6b2b7e5d3b63

Observation f0b73f3c-addd-4bdd-8dfc-27ba463084f6 · inbound

ARM: An AutoRegressive Large Multimodal Model with Unified Discrete Representations cites this paper.

ARM: An AutoRegressive Large Multimodal Model with Unified Discrete Representations ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-03T05:07:38.513361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T13:29:11.526106Z digest=sha256:a29d85a182333b7d64fdff63d9f266dae30296f9b792e07db76ede4cda028eae

Observation ab01bb93-776e-4a24-beeb-28189b032601 · inbound

Unified Multimodal Autoregressive Modeling with Shared Context-Visual Tokenizer is Key to Unification cites this paper.

Unified Multimodal Autoregressive Modeling with Shared Context-Visual Tokenizer is Key to Unification ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:28:55.875492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T01:18:03.846908Z digest=sha256:3c89326a3f1712e31e16fb67e1bb433da57709d0f70f85ebf395d29f5da0775b

Observation fe67454a-814a-44fd-ae3c-dd1fc17f989f · inbound

SPAR: Semantic-Pixel Self-Alignment and Adaptive Routing for Unified Multimodal Models cites this paper.

SPAR: Semantic-Pixel Self-Alignment and Adaptive Routing for Unified Multimodal Models ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T10:09:44.712827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-26T09:08:25.661515Z digest=sha256:87a40b7aaefe55a69e648be6c30f9e5aca889c3f7bd7e91126ccfe1d6afd4f0b

Observation 1de9217b-9bef-4351-9d86-57918e098cee · inbound

SPAR: Semantic-Pixel Self-Alignment and Adaptive Routing for Unified Multimodal Models cites this paper.

SPAR: Semantic-Pixel Self-Alignment and Adaptive Routing for Unified Multimodal Models ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T23:19:02.551077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-03T23:15:09.253879Z digest=sha256:ba2d9e40d7220ac7621b08a2460eb706f32bad6d52564c416921807751a9b729

Observation 641635b8-ad85-4a14-92e2-7750f4291c44 · inbound

IV-CoT: Implicit Visual Chain-of-Thought for Structure-Aware Text-to-Image Generation cites this paper.

IV-CoT: Implicit Visual Chain-of-Thought for Structure-Aware Text-to-Image Generation ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation

Reference 56

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T16:39:58.132069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-26T00:19:49.071495Z digest=sha256:82419041ff63835a56e1ae742ab63e5f41d17a21056e050ea4ac0b55c0338bd9

Observation 103e9039-0169-4e0a-ab9d-31bb526dea8c · inbound

Nemotron-Labs-Diffusion-Image: Advancing Masked Discrete Diffusion for High-Resolution Image Synthesis cites this paper.

Nemotron-Labs-Diffusion-Image: Advancing Masked Discrete Diffusion for High-Resolution Image Synthesis ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-06-30T08:14:26.639738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T06:07:55.600338Z digest=sha256:7a961164e87e5ae8477b43cd63679b8b9a829035c9b3e3d296d28798ab0c8c44

Observation ecef6b90-21a0-4183-9f66-858677c0af61 · inbound

Nemotron-Labs-Diffusion-Image: Advancing Masked Discrete Diffusion for High-Resolution Image Synthesis cites this paper.

Nemotron-Labs-Diffusion-Image: Advancing Masked Discrete Diffusion for High-Resolution Image Synthesis ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-02T09:39:38.579281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:39:38.579281Z digest=sha256:6a0aceb8aa04ee93d759546a5da85816e44c71e86ed783603e608f2675631e35

Observation ddcc576d-155f-4567-ac5a-5e08fe62b0e1 · inbound

Bridging Video Understanding and Generation in a Unified Framework cites this paper.

Bridging Video Understanding and Generation in a Unified Framework ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-01T10:05:40.348481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-01T05:57:54.653504Z digest=sha256:9ec7d7ff3027acb4bdc8d9d7211d87bc0fc873bbad465eaf5e830718f6503649

Observation 7c2e7998-274a-41b1-a194-d37e832b4a2d · inbound

Amortized Moment Matching for Visual Generation cites this paper.

Amortized Moment Matching for Visual Generation ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation

Reference 150

Resolution
unresolved
no resolver link, observed 2026-07-30T18:58:28.492234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T18:58:28.492234Z digest=sha256:791cc5189fc60808e03395bce66a8e653878663e94442df3355e638bca0ddaa4