Pith. sign in

Paper Citation Record · LEDGER

Emu3.5: Native Multimodal Models are World Learners

As of 21 August 2026, this Paper Citation Record lists 100 of 131 outbound references and 68 inbound Pith citation observations for arXiv:2510.26583.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2510.26583 v1

Coverage vector

measured 100 of 131 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-18T01:12:13.426640Z

measured 168 of 168 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 68 of 68 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-14T04:34:59.990040Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

100 of 131 outbound references displayed

  • verified exact44
  • verified fuzzy55
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 22c3e724-4a22-4659-acc2-a1b148d5b13c · outbound

This paper cites GPT-4 Technical Report.

Emu3.5: Native Multimodal Models are World Learners GPT-4 Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-18T01:12:13.685502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:a8cf85b2bc46f71096e42ad4d94091eef8a70a55714d412f9873e0596610952a

Observation f7c019b8-984e-4390-a746-08f421e500ad · outbound

This paper cites GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints.

Emu3.5: Native Multimodal Models are World Learners GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-18T01:12:13.629916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:0a7b5bda0d15144145213208d7ccde81c118e58f9226e49fc65e74af47a716f5

Observation 31ac4d8c-5c96-431f-9dad-d90d2164b4b9 · outbound

This paper cites Claude 3.5: An ai assistant by anthropic.

Emu3.5: Native Multimodal Models are World Learners Claude 3.5: An ai assistant by anthropic

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T01:12:13.721953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:d1cae547fa64ffcebc8d8f73f5fd65bedb56177432eb44b0b8afabcad1b6be01

Observation fa6a8267-4b96-43b8-b9f4-d5d8103fabf1 · outbound

This paper cites The chosen one: Consistent characters in text-to-image diffusion models.

Emu3.5: Native Multimodal Models are World Learners The chosen one: Consistent characters in text-to-image diffusion models

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T01:12:13.724912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:944d1c60c8bb82af1f82da936ab66cd08ec766248f1e7bc46bfcbc4b2f6422aa

Observation 9a960c7f-f3ff-4341-be5c-2f77286b509c · outbound

This paper cites Qwen2.5-VL Technical Report.

Emu3.5: Native Multimodal Models are World Learners Qwen2.5-VL Technical Report

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-18T01:12:13.656605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:fdea4bf6351835489647ba20ad2028f3bcd7b7dc3ecb257d3f7ce9ef01bf1017

Observation c77f1912-16e6-47fd-bfff-100e39170382 · outbound

This paper cites Improving image generation with better captions.https://cdn.openai.com/ papers/dall-e-3.pdf.

Emu3.5: Native Multimodal Models are World Learners Improving image generation with better captions.https://cdn.openai.com/ papers/dall-e-3.pdf

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T01:12:13.727363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:3c10cf7c7b1f8cabcd8d2c3ed708a591ffef163d944d99b9846c1eedd739a5da

Observation d26faa91-7994-4244-a6c5-96c5e73265dd · outbound

This paper cites Instructpix2pix: Learning to follow image editing instructions.

Emu3.5: Native Multimodal Models are World Learners Instructpix2pix: Learning to follow image editing instructions

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T01:12:13.730610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:f0e8dee7e5d7babc51f2a77499101a8060d5404d466ca7316cf98f72d441dfa8

Observation b3c29c2b-d468-48a9-8659-e1e421197366 · outbound

This paper cites AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems.

Emu3.5: Native Multimodal Models are World Learners AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-18T01:12:13.482985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:fa631e95d0d08b865600203c3c03be0678ba2c185ec0ad6f22b34f9d755ab293

Observation 073cd66e-f006-41e3-8698-411ccea589ea · outbound

This paper cites Coyo-700m: Image-text pair dataset.https://github.com/kakaobrain/coyo-dataset.

Emu3.5: Native Multimodal Models are World Learners Coyo-700m: Image-text pair dataset.https://github.com/kakaobrain/coyo-dataset

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T01:12:13.733812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:85f448806e7c9fcc629fdd4d876992c158c433e157613402fc0f961adcf945be

Observation f4217f92-b9a0-4c3a-bfd6-c74a315df1b0 · outbound

This paper cites HiDream-I1: A High-Efficient Image Generative Foundation Model with Sparse Diffusion Transformer.

Emu3.5: Native Multimodal Models are World Learners HiDream-I1: A High-Efficient Image Generative Foundation Model with Sparse Diffusion Transformer

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-18T01:12:13.637118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:4abb1845b1128544f55b962934b35098895422014d53ded71685160005e9764d

Observation 692f1dc0-35f6-49af-8546-cda699a05ef7 · outbound

This paper cites Flash diffusion: Acceler- ating any conditional diffusion model for few steps image generation.

Emu3.5: Native Multimodal Models are World Learners Flash diffusion: Acceler- ating any conditional diffusion model for few steps image generation

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T01:12:13.736854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:20aa9ea5bd2229259ce88ca65aa842861e24e9a9f86435fafd8a3508c5964816

Observation 7074b797-ab45-45af-9c31-e98a5a625764 · outbound

This paper cites OneIG-Bench: Omni-dimensional Nuanced Evaluation for Image Generation.

Emu3.5: Native Multimodal Models are World Learners OneIG-Bench: Omni-dimensional Nuanced Evaluation for Image Generation

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:12:13.675245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:a1605329da1b034d6aafe55161119a5dcda6af752fbd5ef29695178c8cf012b0

Observation 19919518-d516-461f-825c-fd9d6b083d99 · outbound

This paper cites Conceptual 12m: Pushing web- scale image-text pre-training to recognize long-tail visual concepts.

Emu3.5: Native Multimodal Models are World Learners Conceptual 12m: Pushing web- scale image-text pre-training to recognize long-tail visual concepts

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T01:12:13.740368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:8d9fd827fdf767645b8ea4228ace6c9eff30244cf64acf2d7c78afb4f8bd3e31

Observation 23603f50-c379-484c-ab6c-d34654ecbce7 · outbound

This paper cites Interleaved Scene Graphs for Interleaved Text-and-Image Generation Assessment.

Emu3.5: Native Multimodal Models are World Learners Interleaved Scene Graphs for Interleaved Text-and-Image Generation Assessment

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:12:13.497603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:b1947ef75e3e48df2e68753b448c7299af86b6ee663da857ebb903994221aab2

Observation cdb517c1-64f1-4e75-b036-5b9d17c4ff32 · outbound

This paper cites BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset.

Emu3.5: Native Multimodal Models are World Learners BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-18T01:12:13.520072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:319f59151aa2b66803009d015dde18f5b19abd916634649cddef637111c962b5

Observation 5aa384f7-4c39-4008-bbee-35e3365c1c77 · outbound

This paper cites ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation.

Emu3.5: Native Multimodal Models are World Learners ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:12:13.530814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:b8ee7dbe4b19165d1012b6480c0fae140a80958126430715e32a6307db6c9027

Observation 45131fa5-f0a2-4284-ada6-c5bd38d6af30 · outbound

This paper cites MultiRef: Controllable Image Generation with Multiple Visual References.

Emu3.5: Native Multimodal Models are World Learners MultiRef: Controllable Image Generation with Multiple Visual References

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:12:13.556070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:ffeb3d6e6fdaa12aac328415d9e3d95489721cb40112f879a8b96a9702dbe322

Observation e11ce4b4-d839-4d1d-9251-13d47c24cab1 · outbound

This paper cites PosterCraft: Rethinking High-Quality Aesthetic Poster Generation in a Unified Framework.

Emu3.5: Native Multimodal Models are World Learners PosterCraft: Rethinking High-Quality Aesthetic Poster Generation in a Unified Framework

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:12:13.611804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:17f9d84be56f0ddb48e108032d7d22212a26fe135d75a6a77a5dd688a132e56b

Observation d1993717-6c22-45a7-8da8-70a31188512b · outbound

This paper cites Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling.

Emu3.5: Native Multimodal Models are World Learners Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-18T01:12:13.626429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:fad8289380be8494850f8accce9f1ce93cf82a4a76c321b8ca03e9da50738649

Observation aa97f3af-a802-4178-a1b1-fa0b2b49ad38 · outbound

This paper cites Bagel: a web- based bacteriocin genome mining tool.Nucleic acids research, 34(suppl_2):W273–W279.

Emu3.5: Native Multimodal Models are World Learners Bagel: a web- based bacteriocin genome mining tool.Nucleic acids research, 34(suppl_2):W273–W279

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T01:12:13.743456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:b29895c1ae455cd8fe5643fbaf62b61c0d8a82454ff5a4f106fce0654f3a80af

Observation 108c810a-5ce7-4e9b-9833-5c38a2b2e838 · outbound

This paper cites insightface.https://github.com/deepinsight/insightface.

Emu3.5: Native Multimodal Models are World Learners insightface.https://github.com/deepinsight/insightface

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T01:12:13.746308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:c64d576d3e7e47c11d97ae7216be5574a4b5efcf1fc1cd309115b5817bf38cea

Observation 108fac73-3b76-4455-9a43-09c042f7fe64 · outbound

This paper cites Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning.

Emu3.5: Native Multimodal Models are World Learners Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T01:12:13.749078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:b877f46631babe7fffeacfb31d57c3e094ea642b0a5f6f83a86a604ccaa5ae69

Observation 78e84d4c-532e-4ba0-ab53-b35db87139bb · outbound

This paper cites Scaling vision transformers to 22 billion parameters.

Emu3.5: Native Multimodal Models are World Learners Scaling vision transformers to 22 billion parameters

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T01:12:13.752334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:0efe2776a1d3752dc43f774bfc6334baf14d7f36f64b4a1e4f071da2741cb6b4

Observation 11fc43eb-8bd3-43e4-b5f8-463c8ad640ae · outbound

This paper cites Uniform discrete diffusion with metric path for video generation.

Emu3.5: Native Multimodal Models are World Learners Uniform discrete diffusion with metric path for video generation

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:12:13.471400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:ba7edfcc62de2cf7147e74812e97ac0b69e8873b7f540b5ab2c80f786146e2e4

Observation 5424ae77-2b00-428c-9543-9bf10ead96bb · outbound

This paper cites Retinaface: Single-shot multi-level face localisation in the wild.

Emu3.5: Native Multimodal Models are World Learners Retinaface: Single-shot multi-level face localisation in the wild

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T01:12:13.755853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:2ceae8aa65e1982b5f417d8ffe7a6bd8215729aa6b2e44b27f7e09ac3ef3b57f

Observation a2ee46ab-e5ee-46b6-b264-8e342353d1da · outbound

This paper cites Textcrafter: Accurately rendering multiple texts in complex visual scenes.

Emu3.5: Native Multimodal Models are World Learners Textcrafter: Accurately rendering multiple texts in complex visual scenes

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:12:13.490601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:2e0a78e1f81ffdc36c4e5270d2821d5cef917bd6408074a2e137f1addb8b253d

Observation cd3675af-ace2-49ff-b443-1acb0bd85838 · outbound

This paper cites Scaling rectified flow transformers for high-resolution image synthesis.

Emu3.5: Native Multimodal Models are World Learners Scaling rectified flow transformers for high-resolution image synthesis

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T01:12:13.758767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:c3f83310f11f570a0095b8b827cbdf2de226d3a7c55ef17dbaa4d547cc46005c

Observation 0c832447-bb02-4501-a2cc-82b42bdc220d · outbound

This paper cites Taming transformers for high-resolution image synthesis.

Emu3.5: Native Multimodal Models are World Learners Taming transformers for high-resolution image synthesis

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T01:12:13.762096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:c6c84ff031407db41064bdb061cc8bf371a49c47a95edabd72bf4c09c49365e2

Observation c384ecdd-b071-44c0-b1ec-afd351e22510 · outbound

This paper cites Datacomp: In search of the next generation of multimodal datasets.Advances in Neural Information Processing Systems, 36:27092– 27112.

Emu3.5: Native Multimodal Models are World Learners Datacomp: In search of the next generation of multimodal datasets.Advances in Neural Information Processing Systems, 36:27092– 27112

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T01:12:13.765390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:2277fdc00023ae68a2a600589b73512785851f3454d1c2779e0f653e688671bb

Observation 2f61a95d-cb7e-49ee-9db5-03930dc3cb3b · outbound

This paper cites Seedream 3.0 Technical Report.

Emu3.5: Native Multimodal Models are World Learners Seedream 3.0 Technical Report

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-18T01:12:13.541450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:2a85f01e986218a0b684d0eb87ce2932c8e64c2f13c7ae607277ef122d535194

Observation 891cca51-81e1-4132-a405-5057cb9a82b6 · outbound

This paper cites Discrete flow matching.Advances in Neural Information Processing Systems, 37:133345–133385.

Emu3.5: Native Multimodal Models are World Learners Discrete flow matching.Advances in Neural Information Processing Systems, 37:133345–133385

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T01:12:13.768172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:a2f66ed6f1f6a37f475c75af0eae4ba8163f7c45700911ebbdc12aba2ce46337

Observation 1006dae3-4b1b-47cf-856a-808dbe738fe7 · outbound

This paper cites SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation.

Emu3.5: Native Multimodal Models are World Learners SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-18T01:12:13.566047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:1371381b504be53db6c3570fc0a32dd1ad8b9c5e3f10c21a3b6a73bfdcea072d

Observation a0bf2467-64fc-4bdc-a16e-27a87faa2575 · outbound

This paper cites X-Omni: Reinforcement Learning Makes Discrete Autoregressive Image Generative Models Great Again.

Emu3.5: Native Multimodal Models are World Learners X-Omni: Reinforcement Learning Makes Discrete Autoregressive Image Generative Models Great Again

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:12:13.587424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:37eabb4fd426c2ee50156cf885097caab90f96c6ac6117ab9eff05440d6676d4

Observation a573a1be-f126-48ee-9657-904ff3c374f9 · outbound

This paper cites Geneval: An object-focused framework for evaluating text-to-image alignment.Advances in Neural Information Processing Systems, 36.

Emu3.5: Native Multimodal Models are World Learners Geneval: An object-focused framework for evaluating text-to-image alignment.Advances in Neural Information Processing Systems, 36

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T01:12:13.771449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:30bd26fcaed2fa3746edf38f556b6805aec22f9edcf82c6b1f25a3b4e295a36c

Observation f7d3aedd-eca1-43c0-8745-e8f2d4f75e62 · outbound

This paper cites Gemini 2.0 flash.https://developers.googleblog.com/en/ experiment-with-gemini-20-flash-native-image-generation.

Emu3.5: Native Multimodal Models are World Learners Gemini 2.0 flash.https://developers.googleblog.com/en/ experiment-with-gemini-20-flash-native-image-generation

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T01:12:13.774805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:c0bf4055eb6395ae8a51ae67a22aadc2898b985b7b9a49150ca7080ac33630fe

Observation f2cd30cf-4bad-447a-8c0a-855d26c1d5c5 · outbound

This paper cites Imagen 3.

Emu3.5: Native Multimodal Models are World Learners Imagen 3

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T01:12:13.777738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:791a0224058c904216f4a0869c78f9c599c158c1570305b0d4cf049409a92cbb

Observation 0f118b8a-a3f0-449e-9018-9d6034243968 · outbound

This paper cites Imagen 4.

Emu3.5: Native Multimodal Models are World Learners Imagen 4

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T01:12:13.780361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:40370366fe42318e12567e266c4f51e4bb3b165e1e484a962418594551d1cc72

Observation 933d34f6-4ffb-4fd8-ae93-ffb88ebe1745 · outbound

This paper cites Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data.

Emu3.5: Native Multimodal Models are World Learners Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:12:13.653172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:2bb4ad410051e2d418988a28a74c355a32417a726c68ec62e0f5ae8cb6b1e654

Observation 2c3082ae-d2ce-416b-a07b-17141474286e · outbound

This paper cites Infinity: Scaling bitwise autoregressive modeling for high-resolution image synthesis.

Emu3.5: Native Multimodal Models are World Learners Infinity: Scaling bitwise autoregressive modeling for high-resolution image synthesis

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T01:12:13.783517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:77684ed25ae132786e6300b428cc4dcf9c3c88a8253d5d9c1a83fcc75e9d8a3c

Observation 59721299-0f79-4859-ba53-888479184049 · outbound

This paper cites Measuring colorfulness in natural images.

Emu3.5: Native Multimodal Models are World Learners Measuring colorfulness in natural images

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T01:12:13.786783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:870c618f06648d5f9e280bd743049229db56cc0aae7f94d2a45c4f4ec665d71c

Observation 2f28fe49-1ccd-4cc7-a525-605679157759 · outbound

This paper cites ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment.

Emu3.5: Native Multimodal Models are World Learners ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-05-18T01:12:13.682211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:d4d0b1412727ffb9627c78a61fce1f32fbd4d7df8376476f727e41b0dd73870c

Observation 6d142b67-4384-4967-ab09-8a0121a0e674 · outbound

This paper cites Image-to-image translation with conditional adversarial networks.

Emu3.5: Native Multimodal Models are World Learners Image-to-image translation with conditional adversarial networks

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T01:12:13.789672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:cb73bd290b14269ce9b7ca787888733ee5e0a1f135a1b9ae4b7b2aa56e37925a

Observation 19bcfd07-38fd-4444-89d6-fac2afc1f9a2 · outbound

This paper cites Lego-edit: A general image editing framework with model-level bricks and mllm builder.

Emu3.5: Native Multimodal Models are World Learners Lego-edit: A general image editing framework with model-level bricks and mllm builder

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:12:13.475603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:471badc284541c2e021c49cdb4269a777362e343c3bc8b1dd199387bf71afe12

Observation ce6bc978-0ffb-4818-844a-414e6ead6f70 · outbound

This paper cites InfiniteYou: Flexible Photo Recrafting While Preserving Your Identity.

Emu3.5: Native Multimodal Models are World Learners InfiniteYou: Flexible Photo Recrafting While Preserving Your Identity

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:12:13.479387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:2ac05c07c20dec76446bc354cae32f8617dd759b0c46f464bd2a19e282c069f3

Observation 1ac987de-fa1e-4ad1-9375-c2cf1bc9ad6f · outbound

This paper cites Musiq: Multi-scale image quality transformer.

Emu3.5: Native Multimodal Models are World Learners Musiq: Multi-scale image quality transformer

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T01:12:13.792561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:66ffc39024c7eba39addcabbfa4c94d73cd26f119149f508cb91662293d70f4e

Observation 622f510a-35ac-42ce-87a3-7b71bed633cc · outbound

This paper cites an unresolved cited work.

Emu3.5: Native Multimodal Models are World Learners Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-05-18T01:12:13.795238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:5510f145798df1665cc8b061f62c6a303e5d346c7ef42bd0afb2e4394e6ab8d9

Observation 62e3ca79-76d4-41bd-8f5e-9e4373267ef0 · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

Emu3.5: Native Multimodal Models are World Learners Gonzalez, Hao Zhang, and Ion Stoica

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T01:12:13.798118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:c639f39fecc3b1c6615927ffc374da291ce6db5ce8c3a4843e35d681c6fad676

Observation 088f4e04-7532-45d3-9ac1-529fded71c20 · outbound

This paper cites Flux.https://github.com/black-forest-labs/flux.

Emu3.5: Native Multimodal Models are World Learners Flux.https://github.com/black-forest-labs/flux

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T01:12:13.800832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:266f51ae2292b0db17746028e4c5169ef4b57d443d53f63f7b236345da10aca3

Observation 39c2e95b-f16d-41ca-a164-b12167a709ce · outbound

This paper cites FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space.

Emu3.5: Native Multimodal Models are World Learners FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-05-18T01:12:13.527168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:8679892a47f4db141d4fed2fa206df7636aacc8db648e5588482f458581c68ea

Observation 954be859-f10a-466a-b4d2-1f9796993024 · outbound

This paper cites Grounding image matching in 3d with mast3r.

Emu3.5: Native Multimodal Models are World Learners Grounding image matching in 3d with mast3r

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T01:12:13.804599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:d7c551cca3e1c46146ca3716428771771b5ce3244e6c135150dae7094d9f9e5a

Observation b941330f-2654-44f9-ae65-3526696db315 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Emu3.5: Native Multimodal Models are World Learners LLaVA-OneVision: Easy Visual Task Transfer

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-05-18T01:12:13.537985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:3196e7ed7d6d7eaa44fe218d947e37b8020fdbcf5358a8df195d6a3a5b422879

Observation 10165464-9ec3-435a-9dfa-43bb34e3e5fe · outbound

This paper cites Infinity instruct: Scaling instruction selection and synthesis to enhance language models.

Emu3.5: Native Multimodal Models are World Learners Infinity instruct: Scaling instruction selection and synthesis to enhance language models

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T01:12:13.807953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:fee4cb4a5e248b0b057f62114705176d3b248e52deb89a51fa48f839bdefd47f

Observation 1ddd2c17-163f-4151-a9a9-3f8db59b066e · outbound

This paper cites Sekai: A video dataset towards world exploration.

Emu3.5: Native Multimodal Models are World Learners Sekai: A video dataset towards world exploration

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:12:13.549220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:98383f4e2c80697695f288ec7b3be1f48947023cb9c77a7d84157b90711e70d5

Observation 8e9c6323-696b-4644-a8d2-a3115c9915aa · outbound

This paper cites UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation.

Emu3.5: Native Multimodal Models are World Learners UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-05-18T01:12:13.552554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:785e27eb7e39bec34357884319bc25d0e5e77144848eb7f2dc490f3e13002dd8

Observation 8d00c863-c005-42f1-8086-6686d7dabc88 · outbound

This paper cites Focal loss for dense object detection.

Emu3.5: Native Multimodal Models are World Learners Focal loss for dense object detection

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T01:12:13.810928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:cae085cc0425677d0b3039427636c8d454e7b042271a345067821f86001b8f85

Observation 118e3dbe-cb2d-437b-979c-9ee26854c37e · outbound

This paper cites Flow Matching for Generative Modeling.

Emu3.5: Native Multimodal Models are World Learners Flow Matching for Generative Modeling

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-05-18T01:12:13.559401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:aec600ce338fa34b3c625fbfa6de8555058fe2698922544525268329eb5457ff

Observation 4fe847ad-4b8e-4dd8-bfdf-12fd48fd0cba · outbound

This paper cites Improved baselines with visual instruction tuning.

Emu3.5: Native Multimodal Models are World Learners Improved baselines with visual instruction tuning

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T01:12:13.813942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:8ac44a6e4561c654622613ae21d77b0095fb6b5acc80b63b4f880b7619fbe4ec

Observation 386bc1bb-2611-4e77-889e-0fdac997a5ce · outbound

This paper cites Step1X-Edit: A Practical Framework for General Image Editing.

Emu3.5: Native Multimodal Models are World Learners Step1X-Edit: A Practical Framework for General Image Editing

Reference 58

Resolution
verified exact
local_arxiv, observed 2026-05-18T01:12:13.576502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:78c45d5300a005b8e8ebbf8d82547a04523a21f80cdc6ad90598559844e5f18b

Observation 089739e2-63ab-436a-b08e-e4c46c79e9af · outbound

This paper cites Dpm-solver: A fast ode solver for diffusion probabilistic model sampling in around 10 steps.Advances in neural information processing systems, 35:5775–5787.

Emu3.5: Native Multimodal Models are World Learners Dpm-solver: A fast ode solver for diffusion probabilistic model sampling in around 10 steps.Advances in neural information processing systems, 35:5775–5787

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T01:12:13.817237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:974cd84bc92ec975b855ebc8ccf4bc3adef4a59a4e0947829270a31adba40073

Observation 5b1bf519-4281-4e10-b7fa-ecf85beb41a6 · outbound

This paper cites Open-MAGVIT2: An Open-Source Project Toward Democratizing Auto-regressive Visual Generation.

Emu3.5: Native Multimodal Models are World Learners Open-MAGVIT2: An Open-Source Project Toward Democratizing Auto-regressive Visual Generation

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:12:13.590977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:cebc27f6c3779b003658c362a6f7b537a10c39dc4eca2aeb3515645b4e68e066

Observation 9cae939b-c162-4840-84b3-a6cf7cffaa01 · outbound

This paper cites Sit: Exploring flow and diffusion-based generative models with scalable interpolant transform- ers.

Emu3.5: Native Multimodal Models are World Learners Sit: Exploring flow and diffusion-based generative models with scalable interpolant transform- ers

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T01:12:13.822757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:de68ffea453c1d3802d35a1a98a2b12c1f5273609acb9c448733326f3d2244fa

Observation ef8df658-3626-4a2b-8800-6ba11319e378 · outbound

This paper cites Midjourney.

Emu3.5: Native Multimodal Models are World Learners Midjourney

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T01:12:13.825443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:e3d8b6bba279005f56341ef5d59ec1591aca3af89da0c349cfe688e584795e55

Observation 3f0db7fb-514b-49df-adaa-bc3f883a4e7e · outbound

This paper cites Gpt-4o.https://openai.com/index/introducing-4o-image-generation.

Emu3.5: Native Multimodal Models are World Learners Gpt-4o.https://openai.com/index/introducing-4o-image-generation

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T01:12:13.828223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:a126e12a7927e6cf33fceb57a7b040db421a789f1f9fbc0251f423a4f80f76cc

Observation 818aa069-da24-46d4-a6b6-0b8dc183df5d · outbound

This paper cites Image generation API.https://openai.com/index/image-generation-api/.

Emu3.5: Native Multimodal Models are World Learners Image generation API.https://openai.com/index/image-generation-api/

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T01:12:13.831098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:b9cbdc4be10aefd2cb43296d78d776acc89a17a520eed778d222db9b47d9e274

Observation 27d1099d-0c6e-468e-afea-399088f35863 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

Emu3.5: Native Multimodal Models are World Learners DINOv2: Learning Robust Visual Features without Supervision

Reference 65

Resolution
verified exact
local_arxiv, observed 2026-05-18T01:12:13.640216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:1ae4859102d793cde137a9e2cc11900741661b0ce8aeadd5c4ed3aafd835313e

Observation 8da3dbca-297d-467a-897c-74210863a038 · outbound

This paper cites Open x-embodiment: Robotic 40 learning datasets and rt-x models: Open x-embodiment collaboration 0.

Emu3.5: Native Multimodal Models are World Learners Open x-embodiment: Robotic 40 learning datasets and rt-x models: Open x-embodiment collaboration 0

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T01:12:13.834011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:f9132ae62cce4f18ddea000d49faa47390b58d312751aff5be9b2b523a795035

Observation 0cf2164b-b9b0-46ef-8a1a-53a5b7a0d03a · outbound

This paper cites Journeydb: A benchmark for generative image understanding.

Emu3.5: Native Multimodal Models are World Learners Journeydb: A benchmark for generative image understanding

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T01:12:13.836812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:ac99550e8196d929b27dc28a2d07de8d2bef762fcacdcd67de69b93974504626

Observation be314694-5e33-48de-a68d-eded2df1b19a · outbound

This paper cites ICE-Bench: A Unified and Comprehensive Benchmark for Image Creating and Editing.

Emu3.5: Native Multimodal Models are World Learners ICE-Bench: A Unified and Comprehensive Benchmark for Image Creating and Editing

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:12:13.660625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:c7afb05bc6deadde92cd3a3d2c5fcf22769e2d52d018764f93237d58561e84dd

Observation 5d59952a-bf9e-46c5-8b02-12ff9887d1a7 · outbound

This paper cites Scalable Diffusion Models with Transformers.

Emu3.5: Native Multimodal Models are World Learners Scalable Diffusion Models with Transformers

Reference 69

Resolution
verified exact
local_arxiv, observed 2026-05-18T01:12:13.664255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:30f840054325ca7634f8d2eb75b0e7d49956b2abcf1f52197e422b7a503dbf74

Observation b2edb98c-1047-4507-ae84-f427a3d890cb · outbound

This paper cites Tokenflow: Unified image tokenizer for multimodal understanding and generation.

Emu3.5: Native Multimodal Models are World Learners Tokenflow: Unified image tokenizer for multimodal understanding and generation

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T01:12:13.840924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:ca22a285691df837ba30b94fa831c773e35bcd94213ec7d4ee02dff6494f3ae8

Observation 3d1e6416-407b-452f-be09-65236acf5898 · outbound

This paper cites Robust speech recognition via large-scale weak supervision.

Emu3.5: Native Multimodal Models are World Learners Robust speech recognition via large-scale weak supervision

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T01:12:13.844425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:e71d0b620a44eb2243d9dcd15aac56a1ff06ede969f420755481447a1bce502c

Observation cafbb9c9-9ab5-4935-a2bc-47a3d2800170 · outbound

This paper cites Recraft.https://www.recraft.ai/.

Emu3.5: Native Multimodal Models are World Learners Recraft.https://www.recraft.ai/

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T01:12:13.847335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:8dc34498108df3456d71984d312fdd1d004c8b587da4b207d082bff6e65db16d

Observation cf75fe7b-e3b7-4854-9f1b-a45e129e40af · outbound

This paper cites Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation.

Emu3.5: Native Multimodal Models are World Learners Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T01:12:13.850376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:d45d5cdd8b7ea60cdf95d70211b66a40651983e3dd9e8301a8b6dceb1fd572ee

Observation b6cc72de-785a-4288-88b9-18fbfb93c248 · outbound

This paper cites Imagenet large scale visual recognition challenge.International journal of computer vision, 115(3):211–252.

Emu3.5: Native Multimodal Models are World Learners Imagenet large scale visual recognition challenge.International journal of computer vision, 115(3):211–252

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T01:12:13.856479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:b1a79799197e13a284df52a8e019f67cb6291fb60dccfa180a820f5794ca3041

Observation e6bfb617-6f55-4f8f-864f-092c2fcace0a · outbound

This paper cites Laion-5b: An open large-scale dataset for training next generation image-text models.Advances in neural information processing systems, 35:25278–25294.

Emu3.5: Native Multimodal Models are World Learners Laion-5b: An open large-scale dataset for training next generation image-text models.Advances in neural information processing systems, 35:25278–25294

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T01:12:13.859786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:7306ea6c87aa24025ef368c0b2b5aea3d0dd1b5ee57cff4baab204802e8b7a56

Observation 568a6c2d-a24e-466b-9b2c-595419b1dd68 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Emu3.5: Native Multimodal Models are World Learners DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 76

Resolution
verified exact
local_arxiv, observed 2026-05-18T01:12:13.486509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:5057d1091d497d12decdb7015f60f2c420c90308b8bdee057518c992adb3e504

Observation a77e738e-23ee-424b-84aa-9478fe595362 · outbound

This paper cites Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning.

Emu3.5: Native Multimodal Models are World Learners Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T01:12:13.863114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:2f88a9758ceea475eb10d85bb075d53ba3565faf7d26ed7070235f896d405ed4

Observation c854674e-5a63-48a3-904f-c70247a1bd89 · outbound

This paper cites GLU Variants Improve Transformer.

Emu3.5: Native Multimodal Models are World Learners GLU Variants Improve Transformer

Reference 78

Resolution
verified exact
local_arxiv, observed 2026-05-18T01:12:13.494054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:11583ad4c17e4c0ffd3e2db9a82c4cb423e49f400fc6535c4621688948c1b843

Observation ed0dd9fe-d810-4cb2-8e57-7ba100a0033d · outbound

This paper cites Storygpt-v: Large language models as consistent story vi- sualizers.

Emu3.5: Native Multimodal Models are World Learners Storygpt-v: Large language models as consistent story vi- sualizers

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T01:12:13.866281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:3a0798f771c2690d835f83733dfec3c6667fb0fec2b54ead57d145c22b925322

Observation eca8cc75-7f27-4bca-857c-4bd97203bca3 · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

Emu3.5: Native Multimodal Models are World Learners HybridFlow: A Flexible and Efficient RLHF Framework

Reference 80

Resolution
verified exact
local_arxiv, observed 2026-05-18T01:12:13.500996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:ff1edf34b781fb34c3c15b1a82b354f77b40aacd87f18e2e08bd940c50188e95

Observation ae32dae3-3d6b-46d0-a3d7-3906182ad645 · outbound

This paper cites Scalable Image Tokenization with Index Backpropagation Quantization.

Emu3.5: Native Multimodal Models are World Learners Scalable Image Tokenization with Index Backpropagation Quantization

Reference 81

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:12:13.504973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:0820526ca683b82cabc831fd0279e0c5bb9eba39f692671d6708d48693be3211

Observation c9ba1edb-2354-418b-9c38-f0f87cc495e1 · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

Emu3.5: Native Multimodal Models are World Learners Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 82

Resolution
verified exact
local_arxiv, observed 2026-05-18T01:12:13.508411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:793a76dcb93bf9e86143293c20a63d72c72b9c1cd9c34607fced8081fa803cd4

Observation 6970193c-06e6-4feb-84eb-96d084a1ab5f · outbound

This paper cites Denoising Diffusion Implicit Models.

Emu3.5: Native Multimodal Models are World Learners Denoising Diffusion Implicit Models

Reference 83

Resolution
verified exact
local_arxiv, observed 2026-05-18T01:12:13.512775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:aff558c8a326c2f7c99b3c4851d911f0488c74925de3fa94907146146f7dff4f

Observation 9b4fce4f-03f0-4b94-adbd-7c319f45989b · outbound

This paper cites Nemotron-CC: Transforming Common Crawl into a Refined Long-Horizon Pretraining Dataset.

Emu3.5: Native Multimodal Models are World Learners Nemotron-CC: Transforming Common Crawl into a Refined Long-Horizon Pretraining Dataset

Reference 84

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:12:13.516657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:155b068000536374560c8d0c6c38447c6f7c74fdc4a65d306caae836a9aeb62c

Observation 503ce59f-d8cd-47da-8603-3876000d39f8 · outbound

This paper cites Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063.

Emu3.5: Native Multimodal Models are World Learners Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T01:12:13.869591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:947dd8fcc84d892a4158a07043001702f65d1dc8227898e7f07a67b43aca4984

Observation bd0b1bc5-5322-4bde-a831-09c7f9c3801e · outbound

This paper cites Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation.

Emu3.5: Native Multimodal Models are World Learners Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation

Reference 86

Resolution
verified exact
local_arxiv, observed 2026-05-18T01:12:13.523756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:8c89991508ad652cbc2206fba5c6626d4663e05b34e4da217a0cec3ff21eaef3

Observation 2c7aae4d-ae16-4695-b11c-cf66d13c5a91 · outbound

This paper cites Generative multimodal models are in-context learn- ers.

Emu3.5: Native Multimodal Models are World Learners Generative multimodal models are in-context learn- ers

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T01:12:13.872513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:814c61fff8fe630841c39ea222877e7bbf8f22da826dd77bb70369f72479bc2c

Observation a74d0875-1b45-4f8f-a7d0-2e3e956c6f00 · outbound

This paper cites Emu: Generative pretraining in multimodality.

Emu3.5: Native Multimodal Models are World Learners Emu: Generative pretraining in multimodality

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T01:12:13.875376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:198099bc9c8ac8cb94a927b8bf2643c329a536f7df3dc7e0674e149d9785c048

Observation 2a9b32f0-c245-4166-b57e-bd7717d0dc84 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

Emu3.5: Native Multimodal Models are World Learners Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 89

Resolution
verified exact
local_arxiv, observed 2026-05-18T01:12:13.534612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:fe6d2ba974961e9f6b5926d658ae220dbc38165e1d58a703927d0cb63edd21df

Observation a5c21e30-15af-46d0-a959-379a48b053e4 · outbound

This paper cites FlagScale: A unified meta-framework enabling adaptive heterogeneous computing for the llm ecosystem.https://github.com/FlagOpen/FlagScale.

Emu3.5: Native Multimodal Models are World Learners FlagScale: A unified meta-framework enabling adaptive heterogeneous computing for the llm ecosystem.https://github.com/FlagOpen/FlagScale

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T01:12:13.892238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:db327588e98d1e357e8268c31a418d981981716a77155d31b452beb877588e52

Observation 0d60490a-40c8-4c67-9ee6-f7fa421402f5 · outbound

This paper cites Gemini 2.5 flash & gemini 2.5 flash image model card.https://storage.

Emu3.5: Native Multimodal Models are World Learners Gemini 2.5 flash & gemini 2.5 flash image model card.https://storage

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T01:12:13.895266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:390a89417467f8196b056e60672e8c4dd99dd25ec0e7c2e00d12a1e0d030cc7f

Observation cbc2613b-9e37-4c01-8bce-edd438dd9581 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

Emu3.5: Native Multimodal Models are World Learners Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 92

Resolution
verified exact
local_arxiv, observed 2026-05-18T01:12:13.545324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:3530dea3e8bcfe9f82b2e60d437c8d22eed99e33a5320ed7c2315a8ec3a71a41

Observation 4bd5cfe4-5161-4574-9f24-851b5a3fe7b3 · outbound

This paper cites Kolors 2.0.https://app.klingai.com/cn/.

Emu3.5: Native Multimodal Models are World Learners Kolors 2.0.https://app.klingai.com/cn/

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T01:12:13.897797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:b23068051e509ee58c82953df8adbaf9669190631a091cdbf180a2a7ceae85de

Observation 2afe1780-1b23-4a20-aeae-8031f00e81a9 · outbound

This paper cites Raft: Recurrent all-pairs field transforms for optical flow.

Emu3.5: Native Multimodal Models are World Learners Raft: Recurrent all-pairs field transforms for optical flow

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T01:12:13.900982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:a2171874655e2558537fe83f449dca965a502a4c53a0091f08ae377ec2bd0ae4

Observation 92e309a8-d75d-46a5-b963-7013954c790d · outbound

This paper cites Training-free consistent text-to-image generation.ACM Transactions on Graphics (TOG), 43(4):1– 18.

Emu3.5: Native Multimodal Models are World Learners Training-free consistent text-to-image generation.ACM Transactions on Graphics (TOG), 43(4):1– 18

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T01:12:13.904012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:54f6d3306753fcf797305050307f2f278784e312be01d4581d24da842b4cb1eb

Observation ba33742d-c700-4b24-8781-215ec2284199 · outbound

This paper cites Visual autoregressive model- ing: Scalable image generation via next-scale prediction.Advances in neural information processing systems, 37:84839–84865.

Emu3.5: Native Multimodal Models are World Learners Visual autoregressive model- ing: Scalable image generation via next-scale prediction.Advances in neural information processing systems, 37:84839–84865

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T01:12:13.907326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:eaf9d79e7734e7bb02b84c8e356ac6487a1b7508e251be712408b237eefdd2a8

Observation 63a4f167-f0c3-4913-8b06-110ecc88c93c · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Emu3.5: Native Multimodal Models are World Learners Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 97

Resolution
verified exact
local_arxiv, observed 2026-05-18T01:12:13.562571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:dd4504653f07013f07f6b9ea3678f2b44676010cff8c570f2987c6605e3f5589

Observation 0878c272-1884-463f-96a2-19246211155a · outbound

This paper cites Wan: Open and advanced large-scale video generative models.

Emu3.5: Native Multimodal Models are World Learners Wan: Open and advanced large-scale video generative models

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T01:12:13.910441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:4468ec0c22a98d375e8922b7d56f49a359f13ca9aed35ff9b952996d1f11f292

Observation 6ff02927-3e71-458c-97da-be03eea707d2 · outbound

This paper cites Textatlas5m: A large-scale dataset for dense text image generation.

Emu3.5: Native Multimodal Models are World Learners Textatlas5m: A large-scale dataset for dense text image generation

Reference 99

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:12:13.569579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:2f8a9563386a40465ea76b5eb8aaeca449c50aa2e95b196184a88309c28f7839

Observation 52937886-7a67-4471-b533-8e1c5411bc4d · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Emu3.5: Native Multimodal Models are World Learners Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 100

Resolution
verified exact
local_arxiv, observed 2026-05-18T01:12:13.573083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:8aa8979eb616ace66999d53c479154f437453c300e2ab2188e5d7210ce6ad50f

Pith citing papers

Observation f6d047db-2a61-435b-af07-280e9e5a2503 · inbound

Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm cites this paper.

Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm Emu3.5: Native Multimodal Models are World Learners

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:12:13.931083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T00:54:24.641649Z digest=sha256:3e1da1cc6cddd3d8b4f32b2a5c612caf77d9d08a32cddce0a8182266702d58b6

Observation e85d8781-85a1-40a8-981f-bcbd8cbbbdc8 · inbound

CG-MLLM: Captioning and Generating 3D content via Multi-modal Large Language Models cites this paper.

CG-MLLM: Captioning and Generating 3D content via Multi-modal Large Language Models Emu3.5: Native Multimodal Models are World Learners

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-05-21T14:50:14.634653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-21T14:48:21.787919Z digest=sha256:d2ea315a7818a5deda1e4d2f9c993a487a1048a134c0bd15e33ceeb425fcfa2d

Observation 7b9833f7-cfc6-4a0f-a2bd-19d1152d0294 · inbound

MentisOculi: Revealing the Limits of Reasoning with Mental Imagery cites this paper.

MentisOculi: Revealing the Limits of Reasoning with Mental Imagery Emu3.5: Native Multimodal Models are World Learners

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T05:23:31.119886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:23:31.119886Z digest=sha256:93421d7138e1ee6f637eeb17814ab2d2d7f14f748e58e1cd8b74cd455aed448f

Observation 39446122-d597-4d44-8273-4b70f81fc342 · inbound

LLaMo: Scaling Pretrained Language Models for Unified Motion Understanding and Generation with Continuous Autoregressive Tokens cites this paper.

LLaMo: Scaling Pretrained Language Models for Unified Motion Understanding and Generation with Continuous Autoregressive Tokens Emu3.5: Native Multimodal Models are World Learners

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:12:13.931083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-16T05:01:11.880003Z digest=sha256:d6bc222e05eaa4bcafd76f212c047e841aa9d27b6d1a6df2798810bb3fabadc0

Observation 3a721f1f-b528-43cd-9b49-b910113cc8c1 · inbound

VLANeXt: Recipes for Building Strong VLA Models cites this paper.

VLANeXt: Recipes for Building Strong VLA Models Emu3.5: Native Multimodal Models are World Learners

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:00:09.819893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-21T12:58:30.777235Z digest=sha256:0edcee423cc0bda8732b679673b9bf719f3bb568fdb371dfa8be72c6170f924f

Observation 35772a41-bbc4-4737-b365-352d08d46ff9 · inbound

WeEdit: A Dataset, Benchmark and Glyph-Guided Framework for Text-centric Image Editing cites this paper.

WeEdit: A Dataset, Benchmark and Glyph-Guided Framework for Text-centric Image Editing Emu3.5: Native Multimodal Models are World Learners

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T18:25:52.604284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:25:52.604284Z digest=sha256:2f854b4fbe108496c846ed194c4829453026d2e820a0b0b069a6af7ed1bf96a3

Observation 900bf7e2-3b1b-486b-aa15-3b9281c21bf4 · inbound

TRACE: High-Fidelity 3D Scene Editing via Tangible Reconstruction and Geometry-Aligned Contextual Video Masking cites this paper.

TRACE: High-Fidelity 3D Scene Editing via Tangible Reconstruction and Geometry-Aligned Contextual Video Masking Emu3.5: Native Multimodal Models are World Learners

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T02:30:57.284292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:30:57.284292Z digest=sha256:68367b8d45c7e537af3d866e7b1c8d2b9a046fbb74f3cd9b0cc34b718d337b9c

Observation e272cddb-bc6a-4a9a-a8dc-f77c87098087 · inbound

OpenWorldLib: A Unified Codebase and Definition of Advanced World Models cites this paper.

OpenWorldLib: A Unified Codebase and Definition of Advanced World Models Emu3.5: Native Multimodal Models are World Learners

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:12:13.931083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T19:36:42.100191Z digest=sha256:537e6d427f1dd1822cceb9fa365c7527b64ad0da61df769cfc8ace5cd6f7c717

Observation 3b0d5a40-e62d-4bcd-a180-c53f7d1e3158 · inbound

OpenWorldLib: A Unified Codebase and Definition of Advanced World Models cites this paper.

OpenWorldLib: A Unified Codebase and Definition of Advanced World Models Emu3.5: Native Multimodal Models are World Learners

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-13T09:42:23.808691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T09:42:23.808691Z digest=sha256:776d47f7404abf9aea8bc5cabeb913388b4994e221d560e476db5d5311fbaa06

Observation fd8475af-1a91-4dc4-829e-7d65da4558c1 · inbound

Learning Preference-Based Objectives from Clinical Narratives for Dynamic Sepsis Treatment cites this paper.

Learning Preference-Based Objectives from Clinical Narratives for Dynamic Sepsis Treatment Emu3.5: Native Multimodal Models are World Learners

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-12T22:22:06.385856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:22:06.385856Z digest=sha256:75d3d7a93b8097f5a2ec75efa32acda93864242f1428b95c7c8bc348ea64d89f

Observation bd1eadb0-4071-4a9d-bdce-b13e5ebca823 · inbound

TorchUMM: A Unified Multimodal Model Codebase for Evaluation, Analysis, and Post-training cites this paper.

TorchUMM: A Unified Multimodal Model Codebase for Evaluation, Analysis, and Post-training Emu3.5: Native Multimodal Models are World Learners

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:12:13.931083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T15:32:01.551829Z digest=sha256:af37925028d7467812aa0a24c251e760dec1bd5b909633f21c23083e3ebb6fc0

Observation 5569b540-477f-47fc-b365-ba8de8784322 · inbound

TorchUMM: A Unified Multimodal Model Codebase for Evaluation, Analysis, and Post-training cites this paper.

TorchUMM: A Unified Multimodal Model Codebase for Evaluation, Analysis, and Post-training Emu3.5: Native Multimodal Models are World Learners

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-21T08:59:55.280611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-21T08:59:10.877437Z digest=sha256:472a154286861b2cd102015de11aae00facc5c9479859c4fb3ef7776519e7ee0

Observation 5550148e-c7ce-44ce-9847-34e2951046c7 · inbound

UniGeo: Unifying Geometric Guidance for Camera-Controllable Image Editing via Video Models cites this paper.

UniGeo: Unifying Geometric Guidance for Camera-Controllable Image Editing via Video Models Emu3.5: Native Multimodal Models are World Learners

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:12:13.931083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T06:03:06.920592Z digest=sha256:4114000f8eff6668f20bd479c26f08c28fef5c88b35b66fa3499e8c9d065eb02

Observation 0d443724-2970-41cf-83f2-5b1bd40d2526 · inbound

UniGeo: Unifying Geometric Guidance for Camera-Controllable Image Editing via Video Models cites this paper.

UniGeo: Unifying Geometric Guidance for Camera-Controllable Image Editing via Video Models Emu3.5: Native Multimodal Models are World Learners

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:12:13.931083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-12T02:52:58.232585Z digest=sha256:2ab64bddd100355dbfef77b88a4d61230783206bc43cd23f053640daa263f1d9

Observation e455f6a8-2414-4623-9b14-f094bd6ba615 · inbound

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation cites this paper.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation Emu3.5: Native Multimodal Models are World Learners

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:12:13.931083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T05:42:41.112158Z digest=sha256:361a16b6ec8508416ee5e9741d0d5e02a230ea4407f2f5199af04f3038f5f21e

Observation 19177e61-0370-4e00-b196-c64d4bda4c92 · inbound

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation cites this paper.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation Emu3.5: Native Multimodal Models are World Learners

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:12:13.931083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:705487b5cd9b66599f33b1a82d69b2d508f3c217c014d7aaa3f433b076f04de5

Observation 9a89ef6d-d376-4227-83b6-4641923932c2 · inbound

UDM-GRPO: Stable and Efficient Group Relative Policy Optimization for Uniform Discrete Diffusion Models cites this paper.

UDM-GRPO: Stable and Efficient Group Relative Policy Optimization for Uniform Discrete Diffusion Models Emu3.5: Native Multimodal Models are World Learners

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:12:13.931083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T05:32:54.235578Z digest=sha256:d7a85873e149c49042304cbbdb928555bc83ce5209213798c7a384c55a46794f

Observation 0a6cd5dd-dda3-41c9-b163-87feb64903e3 · inbound

UDM-GRPO: Stable and Efficient Group Relative Policy Optimization for Uniform Discrete Diffusion Models cites this paper.

UDM-GRPO: Stable and Efficient Group Relative Policy Optimization for Uniform Discrete Diffusion Models Emu3.5: Native Multimodal Models are World Learners

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-05T11:41:02.738836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-05T11:32:38.636335Z digest=sha256:5817d291bb7340c08837c4c832a0c5ec240f4fd21152319a7171be31709b38fa

Observation a0ee33fa-5819-4e27-96ce-84c4528d6fbe · inbound

Exploring Spatial Intelligence from a Generative Perspective cites this paper.

Exploring Spatial Intelligence from a Generative Perspective Emu3.5: Native Multimodal Models are World Learners

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:12:13.931083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T00:45:46.261005Z digest=sha256:880592ac1cea59db9c163ef3265496e235daffc6d047892013b3db000fb8cb15

Observation f2aacf6d-66d8-4a3b-9069-1c8944a7aa32 · inbound

Context Unrolling in Omni Models cites this paper.

Context Unrolling in Omni Models Emu3.5: Native Multimodal Models are World Learners

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:12:13.931083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-09T22:02:57.841111Z digest=sha256:956f0c866a51fe13528c2f7273cd9e560e21a3f2821a05c2ca46052a76a34520

Observation 2cf5a54f-e46d-4384-9f6d-6eea870b9d79 · inbound

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation cites this paper.

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation Emu3.5: Native Multimodal Models are World Learners

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:12:13.931083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-08T04:31:26.325118Z digest=sha256:10596382c8ac63fbeed8a2a68ec5cdc6d5c27b6953f76c51b5106f8aedf99f51

Observation 43c9d681-bb25-43b8-b687-61e13146f13c · inbound

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation cites this paper.

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation Emu3.5: Native Multimodal Models are World Learners

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-20T23:43:51.071493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T23:41:25.275207Z digest=sha256:e2d9de667a437accfc1341f5aede2c7661ba6a336205e2e3d8d171315f53a2af

Observation 82a09e16-f052-425c-ad12-0dc305fdaeb4 · inbound

Unified 4D World Action Modeling from Video Priors with Asynchronous Denoising cites this paper.

Unified 4D World Action Modeling from Video Priors with Asynchronous Denoising Emu3.5: Native Multimodal Models are World Learners

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:12:13.931083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-07T10:40:04.767657Z digest=sha256:ae56e853ead2153bc3591dc671c50a664223a3819226307e51e8cd34da83dfbb

Observation 580e768a-833a-4f20-bb57-70499ac369c6 · inbound

Unified 4D World Action Modeling from Video Priors with Asynchronous Denoising cites this paper.

Unified 4D World Action Modeling from Video Priors with Asynchronous Denoising Emu3.5: Native Multimodal Models are World Learners

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:12:13.931083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-08T03:20:04.435904Z digest=sha256:9437a31ba7f3b661c7ab27f5fe4448becab05e1a77e5a7e82bb587d98bd70c31

Observation 1475be14-4922-4694-ad74-01d9705f3368 · inbound

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation cites this paper.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Emu3.5: Native Multimodal Models are World Learners

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:12:13.931083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-08T17:57:08.606559Z digest=sha256:9426de87592b9070995649d6cfabee6a16a5fca5257660e65090ee1fbc2743ba

Observation 389f1f92-6c60-4f67-a9ce-bb8d79c483e7 · inbound

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation cites this paper.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Emu3.5: Native Multimodal Models are World Learners

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-21T08:19:52.730252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:f282c09edd7f1b80e40575c3e4a94bce371c86867271a88e9c717e5b79dfd1e2

Observation 0b36d08d-7b99-4d1e-a604-508c19f83d2f · inbound

MULTITEXTEDIT: Benchmarking Cross-Lingual Degradation in Text-in-Image Editing cites this paper.

MULTITEXTEDIT: Benchmarking Cross-Lingual Degradation in Text-in-Image Editing Emu3.5: Native Multimodal Models are World Learners

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:12:13.931083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-12T01:11:35.480399Z digest=sha256:1f80fdcf4ddd521dbe61ad2b4d855b32c6c96b42c55cfcdaa3c30b032bea7981

Observation fa20c607-f838-4b49-99d9-706f2225c779 · inbound

Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm cites this paper.

Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Emu3.5: Native Multimodal Models are World Learners

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:12:13.931083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-13T06:02:39.114660Z digest=sha256:0fead7ef65bad8195ec562630fc488e9f7efbd90cc4bd629ebddeb9a798c4647

Observation e33e2eea-d06b-4c15-ba1d-632fae2e7cf2 · inbound

Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm cites this paper.

Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Emu3.5: Native Multimodal Models are World Learners

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:05:46.685305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T22:18:55.933311Z digest=sha256:19ce0b9d16ebd160a305a1370eb0355c56b439ad07bfa7234d70549a23099bcd

Observation 15f8e895-b4ee-41ae-abcc-e9c546de505b · inbound

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture cites this paper.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Emu3.5: Native Multimodal Models are World Learners

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:12:13.931083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:ff90650cc0c40a7ddc57d80c69da1ec6611b258d3b8ba31209e38ec7a47938e2

Observation 79afed4e-be56-4a06-b432-c406a2b5fb51 · inbound

Edit-Compass & EditReward-Compass: A Unified Benchmark for Image Editing and Reward Modeling cites this paper.

Edit-Compass & EditReward-Compass: A Unified Benchmark for Image Editing and Reward Modeling Emu3.5: Native Multimodal Models are World Learners

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:12:13.931083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:22:20.464966Z digest=sha256:d755b1cd533a884b7f17813623fe120f1e9b82d753e8c41761a991da250afbea

Observation b7561dea-ddb6-4f83-bada-794f63a097c7 · inbound

InsightTok: Improving Text and Face Fidelity in Discrete Tokenization for Autoregressive Image Generation cites this paper.

InsightTok: Improving Text and Face Fidelity in Discrete Tokenization for Autoregressive Image Generation Emu3.5: Native Multimodal Models are World Learners

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:12:13.931083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-15T02:10:18.004724Z digest=sha256:9ba00a99424af388eddce447c2a1ee9a2a2066e440dde240ab8e564bec9ecfc2

Observation 773d1a32-1274-4cd2-bb7f-56a356d6077e · inbound

Sketch Then Paint: Hierarchical Reinforcement Learning for Diffusion Multi-Modal Large Language Models cites this paper.

Sketch Then Paint: Hierarchical Reinforcement Learning for Diffusion Multi-Modal Large Language Models Emu3.5: Native Multimodal Models are World Learners

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-19T21:22:48.576713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-19T21:18:03.005508Z digest=sha256:f0f3c5a5b1745e5b137673580a8ae06a9009e7d4b250be8e33de30eaf3ee131e

Observation 5d8a8294-0514-4bf0-aaba-bbad823299a8 · inbound

LatentUMM: Dual Latent Alignment for Unified Multimodal Models cites this paper.

LatentUMM: Dual Latent Alignment for Unified Multimodal Models Emu3.5: Native Multimodal Models are World Learners

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-20T12:43:17.378868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T12:39:28.058049Z digest=sha256:f300fc20b3f5d09ae03301b43de4ed76ae8be807942b014460b812df96a4ca1a

Observation 14d14ed3-5410-4461-8c61-f7459cdc3b32 · inbound

Lance: Unified Multimodal Modeling by Multi-Task Synergy cites this paper.

Lance: Unified Multimodal Modeling by Multi-Task Synergy Emu3.5: Native Multimodal Models are World Learners

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-20T11:48:14.949487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T11:46:52.658984Z digest=sha256:de1e3313c54f00d2d343cdd248e82e0eb0bbb23f058f6ce2e1694fc5ddd10746

Observation 9813bfcc-dc4c-435b-92d1-de6b9fce24e2 · inbound

Lance: Unified Multimodal Modeling by Multi-Task Synergy cites this paper.

Lance: Unified Multimodal Modeling by Multi-Task Synergy Emu3.5: Native Multimodal Models are World Learners

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:59:50.460568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:529bdfeb22ee6f3da5028696c637d6f0feeac6fc3ab96d8010b0dbe6fbea485d

Observation e880fb75-7176-44fb-bcbf-ac24bb537257 · inbound

TextSculptor: Training and Benchmarking Scene Text Editing cites this paper.

TextSculptor: Training and Benchmarking Scene Text Editing Emu3.5: Native Multimodal Models are World Learners

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-21T05:19:39.390736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-21T05:16:43.756525Z digest=sha256:71f6753e9e018160a2b76a63cc280f8d13385358c8c906a6b294449f1764da49

Observation cdc50307-8237-4bd7-b4ab-b319acbfe5f8 · inbound

Uni-Edit: Intelligent Editing Is A General Task For Unified Model Tuning cites this paper.

Uni-Edit: Intelligent Editing Is A General Task For Unified Model Tuning Emu3.5: Native Multimodal Models are World Learners

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-21T04:33:57.702602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-21T04:33:51.535737Z digest=sha256:5e791153c578ade959d28d4a5e1a6c39ae3c5830b9d72f513a3583974a080642

Observation aa556b57-17a9-4ca9-a115-0d8dcdca9832 · inbound

Uni-Edit: Intelligent Editing Is A General Task For Unified Model Tuning cites this paper.

Uni-Edit: Intelligent Editing Is A General Task For Unified Model Tuning Emu3.5: Native Multimodal Models are World Learners

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-25T05:45:24.103941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-25T05:43:38.328972Z digest=sha256:347949200e56de99180627d100036c0af12cc6620ba3259ab9607103fb375ecc

Observation cbbe4208-39bb-4712-98d5-1a512b26d723 · inbound

Bernini: Latent Semantic Planning for Video Diffusion cites this paper.

Bernini: Latent Semantic Planning for Video Diffusion Emu3.5: Native Multimodal Models are World Learners

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-22T06:41:10.530534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-22T06:39:47.124605Z digest=sha256:70b203203b2e9580abbdf242387c0a259f6a3fafea5e539cd64efe91a7f46755

Observation 140741fe-ff67-4980-b9e1-07e38a88ec69 · inbound

Guess the Unified Model: How Much Can We Recover from Generated Images? cites this paper.

Guess the Unified Model: How Much Can We Recover from Generated Images? Emu3.5: Native Multimodal Models are World Learners

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T11:54:38.156916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T11:53:36.715338Z digest=sha256:86c86a1ad7133258bc62a5db18006406cc0beb37453def75ab4461c0eb4b6f02

Observation a9c42e88-73f1-40ec-a5e0-8feb6ef22e79 · inbound

Toward Native Multimodal Modeling: A Roadmap cites this paper.

Toward Native Multimodal Modeling: A Roadmap Emu3.5: Native Multimodal Models are World Learners

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-06-29T23:04:02.010989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-29T22:58:38.610609Z digest=sha256:c8725435fb47e55516fcfcf95a929ff06721876b39997747523910cb031a1bf5

Observation 8a02ef79-98bd-4a09-9d19-104b173be02b · inbound

OmniVerifier-M1: Multimodal Meta-Verifier with Explicit Structured Recalibration cites this paper.

OmniVerifier-M1: Multimodal Meta-Verifier with Explicit Structured Recalibration Emu3.5: Native Multimodal Models are World Learners

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:03:26.008261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-29T13:01:08.529979Z digest=sha256:752b4b3d2aa024956b612b30085311d16725474191954a374932edf7ef6440eb

Observation c48ecdcb-6dd0-44f8-8d91-44428270f8f5 · inbound

GenClaw: Code-Driven Agentic Image Generation cites this paper.

GenClaw: Code-Driven Agentic Image Generation Emu3.5: Native Multimodal Models are World Learners

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-06-29T07:43:13.990845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-29T07:36:54.292448Z digest=sha256:293352de3f841aae0e843560eae87043f1cd7b4e1845270cd7dc20da07bf7682

Observation 771fcf27-c3a8-4744-b39e-0acfddc9fc4b · inbound

Echo-Memory: A Controlled Study of Memory in Action World Models cites this paper.

Echo-Memory: A Controlled Study of Memory in Action World Models Emu3.5: Native Multimodal Models are World Learners

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-03T00:57:29.767850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-27T16:58:37.552036Z digest=sha256:f2f396a41ed29ef3bf787fead895d305c8531b48be38f6f3650cb9a8d311b3ef

Observation ec2e823d-d79b-49a7-bf2f-91a6de4b514d · inbound

SVoT: State-aware Visualization-of-Thought for Spatial Reasoning via Reinforcement Learning cites this paper.

SVoT: State-aware Visualization-of-Thought for Spatial Reasoning via Reinforcement Learning Emu3.5: Native Multimodal Models are World Learners

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-07-03T10:27:56.940777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-27T09:59:02.899488Z digest=sha256:604d4d8d70fb06c9333c78bf3628f49a62eb40a9679a5fcf25c89b7c5b7b82e3

Observation 5c1f5286-4b7a-4218-935f-241361e74011 · inbound

InterleaveThinker: Reinforcing Agentic Interleaved Generation cites this paper.

InterleaveThinker: Reinforcing Agentic Interleaved Generation Emu3.5: Native Multimodal Models are World Learners

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-03T15:08:32.891341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-27T06:42:34.126336Z digest=sha256:65aee1387daa6455a2419df2d42943cf9905c638247b2c112cb21519db749001

Observation 93da6e2c-168e-4ab1-b890-5b0c353b7d09 · inbound

ThinkingVLA: Interleaved Vision and Language Reasoning for Robotic Manipulation cites this paper.

ThinkingVLA: Interleaved Vision and Language Reasoning for Robotic Manipulation Emu3.5: Native Multimodal Models are World Learners

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-03T21:08:58.669938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-27T00:51:05.494085Z digest=sha256:1ba62d00dea3afc96d360cecb91be16c94e05ab11af8737742a5bbd0a634c0ef

Observation 15e506a9-738a-4495-880a-c09e55965272 · inbound

Spotlight: Synergizing Seed Exploration and Spot GPUs for DiT RL Post-Training cites this paper.

Spotlight: Synergizing Seed Exploration and Spot GPUs for DiT RL Post-Training Emu3.5: Native Multimodal Models are World Learners

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T02:49:24.716057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-26T19:06:16.100762Z digest=sha256:4360af29a07fbfc674fec3ca83f6b171784833cd7ac222dcf301fa2edcd283a2

Observation 1e143c10-f618-42d0-89f9-0fa70d00660e · inbound

IV-CoT: Implicit Visual Chain-of-Thought for Structure-Aware Text-to-Image Generation cites this paper.

IV-CoT: Implicit Visual Chain-of-Thought for Structure-Aware Text-to-Image Generation Emu3.5: Native Multimodal Models are World Learners

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-07-04T16:39:58.262161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-26T00:19:49.071495Z digest=sha256:5477a370323e4e3efe59e0f0b6c9225056681b2da086e20b9a16ecc3d4e32b9b

Observation 2d629f4d-3d9b-4062-835c-089727b01795 · inbound

Nemotron-Labs-Diffusion-Image: Advancing Masked Discrete Diffusion for High-Resolution Image Synthesis cites this paper.

Nemotron-Labs-Diffusion-Image: Advancing Masked Discrete Diffusion for High-Resolution Image Synthesis Emu3.5: Native Multimodal Models are World Learners

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-06-30T08:14:26.581955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T06:07:55.600338Z digest=sha256:25fc89d87a3ba1919ecfe3855a5748598fcb24fd2ba5a3897430fefd8361c012

Observation afc89459-cb70-47f4-b347-1321ce527ceb · inbound

Nemotron-Labs-Diffusion-Image: Advancing Masked Discrete Diffusion for High-Resolution Image Synthesis cites this paper.

Nemotron-Labs-Diffusion-Image: Advancing Masked Discrete Diffusion for High-Resolution Image Synthesis Emu3.5: Native Multimodal Models are World Learners

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T09:39:33.426324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:39:33.426324Z digest=sha256:8da37049951c79a546bfed4460e2293687c728833698c6d300dbb886a0216bf9

Observation fc548f2a-8da7-4b01-ad31-a5bd8568a2e9 · inbound

Illuminating Unified Multimodal Model for Free-form Interleaved Text-Image Generation cites this paper.

Illuminating Unified Multimodal Model for Free-form Interleaved Text-Image Generation Emu3.5: Native Multimodal Models are World Learners

Reference 12

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T06:04:20.886608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T06:04:07.327934Z digest=sha256:bed5c1508f88e77c2d0eb72ab650d005749671c7bf5fc10b013f619b89a27b20

Observation c98223a1-3a5a-450e-be53-536016f92f0f · inbound

MemLearner: Learning to Query Context memory for Video World Models cites this paper.

MemLearner: Learning to Query Context memory for Video World Models Emu3.5: Native Multimodal Models are World Learners

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-01T10:25:41.921325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-01T05:30:56.140465Z digest=sha256:ddbb03c1038d94d74f7951db9f1060312aa5383facf1410974330c56dbf611bf

Observation e011d6fd-0809-420a-9206-5f96753abb7d · inbound

WorldBagel: Uncovering the Power of Unified Multimodal Models for Vision-Language-Action-World Modeling cites this paper.

WorldBagel: Uncovering the Power of Unified Multimodal Models for Vision-Language-Action-World Modeling Emu3.5: Native Multimodal Models are World Learners

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-12T02:19:35.281990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T02:19:35.281990Z digest=sha256:85082e172136bf4dfa080c8b946f59f31ad123fe83da33b029a91206b97dac89

Observation cc9f7a43-ac6f-474c-b24a-723647ad2864 · inbound

Bridging Interleaved Multi-Modal Reasoning as a Unified Decision Process cites this paper.

Bridging Interleaved Multi-Modal Reasoning as a Unified Decision Process Emu3.5: Native Multimodal Models are World Learners

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-12T00:14:19.104498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:14:19.104498Z digest=sha256:678f7f19d5dcd138559fb023f2f9ec991e2b1e1810b2b4cb8bec415375996f26

Observation cb592d58-ace3-4024-9f75-e9503c792d2b · inbound

Transferability Between Understanding and Generation in Unified Multimodal Models cites this paper.

Transferability Between Understanding and Generation in Unified Multimodal Models Emu3.5: Native Multimodal Models are World Learners

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-11T19:17:19.634242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:17:19.634242Z digest=sha256:cbe8968d3a2cea8ca7ac20e380358fea6c8e7939d8037d486c7fa82a602e0d35

Observation 3006de00-dcad-4461-b36d-110c524b0510 · inbound

DeltaV: Thinking with Visual State Updates in Unified Large Multimodal Models cites this paper.

DeltaV: Thinking with Visual State Updates in Unified Large Multimodal Models Emu3.5: Native Multimodal Models are World Learners

Reference 17

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T07:36:58.019384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-10T07:31:26.225257Z digest=sha256:e295425dc3f1364c0ac4f650dcc259acb1593eeb1913dd937df1331d9891ef83

Observation 92e9c856-d691-4660-876c-8410059df1e1 · inbound

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model cites this paper.

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model Emu3.5: Native Multimodal Models are World Learners

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-14T04:10:14.360463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:10:14.360463Z digest=sha256:8fd249d66a7ac4dc1a71f871b3fcda85a4f83b304005cbdfaeaa6550ffe5ece3

Observation 34bac566-138c-4b8b-bdb0-dcae302bd551 · inbound

StructGen: Disambiguating Multi-Reference Image Generation via Structured Context Modeling cites this paper.

StructGen: Disambiguating Multi-Reference Image Generation via Structured Context Modeling Emu3.5: Native Multimodal Models are World Learners

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T22:47:41.067461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:47:41.067461Z digest=sha256:ea10e17b355880d52d6236f00014fce97a09d72fd8de8f697d04096e14ae0a51

Observation 42c2eeee-3ea5-4ec5-abb1-324f0b42951a · inbound

Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing cites this paper.

Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing Emu3.5: Native Multimodal Models are World Learners

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-01T13:38:59.844112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:38:59.844112Z digest=sha256:7f8663f2135fe52773336fa7c6e16ec1c28c4ee5698a5ee0895e90206f95e4f3

Observation cfe40d53-b601-4d15-800a-1e0873088288 · inbound

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text cites this paper.

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text Emu3.5: Native Multimodal Models are World Learners

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T08:39:36.784561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:39:36.784561Z digest=sha256:0403c9f4a5fc18cc9b7bf441142dfc3b76b1ba8946f534f3c130f2989d2df3ce

Observation 72283a80-6ac1-414d-989d-5a71c8d3f60f · inbound

Scaling Native Multimodal Pre-Training From Scratch cites this paper.

Scaling Native Multimodal Pre-Training From Scratch Emu3.5: Native Multimodal Models are World Learners

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-01T06:05:50.098846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:05:50.098846Z digest=sha256:2b076f9c3e0ae9710fd017de5839baec5c472a8fbf53b829cb9a9acd24dc2a94

Observation 5c246877-d4d8-4903-bf19-87dc913028fe · inbound

OmniPhys: Knowledge-Graph-Driven Benchmarking and Collective Optimization for Physical Commonsense in Text-to-Image Generation cites this paper.

OmniPhys: Knowledge-Graph-Driven Benchmarking and Collective Optimization for Physical Commonsense in Text-to-Image Generation Emu3.5: Native Multimodal Models are World Learners

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T01:54:37.817600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T01:54:37.817600Z digest=sha256:d35229b55c1eef130fb2e4339f038ff1f5fdd23904c9d1adba451f99d7be5cac

Observation 21ec4f14-3bc0-4c65-89fa-f8a76171e5b3 · inbound

Test-Time Curriculum for Open-Set AIGC Detection cites this paper.

Test-Time Curriculum for Open-Set AIGC Detection Emu3.5: Native Multimodal Models are World Learners

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-05T00:46:23.384170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:46:23.384170Z digest=sha256:f3168ba2e36f2dfb23db6b7312d6af964dfe126f1fdf2efe973e99f5bb360200

Observation 4d53c8cd-bbf0-441d-b4e5-540f6a860e8f · inbound

Hunyuan3D-Buffalo 1.0: A Unified Multimodal Model for Scalable 3D Generation, Understanding, and Editing cites this paper.

Hunyuan3D-Buffalo 1.0: A Unified Multimodal Model for Scalable 3D Generation, Understanding, and Editing Emu3.5: Native Multimodal Models are World Learners

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T00:13:38.055689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:13:38.055689Z digest=sha256:2c0822e7c5ab1f50bbcaa31e424ed5fb1048ecda2c6fb7148724ba8918e446de

Observation 3ae2b93c-2391-4649-8aea-bf48113677a2 · inbound

ToolArtist: Tool-Using Unified Multimodal Models for Agentic Image Generation cites this paper.

ToolArtist: Tool-Using Unified Multimodal Models for Agentic Image Generation Emu3.5: Native Multimodal Models are World Learners

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T00:15:37.958795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:15:37.958795Z digest=sha256:4ad6c3b4da2accc6fdfdc184a5719cde7f31b1ce23ce5ce059b046c5fa588154

Observation d8a3a56c-eb6e-4729-a48a-1510864321de · inbound

UniSpace: Unified Visual Representation and Scalable Multimodal Modeling cites this paper.

UniSpace: Unified Visual Representation and Scalable Multimodal Modeling Emu3.5: Native Multimodal Models are World Learners

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-14T04:34:59.990040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:34:59.990040Z digest=sha256:cb5d72cea149bf2454ef2327c0c31b5b4b5d2a9b72f089807c8bf6331648c682