Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-18T01:12:13.426640Z
Paper Citation Record · LEDGER
As of 21 August 2026, this Paper Citation Record lists 100 of 131 outbound references and 68 inbound Pith citation observations for arXiv:2510.26583.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-18T01:12:13.426640Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-14T04:34:59.990040Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
100 of 131 outbound references displayed
External citation measurements
0
pith, observed 2026-08-05T02:28:24.338817Z
Observation 22c3e724-4a22-4659-acc2-a1b148d5b13c · outbound
Emu3.5: Native Multimodal Models are World Learners GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation f7c019b8-984e-4390-a746-08f421e500ad · outbound
Emu3.5: Native Multimodal Models are World Learners GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 31ac4d8c-5c96-431f-9dad-d90d2164b4b9 · outbound
Emu3.5: Native Multimodal Models are World Learners Claude 3.5: An ai assistant by anthropic
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation fa6a8267-4b96-43b8-b9f4-d5d8103fabf1 · outbound
Emu3.5: Native Multimodal Models are World Learners The chosen one: Consistent characters in text-to-image diffusion models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 9a960c7f-f3ff-4341-be5c-2f77286b509c · outbound
Emu3.5: Native Multimodal Models are World Learners Qwen2.5-VL Technical Report
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation c77f1912-16e6-47fd-bfff-100e39170382 · outbound
Emu3.5: Native Multimodal Models are World Learners Improving image generation with better captions.https://cdn.openai.com/ papers/dall-e-3.pdf
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation d26faa91-7994-4244-a6c5-96c5e73265dd · outbound
Emu3.5: Native Multimodal Models are World Learners Instructpix2pix: Learning to follow image editing instructions
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation b3c29c2b-d468-48a9-8659-e1e421197366 · outbound
Emu3.5: Native Multimodal Models are World Learners AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 073cd66e-f006-41e3-8698-411ccea589ea · outbound
Emu3.5: Native Multimodal Models are World Learners Coyo-700m: Image-text pair dataset.https://github.com/kakaobrain/coyo-dataset
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation f4217f92-b9a0-4c3a-bfd6-c74a315df1b0 · outbound
Emu3.5: Native Multimodal Models are World Learners HiDream-I1: A High-Efficient Image Generative Foundation Model with Sparse Diffusion Transformer
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 692f1dc0-35f6-49af-8546-cda699a05ef7 · outbound
Emu3.5: Native Multimodal Models are World Learners Flash diffusion: Acceler- ating any conditional diffusion model for few steps image generation
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 7074b797-ab45-45af-9c31-e98a5a625764 · outbound
Emu3.5: Native Multimodal Models are World Learners OneIG-Bench: Omni-dimensional Nuanced Evaluation for Image Generation
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 19919518-d516-461f-825c-fd9d6b083d99 · outbound
Emu3.5: Native Multimodal Models are World Learners Conceptual 12m: Pushing web- scale image-text pre-training to recognize long-tail visual concepts
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 23603f50-c379-484c-ab6c-d34654ecbce7 · outbound
Emu3.5: Native Multimodal Models are World Learners Interleaved Scene Graphs for Interleaved Text-and-Image Generation Assessment
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation cdb517c1-64f1-4e75-b036-5b9d17c4ff32 · outbound
Emu3.5: Native Multimodal Models are World Learners BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 5aa384f7-4c39-4008-bbee-35e3365c1c77 · outbound
Emu3.5: Native Multimodal Models are World Learners ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 45131fa5-f0a2-4284-ada6-c5bd38d6af30 · outbound
Emu3.5: Native Multimodal Models are World Learners MultiRef: Controllable Image Generation with Multiple Visual References
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation e11ce4b4-d839-4d1d-9251-13d47c24cab1 · outbound
Emu3.5: Native Multimodal Models are World Learners PosterCraft: Rethinking High-Quality Aesthetic Poster Generation in a Unified Framework
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation d1993717-6c22-45a7-8da8-70a31188512b · outbound
Emu3.5: Native Multimodal Models are World Learners Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation aa97f3af-a802-4178-a1b1-fa0b2b49ad38 · outbound
Emu3.5: Native Multimodal Models are World Learners Bagel: a web- based bacteriocin genome mining tool.Nucleic acids research, 34(suppl_2):W273–W279
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 108c810a-5ce7-4e9b-9833-5c38a2b2e838 · outbound
Emu3.5: Native Multimodal Models are World Learners insightface.https://github.com/deepinsight/insightface
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 108fac73-3b76-4455-9a43-09c042f7fe64 · outbound
Emu3.5: Native Multimodal Models are World Learners Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 78e84d4c-532e-4ba0-ab53-b35db87139bb · outbound
Emu3.5: Native Multimodal Models are World Learners Scaling vision transformers to 22 billion parameters
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 11fc43eb-8bd3-43e4-b5f8-463c8ad640ae · outbound
Emu3.5: Native Multimodal Models are World Learners Uniform discrete diffusion with metric path for video generation
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 5424ae77-2b00-428c-9543-9bf10ead96bb · outbound
Emu3.5: Native Multimodal Models are World Learners Retinaface: Single-shot multi-level face localisation in the wild
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation a2ee46ab-e5ee-46b6-b264-8e342353d1da · outbound
Emu3.5: Native Multimodal Models are World Learners Textcrafter: Accurately rendering multiple texts in complex visual scenes
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation cd3675af-ace2-49ff-b443-1acb0bd85838 · outbound
Emu3.5: Native Multimodal Models are World Learners Scaling rectified flow transformers for high-resolution image synthesis
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 0c832447-bb02-4501-a2cc-82b42bdc220d · outbound
Emu3.5: Native Multimodal Models are World Learners Taming transformers for high-resolution image synthesis
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation c384ecdd-b071-44c0-b1ec-afd351e22510 · outbound
Emu3.5: Native Multimodal Models are World Learners Datacomp: In search of the next generation of multimodal datasets.Advances in Neural Information Processing Systems, 36:27092– 27112
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 2f61a95d-cb7e-49ee-9db5-03930dc3cb3b · outbound
Emu3.5: Native Multimodal Models are World Learners Seedream 3.0 Technical Report
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 891cca51-81e1-4132-a405-5057cb9a82b6 · outbound
Emu3.5: Native Multimodal Models are World Learners Discrete flow matching.Advances in Neural Information Processing Systems, 37:133345–133385
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 1006dae3-4b1b-47cf-856a-808dbe738fe7 · outbound
Emu3.5: Native Multimodal Models are World Learners SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation a0bf2467-64fc-4bdc-a16e-27a87faa2575 · outbound
Emu3.5: Native Multimodal Models are World Learners X-Omni: Reinforcement Learning Makes Discrete Autoregressive Image Generative Models Great Again
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation a573a1be-f126-48ee-9657-904ff3c374f9 · outbound
Emu3.5: Native Multimodal Models are World Learners Geneval: An object-focused framework for evaluating text-to-image alignment.Advances in Neural Information Processing Systems, 36
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation f7d3aedd-eca1-43c0-8745-e8f2d4f75e62 · outbound
Emu3.5: Native Multimodal Models are World Learners Gemini 2.0 flash.https://developers.googleblog.com/en/ experiment-with-gemini-20-flash-native-image-generation
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation f2cd30cf-4bad-447a-8c0a-855d26c1d5c5 · outbound
Emu3.5: Native Multimodal Models are World Learners Imagen 3
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 0f118b8a-a3f0-449e-9018-9d6034243968 · outbound
Emu3.5: Native Multimodal Models are World Learners Imagen 4
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 933d34f6-4ffb-4fd8-ae93-ffb88ebe1745 · outbound
Emu3.5: Native Multimodal Models are World Learners Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 2c3082ae-d2ce-416b-a07b-17141474286e · outbound
Emu3.5: Native Multimodal Models are World Learners Infinity: Scaling bitwise autoregressive modeling for high-resolution image synthesis
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 59721299-0f79-4859-ba53-888479184049 · outbound
Emu3.5: Native Multimodal Models are World Learners Measuring colorfulness in natural images
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 2f28fe49-1ccd-4cc7-a525-605679157759 · outbound
Emu3.5: Native Multimodal Models are World Learners ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 6d142b67-4384-4967-ab09-8a0121a0e674 · outbound
Emu3.5: Native Multimodal Models are World Learners Image-to-image translation with conditional adversarial networks
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 19bcfd07-38fd-4444-89d6-fac2afc1f9a2 · outbound
Emu3.5: Native Multimodal Models are World Learners Lego-edit: A general image editing framework with model-level bricks and mllm builder
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation ce6bc978-0ffb-4818-844a-414e6ead6f70 · outbound
Emu3.5: Native Multimodal Models are World Learners InfiniteYou: Flexible Photo Recrafting While Preserving Your Identity
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 1ac987de-fa1e-4ad1-9375-c2cf1bc9ad6f · outbound
Emu3.5: Native Multimodal Models are World Learners Musiq: Multi-scale image quality transformer
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 622f510a-35ac-42ce-87a3-7b71bed633cc · outbound
Emu3.5: Native Multimodal Models are World Learners Unresolved cited work
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 62e3ca79-76d4-41bd-8f5e-9e4373267ef0 · outbound
Emu3.5: Native Multimodal Models are World Learners Gonzalez, Hao Zhang, and Ion Stoica
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 088f4e04-7532-45d3-9ac1-529fded71c20 · outbound
Emu3.5: Native Multimodal Models are World Learners Flux.https://github.com/black-forest-labs/flux
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 39c2e95b-f16d-41ca-a164-b12167a709ce · outbound
Emu3.5: Native Multimodal Models are World Learners FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 954be859-f10a-466a-b4d2-1f9796993024 · outbound
Emu3.5: Native Multimodal Models are World Learners Grounding image matching in 3d with mast3r
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation b941330f-2654-44f9-ae65-3526696db315 · outbound
Emu3.5: Native Multimodal Models are World Learners LLaVA-OneVision: Easy Visual Task Transfer
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 10165464-9ec3-435a-9dfa-43bb34e3e5fe · outbound
Emu3.5: Native Multimodal Models are World Learners Infinity instruct: Scaling instruction selection and synthesis to enhance language models
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 1ddd2c17-163f-4151-a9a9-3f8db59b066e · outbound
Emu3.5: Native Multimodal Models are World Learners Sekai: A video dataset towards world exploration
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 8e9c6323-696b-4644-a8d2-a3115c9915aa · outbound
Emu3.5: Native Multimodal Models are World Learners UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 8d00c863-c005-42f1-8086-6686d7dabc88 · outbound
Emu3.5: Native Multimodal Models are World Learners Focal loss for dense object detection
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 118e3dbe-cb2d-437b-979c-9ee26854c37e · outbound
Emu3.5: Native Multimodal Models are World Learners Flow Matching for Generative Modeling
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 4fe847ad-4b8e-4dd8-bfdf-12fd48fd0cba · outbound
Emu3.5: Native Multimodal Models are World Learners Improved baselines with visual instruction tuning
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 386bc1bb-2611-4e77-889e-0fdac997a5ce · outbound
Emu3.5: Native Multimodal Models are World Learners Step1X-Edit: A Practical Framework for General Image Editing
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 089739e2-63ab-436a-b08e-e4c46c79e9af · outbound
Emu3.5: Native Multimodal Models are World Learners Dpm-solver: A fast ode solver for diffusion probabilistic model sampling in around 10 steps.Advances in neural information processing systems, 35:5775–5787
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 5b1bf519-4281-4e10-b7fa-ecf85beb41a6 · outbound
Emu3.5: Native Multimodal Models are World Learners Open-MAGVIT2: An Open-Source Project Toward Democratizing Auto-regressive Visual Generation
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 9cae939b-c162-4840-84b3-a6cf7cffaa01 · outbound
Emu3.5: Native Multimodal Models are World Learners Sit: Exploring flow and diffusion-based generative models with scalable interpolant transform- ers
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation ef8df658-3626-4a2b-8800-6ba11319e378 · outbound
Emu3.5: Native Multimodal Models are World Learners Midjourney
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 3f0db7fb-514b-49df-adaa-bc3f883a4e7e · outbound
Emu3.5: Native Multimodal Models are World Learners Gpt-4o.https://openai.com/index/introducing-4o-image-generation
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 818aa069-da24-46d4-a6b6-0b8dc183df5d · outbound
Emu3.5: Native Multimodal Models are World Learners Image generation API.https://openai.com/index/image-generation-api/
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 27d1099d-0c6e-468e-afea-399088f35863 · outbound
Emu3.5: Native Multimodal Models are World Learners DINOv2: Learning Robust Visual Features without Supervision
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 8da3dbca-297d-467a-897c-74210863a038 · outbound
Emu3.5: Native Multimodal Models are World Learners Open x-embodiment: Robotic 40 learning datasets and rt-x models: Open x-embodiment collaboration 0
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 0cf2164b-b9b0-46ef-8a1a-53a5b7a0d03a · outbound
Emu3.5: Native Multimodal Models are World Learners Journeydb: A benchmark for generative image understanding
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation be314694-5e33-48de-a68d-eded2df1b19a · outbound
Emu3.5: Native Multimodal Models are World Learners ICE-Bench: A Unified and Comprehensive Benchmark for Image Creating and Editing
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 5d59952a-bf9e-46c5-8b02-12ff9887d1a7 · outbound
Emu3.5: Native Multimodal Models are World Learners Scalable Diffusion Models with Transformers
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation b2edb98c-1047-4507-ae84-f427a3d890cb · outbound
Emu3.5: Native Multimodal Models are World Learners Tokenflow: Unified image tokenizer for multimodal understanding and generation
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 3d1e6416-407b-452f-be09-65236acf5898 · outbound
Emu3.5: Native Multimodal Models are World Learners Robust speech recognition via large-scale weak supervision
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation cafbb9c9-9ab5-4935-a2bc-47a3d2800170 · outbound
Emu3.5: Native Multimodal Models are World Learners Recraft.https://www.recraft.ai/
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation cf75fe7b-e3b7-4854-9f1b-a45e129e40af · outbound
Emu3.5: Native Multimodal Models are World Learners Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation b6cc72de-785a-4288-88b9-18fbfb93c248 · outbound
Emu3.5: Native Multimodal Models are World Learners Imagenet large scale visual recognition challenge.International journal of computer vision, 115(3):211–252
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation e6bfb617-6f55-4f8f-864f-092c2fcace0a · outbound
Emu3.5: Native Multimodal Models are World Learners Laion-5b: An open large-scale dataset for training next generation image-text models.Advances in neural information processing systems, 35:25278–25294
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 568a6c2d-a24e-466b-9b2c-595419b1dd68 · outbound
Emu3.5: Native Multimodal Models are World Learners DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation a77e738e-23ee-424b-84aa-9478fe595362 · outbound
Emu3.5: Native Multimodal Models are World Learners Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation c854674e-5a63-48a3-904f-c70247a1bd89 · outbound
Emu3.5: Native Multimodal Models are World Learners GLU Variants Improve Transformer
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation ed0dd9fe-d810-4cb2-8e57-7ba100a0033d · outbound
Emu3.5: Native Multimodal Models are World Learners Storygpt-v: Large language models as consistent story vi- sualizers
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation eca8cc75-7f27-4bca-857c-4bd97203bca3 · outbound
Emu3.5: Native Multimodal Models are World Learners HybridFlow: A Flexible and Efficient RLHF Framework
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation ae32dae3-3d6b-46d0-a3d7-3906182ad645 · outbound
Emu3.5: Native Multimodal Models are World Learners Scalable Image Tokenization with Index Backpropagation Quantization
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation c9ba1edb-2354-418b-9c38-f0f87cc495e1 · outbound
Emu3.5: Native Multimodal Models are World Learners Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 6970193c-06e6-4feb-84eb-96d084a1ab5f · outbound
Emu3.5: Native Multimodal Models are World Learners Denoising Diffusion Implicit Models
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 9b4fce4f-03f0-4b94-adbd-7c319f45989b · outbound
Emu3.5: Native Multimodal Models are World Learners Nemotron-CC: Transforming Common Crawl into a Refined Long-Horizon Pretraining Dataset
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 503ce59f-d8cd-47da-8603-3876000d39f8 · outbound
Emu3.5: Native Multimodal Models are World Learners Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation bd0b1bc5-5322-4bde-a831-09c7f9c3801e · outbound
Emu3.5: Native Multimodal Models are World Learners Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 2c7aae4d-ae16-4695-b11c-cf66d13c5a91 · outbound
Emu3.5: Native Multimodal Models are World Learners Generative multimodal models are in-context learn- ers
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation a74d0875-1b45-4f8f-a7d0-2e3e956c6f00 · outbound
Emu3.5: Native Multimodal Models are World Learners Emu: Generative pretraining in multimodality
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 2a9b32f0-c245-4166-b57e-bd7717d0dc84 · outbound
Emu3.5: Native Multimodal Models are World Learners Chameleon: Mixed-Modal Early-Fusion Foundation Models
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation a5c21e30-15af-46d0-a959-379a48b053e4 · outbound
Emu3.5: Native Multimodal Models are World Learners FlagScale: A unified meta-framework enabling adaptive heterogeneous computing for the llm ecosystem.https://github.com/FlagOpen/FlagScale
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 0d60490a-40c8-4c67-9ee6-f7fa421402f5 · outbound
Emu3.5: Native Multimodal Models are World Learners Gemini 2.5 flash & gemini 2.5 flash image model card.https://storage
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation cbc2613b-9e37-4c01-8bce-edd438dd9581 · outbound
Emu3.5: Native Multimodal Models are World Learners Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 4bd5cfe4-5161-4574-9f24-851b5a3fe7b3 · outbound
Emu3.5: Native Multimodal Models are World Learners Kolors 2.0.https://app.klingai.com/cn/
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 2afe1780-1b23-4a20-aeae-8031f00e81a9 · outbound
Emu3.5: Native Multimodal Models are World Learners Raft: Recurrent all-pairs field transforms for optical flow
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 92e309a8-d75d-46a5-b963-7013954c790d · outbound
Emu3.5: Native Multimodal Models are World Learners Training-free consistent text-to-image generation.ACM Transactions on Graphics (TOG), 43(4):1– 18
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation ba33742d-c700-4b24-8781-215ec2284199 · outbound
Emu3.5: Native Multimodal Models are World Learners Visual autoregressive model- ing: Scalable image generation via next-scale prediction.Advances in neural information processing systems, 37:84839–84865
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 63a4f167-f0c3-4913-8b06-110ecc88c93c · outbound
Emu3.5: Native Multimodal Models are World Learners Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 0878c272-1884-463f-96a2-19246211155a · outbound
Emu3.5: Native Multimodal Models are World Learners Wan: Open and advanced large-scale video generative models
Reference 98
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 6ff02927-3e71-458c-97da-be03eea707d2 · outbound
Emu3.5: Native Multimodal Models are World Learners Textatlas5m: A large-scale dataset for dense text image generation
Reference 99
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 52937886-7a67-4471-b533-8e1c5411bc4d · outbound
Emu3.5: Native Multimodal Models are World Learners Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation f6d047db-2a61-435b-af07-280e9e5a2503 · inbound
Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm Emu3.5: Native Multimodal Models are World Learners
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation e85d8781-85a1-40a8-981f-bcbd8cbbbdc8 · inbound
CG-MLLM: Captioning and Generating 3D content via Multi-modal Large Language Models Emu3.5: Native Multimodal Models are World Learners
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 7b9833f7-cfc6-4a0f-a2bd-19d1152d0294 · inbound
MentisOculi: Revealing the Limits of Reasoning with Mental Imagery Emu3.5: Native Multimodal Models are World Learners
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39446122-d597-4d44-8273-4b70f81fc342 · inbound
LLaMo: Scaling Pretrained Language Models for Unified Motion Understanding and Generation with Continuous Autoregressive Tokens Emu3.5: Native Multimodal Models are World Learners
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 3a721f1f-b528-43cd-9b49-b910113cc8c1 · inbound
VLANeXt: Recipes for Building Strong VLA Models Emu3.5: Native Multimodal Models are World Learners
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 35772a41-bbc4-4737-b365-352d08d46ff9 · inbound
WeEdit: A Dataset, Benchmark and Glyph-Guided Framework for Text-centric Image Editing Emu3.5: Native Multimodal Models are World Learners
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 900bf7e2-3b1b-486b-aa15-3b9281c21bf4 · inbound
TRACE: High-Fidelity 3D Scene Editing via Tangible Reconstruction and Geometry-Aligned Contextual Video Masking Emu3.5: Native Multimodal Models are World Learners
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e272cddb-bc6a-4a9a-a8dc-f77c87098087 · inbound
OpenWorldLib: A Unified Codebase and Definition of Advanced World Models Emu3.5: Native Multimodal Models are World Learners
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 3b0d5a40-e62d-4bcd-a180-c53f7d1e3158 · inbound
OpenWorldLib: A Unified Codebase and Definition of Advanced World Models Emu3.5: Native Multimodal Models are World Learners
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd8475af-1a91-4dc4-829e-7d65da4558c1 · inbound
Learning Preference-Based Objectives from Clinical Narratives for Dynamic Sepsis Treatment Emu3.5: Native Multimodal Models are World Learners
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd1eadb0-4071-4a9d-bdce-b13e5ebca823 · inbound
TorchUMM: A Unified Multimodal Model Codebase for Evaluation, Analysis, and Post-training Emu3.5: Native Multimodal Models are World Learners
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 5569b540-477f-47fc-b365-ba8de8784322 · inbound
TorchUMM: A Unified Multimodal Model Codebase for Evaluation, Analysis, and Post-training Emu3.5: Native Multimodal Models are World Learners
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 5550148e-c7ce-44ce-9847-34e2951046c7 · inbound
UniGeo: Unifying Geometric Guidance for Camera-Controllable Image Editing via Video Models Emu3.5: Native Multimodal Models are World Learners
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 0d443724-2970-41cf-83f2-5b1bd40d2526 · inbound
UniGeo: Unifying Geometric Guidance for Camera-Controllable Image Editing via Video Models Emu3.5: Native Multimodal Models are World Learners
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation e455f6a8-2414-4623-9b14-f094bd6ba615 · inbound
Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation Emu3.5: Native Multimodal Models are World Learners
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 19177e61-0370-4e00-b196-c64d4bda4c92 · inbound
Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation Emu3.5: Native Multimodal Models are World Learners
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 9a89ef6d-d376-4227-83b6-4641923932c2 · inbound
UDM-GRPO: Stable and Efficient Group Relative Policy Optimization for Uniform Discrete Diffusion Models Emu3.5: Native Multimodal Models are World Learners
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 0a6cd5dd-dda3-41c9-b163-87feb64903e3 · inbound
UDM-GRPO: Stable and Efficient Group Relative Policy Optimization for Uniform Discrete Diffusion Models Emu3.5: Native Multimodal Models are World Learners
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation a0ee33fa-5819-4e27-96ce-84c4528d6fbe · inbound
Exploring Spatial Intelligence from a Generative Perspective Emu3.5: Native Multimodal Models are World Learners
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation f2aacf6d-66d8-4a3b-9069-1c8944a7aa32 · inbound
Context Unrolling in Omni Models Emu3.5: Native Multimodal Models are World Learners
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 2cf5a54f-e46d-4384-9f6d-6eea870b9d79 · inbound
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation Emu3.5: Native Multimodal Models are World Learners
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 43c9d681-bb25-43b8-b687-61e13146f13c · inbound
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation Emu3.5: Native Multimodal Models are World Learners
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 82a09e16-f052-425c-ad12-0dc305fdaeb4 · inbound
Unified 4D World Action Modeling from Video Priors with Asynchronous Denoising Emu3.5: Native Multimodal Models are World Learners
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 580e768a-833a-4f20-bb57-70499ac369c6 · inbound
Unified 4D World Action Modeling from Video Priors with Asynchronous Denoising Emu3.5: Native Multimodal Models are World Learners
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 1475be14-4922-4694-ad74-01d9705f3368 · inbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Emu3.5: Native Multimodal Models are World Learners
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 389f1f92-6c60-4f67-a9ce-bb8d79c483e7 · inbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Emu3.5: Native Multimodal Models are World Learners
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 0b36d08d-7b99-4d1e-a604-508c19f83d2f · inbound
MULTITEXTEDIT: Benchmarking Cross-Lingual Degradation in Text-in-Image Editing Emu3.5: Native Multimodal Models are World Learners
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation fa20c607-f838-4b49-99d9-706f2225c779 · inbound
Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Emu3.5: Native Multimodal Models are World Learners
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation e33e2eea-d06b-4c15-ba1d-632fae2e7cf2 · inbound
Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Emu3.5: Native Multimodal Models are World Learners
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 15f8e895-b4ee-41ae-abcc-e9c546de505b · inbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Emu3.5: Native Multimodal Models are World Learners
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 79afed4e-be56-4a06-b432-c406a2b5fb51 · inbound
Edit-Compass & EditReward-Compass: A Unified Benchmark for Image Editing and Reward Modeling Emu3.5: Native Multimodal Models are World Learners
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation b7561dea-ddb6-4f83-bada-794f63a097c7 · inbound
InsightTok: Improving Text and Face Fidelity in Discrete Tokenization for Autoregressive Image Generation Emu3.5: Native Multimodal Models are World Learners
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 773d1a32-1274-4cd2-bb7f-56a356d6077e · inbound
Sketch Then Paint: Hierarchical Reinforcement Learning for Diffusion Multi-Modal Large Language Models Emu3.5: Native Multimodal Models are World Learners
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 5d8a8294-0514-4bf0-aaba-bbad823299a8 · inbound
LatentUMM: Dual Latent Alignment for Unified Multimodal Models Emu3.5: Native Multimodal Models are World Learners
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 14d14ed3-5410-4461-8c61-f7459cdc3b32 · inbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy Emu3.5: Native Multimodal Models are World Learners
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 9813bfcc-dc4c-435b-92d1-de6b9fce24e2 · inbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy Emu3.5: Native Multimodal Models are World Learners
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation e880fb75-7176-44fb-bcbf-ac24bb537257 · inbound
TextSculptor: Training and Benchmarking Scene Text Editing Emu3.5: Native Multimodal Models are World Learners
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation cdc50307-8237-4bd7-b4ab-b319acbfe5f8 · inbound
Uni-Edit: Intelligent Editing Is A General Task For Unified Model Tuning Emu3.5: Native Multimodal Models are World Learners
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation aa556b57-17a9-4ca9-a115-0d8dcdca9832 · inbound
Uni-Edit: Intelligent Editing Is A General Task For Unified Model Tuning Emu3.5: Native Multimodal Models are World Learners
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation cbbe4208-39bb-4712-98d5-1a512b26d723 · inbound
Bernini: Latent Semantic Planning for Video Diffusion Emu3.5: Native Multimodal Models are World Learners
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 140741fe-ff67-4980-b9e1-07e38a88ec69 · inbound
Guess the Unified Model: How Much Can We Recover from Generated Images? Emu3.5: Native Multimodal Models are World Learners
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation a9c42e88-73f1-40ec-a5e0-8feb6ef22e79 · inbound
Toward Native Multimodal Modeling: A Roadmap Emu3.5: Native Multimodal Models are World Learners
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 8a02ef79-98bd-4a09-9d19-104b173be02b · inbound
OmniVerifier-M1: Multimodal Meta-Verifier with Explicit Structured Recalibration Emu3.5: Native Multimodal Models are World Learners
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation c48ecdcb-6dd0-44f8-8d91-44428270f8f5 · inbound
GenClaw: Code-Driven Agentic Image Generation Emu3.5: Native Multimodal Models are World Learners
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 771fcf27-c3a8-4744-b39e-0acfddc9fc4b · inbound
Echo-Memory: A Controlled Study of Memory in Action World Models Emu3.5: Native Multimodal Models are World Learners
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation ec2e823d-d79b-49a7-bf2f-91a6de4b514d · inbound
SVoT: State-aware Visualization-of-Thought for Spatial Reasoning via Reinforcement Learning Emu3.5: Native Multimodal Models are World Learners
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 5c1f5286-4b7a-4218-935f-241361e74011 · inbound
InterleaveThinker: Reinforcing Agentic Interleaved Generation Emu3.5: Native Multimodal Models are World Learners
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 93da6e2c-168e-4ab1-b890-5b0c353b7d09 · inbound
ThinkingVLA: Interleaved Vision and Language Reasoning for Robotic Manipulation Emu3.5: Native Multimodal Models are World Learners
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 15e506a9-738a-4495-880a-c09e55965272 · inbound
Spotlight: Synergizing Seed Exploration and Spot GPUs for DiT RL Post-Training Emu3.5: Native Multimodal Models are World Learners
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 1e143c10-f618-42d0-89f9-0fa70d00660e · inbound
IV-CoT: Implicit Visual Chain-of-Thought for Structure-Aware Text-to-Image Generation Emu3.5: Native Multimodal Models are World Learners
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 2d629f4d-3d9b-4062-835c-089727b01795 · inbound
Nemotron-Labs-Diffusion-Image: Advancing Masked Discrete Diffusion for High-Resolution Image Synthesis Emu3.5: Native Multimodal Models are World Learners
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation afc89459-cb70-47f4-b347-1321ce527ceb · inbound
Nemotron-Labs-Diffusion-Image: Advancing Masked Discrete Diffusion for High-Resolution Image Synthesis Emu3.5: Native Multimodal Models are World Learners
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc548f2a-8da7-4b01-ad31-a5bd8568a2e9 · inbound
Illuminating Unified Multimodal Model for Free-form Interleaved Text-Image Generation Emu3.5: Native Multimodal Models are World Learners
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation c98223a1-3a5a-450e-be53-536016f92f0f · inbound
MemLearner: Learning to Query Context memory for Video World Models Emu3.5: Native Multimodal Models are World Learners
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation e011d6fd-0809-420a-9206-5f96753abb7d · inbound
WorldBagel: Uncovering the Power of Unified Multimodal Models for Vision-Language-Action-World Modeling Emu3.5: Native Multimodal Models are World Learners
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc9f7a43-ac6f-474c-b24a-723647ad2864 · inbound
Bridging Interleaved Multi-Modal Reasoning as a Unified Decision Process Emu3.5: Native Multimodal Models are World Learners
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb592d58-ace3-4024-9f75-e9503c792d2b · inbound
Transferability Between Understanding and Generation in Unified Multimodal Models Emu3.5: Native Multimodal Models are World Learners
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3006de00-dcad-4461-b36d-110c524b0510 · inbound
DeltaV: Thinking with Visual State Updates in Unified Large Multimodal Models Emu3.5: Native Multimodal Models are World Learners
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 92e9c856-d691-4660-876c-8410059df1e1 · inbound
Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model Emu3.5: Native Multimodal Models are World Learners
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34bac566-138c-4b8b-bdb0-dcae302bd551 · inbound
StructGen: Disambiguating Multi-Reference Image Generation via Structured Context Modeling Emu3.5: Native Multimodal Models are World Learners
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42c2eeee-3ea5-4ec5-abb1-324f0b42951a · inbound
Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing Emu3.5: Native Multimodal Models are World Learners
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cfe40d53-b601-4d15-800a-1e0873088288 · inbound
Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text Emu3.5: Native Multimodal Models are World Learners
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72283a80-6ac1-414d-989d-5a71c8d3f60f · inbound
Scaling Native Multimodal Pre-Training From Scratch Emu3.5: Native Multimodal Models are World Learners
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c246877-d4d8-4903-bf19-87dc913028fe · inbound
OmniPhys: Knowledge-Graph-Driven Benchmarking and Collective Optimization for Physical Commonsense in Text-to-Image Generation Emu3.5: Native Multimodal Models are World Learners
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21ec4f14-3bc0-4c65-89fa-f8a76171e5b3 · inbound
Test-Time Curriculum for Open-Set AIGC Detection Emu3.5: Native Multimodal Models are World Learners
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d53c8cd-bbf0-441d-b4e5-540f6a860e8f · inbound
Hunyuan3D-Buffalo 1.0: A Unified Multimodal Model for Scalable 3D Generation, Understanding, and Editing Emu3.5: Native Multimodal Models are World Learners
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ae2b93c-2391-4649-8aea-bf48113677a2 · inbound
ToolArtist: Tool-Using Unified Multimodal Models for Agentic Image Generation Emu3.5: Native Multimodal Models are World Learners
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8a3a56c-eb6e-4729-a48a-1510864321de · inbound
UniSpace: Unified Visual Representation and Scalable Multimodal Modeling Emu3.5: Native Multimodal Models are World Learners
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.