Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-21T07:56:34.034047Z
Paper Citation Record · LEDGER
As of 22 August 2026, this Paper Citation Record lists 100 of 151 outbound references and 9 inbound Pith citation observations for arXiv:2605.18678.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-21T07:56:34.034047Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T00:17:33.562911Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
100 of 151 outbound references displayed
External citation measurements
0
pith, observed 2026-08-05T02:28:24.338817Z
Observation 76f04b48-247a-4110-827c-7c13582eb821 · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 995c1c97-5af1-4b03-bae1-01630cf359bd · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy Flamingo: a visual language model for few-shot learning
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 19b9e6fc-8593-4b9d-aa26-c6f3fcacc501 · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation b9956c71-7feb-46b8-92ad-0445151d940d · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy Qwen3-VL Technical Report
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 58cdf9be-a587-42b3-955d-9ac8e02e3d84 · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy Qwen2.5-VL Technical Report
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation df09cb3e-1d81-4d59-84a3-c5e1a487817f · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy Improving image generation with better captions.Computer Science
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation bba7fdbd-3482-4775-b6ba-073f017e7ef1 · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy Diffusion Self-Distillation for Zero-Shot Customized Image Generation
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 17f51a29-40c0-43d4-898c-1482aa940cf2 · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy HunyuanImage 3.0 Technical Report
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 9a3e0a54-70b4-4614-b7d5-739b5fea1eb9 · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy Maskgit: Masked generative image transformer
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 7e478804-4605-41bb-93b6-ca5866600e8a · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy Videocrafter2: Overcoming data limitations for high-quality video diffusion models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation a7da9052-8c29-4148-8e9d-8da1f203e3c9 · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation b0bc4d67-3364-4bac-ad67-bd0102cb5a3c · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy Pixart-σ: Weak-to-strong training of diffusion transformer for 4k text-to-image generation
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation e81e7832-46d1-4918-8585-db1038c7d55e · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy TimeMarker: A Versatile Video-LLM for Long and Short Video Understanding with Superior Temporal Localization Ability
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 611133b9-3f0a-41af-972c-c23a0b138a8b · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 00123fe2-4cba-471b-ab1e-7b52c41676d9 · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 3a05cbe2-d270-4427-8349-e48014c6e614 · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 94e83882-5ab2-49c1-a065-6d4a07dc09ec · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy UMO: Scaling Multi-Identity Consistency for Image Customization via Matching Reward
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation a5d1dbaf-8baa-427d-895d-a97bef0a3d2a · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy PaddleOCR 3.0 Technical Report
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 9813bfcc-dc4c-435b-92d1-de6b9fce24e2 · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy Emu3.5: Native Multimodal Models are World Learners
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 73f0a0e9-3b2c-4d3b-ab98-054ea216cbfa · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy Instructblip: Towards general-purpose vision-language models with instruction tuning
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 02508724-e637-43cf-b990-f1efea13a798 · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy ChatUMM: Robust Context Tracking for Conversational Interleaved Generation
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 1fb140fb-cb35-449a-8f1d-615d3fd193ef · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy Scaling vision transformers to 22 billion parameters
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 48a508c7-b0f0-4546-aa2a-62fb452f7ce1 · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy Emerging Properties in Unified Multimodal Pretraining
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 7fd5e614-ce13-465f-85cd-b7f508f632e5 · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy Cogview: Mastering text-to-image generation via transformers.NIPS, 34:19822–19835
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 02a9cb21-26e2-45a1-916c-69d3f140e8ce · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy Taming transformers for high-resolution image synthesis
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation ede6043b-ad4b-4bbd-9cba-91eedc485ce2 · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy Scaling rectified flow transformers for high-resolution image synthesis
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation d5172ce2-5713-4439-8ec6-0c928cda97c6 · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy Unified Autoregressive Visual Generation and Understanding with Continuous Tokens
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 60c46325-5641-4837-a59b-d9537c6f88d3 · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation b5f3473d-104f-479a-a975-1291945aca4a · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy Dreamlite: A lightweight on-device unified model for image generation and editing
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 292da351-5d4b-44ff-a72d-ee1556019aad · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy Feededit: Text-based image editing with dynamic feedback regulation
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation faa5f6d2-ace7-41a3-bb0d-bfeefba198f0 · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy Layeredit: Disentangled multi-object editing via conflict-aware multi-layer learning
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 6f7d5ead-01ae-4faf-a593-de9c9a5619cd · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy Mini-internvl: a flexible-transfer pocket multi-modal model with 5% parameters and 90% performance
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 1edc4828-c9a1-4244-acf0-7d70a78b3698 · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 7fa33aea-fdaa-43b7-996a-1dfadf59551e · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy Geneval: An object-focused framework for evaluating text-to-image alignment.Advancesin Neural Information Processing Systems, 36:52132–52152
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation d40669be-151e-4975-b21b-2fca56e80afd · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy Gemini 3 Pro Image Model Card
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 91d1bb12-0a51-463b-9cb1-d48933ed202a · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation be4499ca-34fa-4c9a-84c8-9a360a5777c7 · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy Tv2tv: A unified framework for interleaved language and video generation
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 7d918d47-cfee-4b84-8f9e-cc94aa5403cd · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy Emma: Efficient multimodal understanding, generation, and editing with a unified architecture
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 35cd92d7-e045-4e21-9fe5-a80380e5af89 · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy Classifier-Free Diffusion Guidance
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation a6143f1d-c9f3-4681-b171-8e2b9e1a30b4 · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy Denoising diffusion probabilistic models.NIPS, 33:6840–6851
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation f964116d-5729-4333-813b-ba0bf40e55ce · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 663963de-e5fd-4d60-9076-f1f594dce0ce · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation da9ec087-5f62-47b2-b537-7de764840ed1 · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy Dse-gan: Dynamic semantic evolution generative adversarial network for text-to-image generation
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation eb69059c-3617-4839-ae24-2b70e7937980 · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy Towards accurate image coding: Improved autoregressive image generation with dynamic vector quantization
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation a2d11953-8a1e-4a56-bffc-037cbb247d59 · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy Realcustom: Narrowing real text word for real-time open-domain text-to-image customization
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 07d9c40a-6b4f-4272-8c65-171a4b918770 · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy Self forcing: Bridging the train-test gap in autoregressive video diffusion.Advances in Neural Information Processing Systems, 38:167283–167308
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 386afc2f-968d-41e7-a1d1-794e78a54dd1 · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy Vbench: Comprehensive benchmark suite for video generative models
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation a6fffe26-aa79-4e9d-8108-a06aaf7e0ad9 · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy Vace: All-in-one video creation and editing
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation d9baa9e7-5671-43c7-bbaa-3d1cbcb8f471 · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy EditVerse: Unifying Image and Video Editing and Generation with In-Context Learning
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 3c03f3d8-85e4-441d-9848-c4352dbde82c · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy FullDiT: Multi-Task Video Generative Foundation Model with Full Attention
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 45936d82-8f83-44b6-8311-49d512248324 · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy Kling ai.https://klingai.kuaishou.com/
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation a5986faf-0494-466a-9eb0-a7170ce8fea6 · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy VideoPoet: A Large Language Model for Zero-Shot Video Generation
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 5557a68d-39cc-45bc-8e98-f18b1d8cc77d · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy AnyV2V: A Tuning-Free Framework For Any Video-to-Video Editing Tasks
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 93676173-f707-4a8c-8f4a-0de9b130506e · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy Flux: Official inference repository for flux.1 models
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation ad3a9c14-1654-40d4-abda-6ef04d860b96 · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation a509b74f-bebe-44cb-9bf6-eb63dcc5641f · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy Obelics: An open web-scale filtered dataset of interleaved image-text documents.Advancesin Neural Information Processing Systems, 36:71683–71702
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation bbd74472-1af9-475e-8d86-9ca007f2e93c · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy LLaVA-OneVision: Easy Visual Task Transfer
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 29b32ebf-ba79-417c-88f8-bc803f973a36 · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy Blip-diffusion: Pre-trained subject representation for controllable text-to-image generation and editing
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation ae29a190-bdb5-4310-9852-4da3f427fe7c · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy Onecat: Decoder-only auto-regressive model for unified understanding and generation
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 5ae2595d-4e4f-4bdc-8a6e-1eb41e5b8841 · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy Mvbench: A comprehensive multi-modal video understanding benchmark
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 2ec55ac7-abec-463a-862a-d06f8bc4793a · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy Videochat: Chat-centric video understanding.Science China Information Sciences, 68(10):200102
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 9432ac7a-6425-4881-acf3-0704c5976927 · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy Autoregressive image generation without vector quantization.Advancesin Neural Information Processing Systems, 37:56424–56445
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 5af6ba44-4ffa-4bc4-afed-bcf0c6fd1524 · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 9a2ba4bb-835b-4156-b3ec-e3e251c52eb0 · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 302463f0-8da9-4e90-a7d7-c0fd2eb37a62 · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy Video-llava: Learning united visual representation by alignment before projection
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 4a622e07-4db7-4f79-a407-80f213d6b71b · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 26cb367b-8230-424d-a95f-e644d662139d · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy Realgeneral: Unifying visual generation via temporal in-context learning with video models
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation c5ebd359-6ad9-450e-803b-c3b7c4e26826 · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy Flow Matching Guide and Code
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation d3083b19-dcad-4494-a176-49a73d4c1d9c · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy DeepSeek-V3 Technical Report
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 94646241-b5af-4aa6-9c8c-8810edf43c51 · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy World Model on Million-Length Video And Language With Blockwise RingAttention
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 78b334a4-5717-490c-a7d4-db9d56b1475a · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 49cafb6d-6686-4c3d-b18b-ab9f62b6bb35 · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy Improved baselines with visual instruction tuning
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 0078a322-9d15-4eb2-826e-03597d4f9368 · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy Llavanext: Improved reasoning, ocr, and world knowledge
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 534419d3-5744-4d85-99f0-0f81aa9bbb25 · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy MarDini: Masked Autoregressive Diffusion for Video Generation at Scale
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation d6ffab13-b578-4a3f-a279-ded3546a1c17 · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy Flow-grpo: Training flow matching models via online rl.Advances in neural information processing systems, 38:40783–40818
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 06a61fe5-935b-4056-8de4-2b86806d98b2 · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy St-llm: Large language models are effective temporal learners
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation e0652860-8976-42a3-9e0f-0419f411aa3f · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy Step1X-Edit: A Practical Framework for General Image Editing
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 331d4a64-a009-48eb-b850-b00ba6c6420c · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy Tuna: Taming unified visual representations for native unified multimodal models
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 1841f1f9-3efe-45ae-ad77-4ed7a70d1b87 · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 8e42cb5d-e659-4374-9e79-958876ff0e86 · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation b002c5a8-d3da-461c-b753-7586e3b20a43 · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy Janusflow: Harmonizing autoregression and rectified flow for unified multimodal understanding and generation
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation ef678626-0e70-4da0-9fe6-fec4d66b39f7 · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy Video-chatgpt: Towards detailed video understanding via large vision and language models
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 3c17b115-17fc-459a-a86b-ee2a4091c5c8 · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy Realcustom++: Representing images as real-word for real-time customization
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation a235a0e1-1aff-400d-b73e-b175301bd075 · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy Realcustom++: Representing images as real textual word for real-time customization.IEEE Transactions on Pattern Analysis and Machine Intelligence
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation eca066d7-229f-49b4-a675-14467d6013f3 · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy Toward accurate image generation via dynamic generative image transformer.IEEE Transactions on Pattern Analysis and Machine Intelligence
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 038933ab-0750-4bee-b5cc-cbb8e10582fe · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy Dreamo: A unified framework for image customization
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 1c1f9a08-ed54-4c58-8ee1-17c898347174 · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy Gpt-4v(ision) system card
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 4e4dc0d2-be2d-407d-ba3a-aa35b69519c3 · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy Introducing 4o image generation.https://openai.com/index/introducing-4o-image-generation/
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation e1a31679-76f4-4a61-ad87-963730dc4738 · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy Unresolved cited work
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation b945f899-6701-408c-9add-6365eeb1cc75 · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy Transfer between Modalities with MetaQueries
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation f4c5776c-7830-42a5-bc9a-3b6f69971c12 · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy Scalable diffusion models with transformers
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation cc3b0bd5-9e00-4cc4-b811-955a4a1ac514 · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy Open-Sora 2.0: Training a Commercial-Level Video Generation Model in $200k
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 3ff34f9f-16da-4616-95a8-b8c900f63293 · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy SDXL: Improving latent diffusion models for high-resolution image synthesis
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 2617c859-5f6d-400f-9b06-351b2420bfab · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy Tokenflow: Unified image tokenizer for multimodal understanding and generation
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation bcf12b67-f523-430b-8bcd-2f1497c1a7f7 · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy Learning transferable visual models from natural language supervision
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 83e70703-ff65-4600-80cd-f9200197a6a1 · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy Zero-shot text-to-image generation
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 8e747262-4032-4002-872e-a828b659cdff · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy High-resolution image synthesis with latent diffusion models
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 7f56b735-27da-4771-a64a-4d282f3fcbab · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy Introducing gen-3 alpha: A new frontier for video generation.https://runwayml.com/research/ introducing-gen-3-alpha, June 2024
Reference 98
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 62f27ca6-ab3f-4be2-b2d4-9c0efd6cab93 · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy Seedance 2.0: Advancing Video Generation for World Complexity
Reference 99
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 2d519c2c-e30a-484f-9084-8cf46ab2c137 · outbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy Seedream 4.0: Toward Next-generation Multimodal Image Generation
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation cd6a6a09-d9f0-4481-adb5-d8c4ca9297f9 · inbound
Toward Native Multimodal Modeling: A Roadmap Lance: Unified Multimodal Modeling by Multi-Task Synergy
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation e4a15058-b41f-4fa0-9ef2-a7d55f221d97 · inbound
S1-Omni-Image: A Unified Model for Scientific Image Understanding, Generation, and Editing Lance: Unified Multimodal Modeling by Multi-Task Synergy
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation efb55953-d674-4c7c-92f9-2ca0dea6b9aa · inbound
HarmVideoBench: Benchmarking Harmful Video Understanding in Large Multimodal Models Lance: Unified Multimodal Modeling by Multi-Task Synergy
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 48f4cefe-1474-4347-8733-e1684dc7ccbc · inbound
HarmVideoBench: Benchmarking Harmful Video Understanding in Large Multimodal Models Lance: Unified Multimodal Modeling by Multi-Task Synergy
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation d2c65a08-fd6b-4f1f-bbf4-a6baa1149595 · inbound
MoRoute: Dynamic Routing for In-Context Multimodal Video Generation Lance: Unified Multimodal Modeling by Multi-Task Synergy
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87316730-cb9d-464f-b5ad-91b868596498 · inbound
Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes Lance: Unified Multimodal Modeling by Multi-Task Synergy
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ed3445c-2933-41fe-91ec-e14877428f08 · inbound
Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes Lance: Unified Multimodal Modeling by Multi-Task Synergy
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 569f7263-26fc-45b3-8025-46d8129b94ac · inbound
Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Lance: Unified Multimodal Modeling by Multi-Task Synergy
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6589330e-2668-4861-94ce-e2b11e0bead1 · inbound
HounsWorld: A Multimodal World Model for Hidden Patient-State Readout, Reconstruction, and Simulation Lance: Unified Multimodal Modeling by Multi-Task Synergy
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.