Pith. sign in

Paper Citation Record · LEDGER

Lance: Unified Multimodal Modeling by Multi-Task Synergy

As of 22 August 2026, this Paper Citation Record lists 100 of 151 outbound references and 9 inbound Pith citation observations for arXiv:2605.18678.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.18678 v2

Coverage vector

measured 100 of 151 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-21T07:56:34.034047Z

measured 109 of 109 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:17:33.562911Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

100 of 151 outbound references displayed

  • verified exact48
  • verified fuzzy51
  • unresolved0
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 76f04b48-247a-4110-827c-7c13582eb821 · outbound

This paper cites GPT-4 Technical Report.

Lance: Unified Multimodal Modeling by Multi-Task Synergy GPT-4 Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:59:50.564591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:3d3e2c113c1a50cd7ceadf30b9276f38835c00799277869880277b236562ea4a

Observation 995c1c97-5af1-4b03-bae1-01630cf359bd · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

Lance: Unified Multimodal Modeling by Multi-Task Synergy Flamingo: a visual language model for few-shot learning

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T08:14:52.355723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:82557f67774651f10d2bb8d60594e96dc4ff9ffe22d026bc8a6e3a9a7c38c51a

Observation 19b9e6fc-8593-4b9d-aa26-c6f3fcacc501 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Lance: Unified Multimodal Modeling by Multi-Task Synergy Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:59:50.691656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:9fa76215752c9dafb73f7e23ebfcb1d09140448abe84f4f7e359b9c84c66b890

Observation b9956c71-7feb-46b8-92ad-0445151d940d · outbound

This paper cites Qwen3-VL Technical Report.

Lance: Unified Multimodal Modeling by Multi-Task Synergy Qwen3-VL Technical Report

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:59:50.583873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:7519ffa6cd10bc5dde5887a76b4f6301aefa323d59a44fb879d6813df1965611

Observation 58cdf9be-a587-42b3-955d-9ac8e02e3d84 · outbound

This paper cites Qwen2.5-VL Technical Report.

Lance: Unified Multimodal Modeling by Multi-Task Synergy Qwen2.5-VL Technical Report

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:59:50.720587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:d847033c276d9721eac7c939becdb847862c320c14d14e4bda6ecd8e08222937

Observation df09cb3e-1d81-4d59-84a3-c5e1a487817f · outbound

This paper cites Improving image generation with better captions.Computer Science.

Lance: Unified Multimodal Modeling by Multi-Task Synergy Improving image generation with better captions.Computer Science

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T08:14:52.352672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:7605ac550e858fb481ab08b20848e2779ff958f105e0c1eae84adff5f252d53d

Observation bba7fdbd-3482-4775-b6ba-073f017e7ef1 · outbound

This paper cites Diffusion Self-Distillation for Zero-Shot Customized Image Generation.

Lance: Unified Multimodal Modeling by Multi-Task Synergy Diffusion Self-Distillation for Zero-Shot Customized Image Generation

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:59:50.759531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:e7e4ce52296438ae952faf6f441bc77e4ba713746c93d5b0a63a97512c0ceab8

Observation 17f51a29-40c0-43d4-898c-1482aa940cf2 · outbound

This paper cites HunyuanImage 3.0 Technical Report.

Lance: Unified Multimodal Modeling by Multi-Task Synergy HunyuanImage 3.0 Technical Report

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:59:50.574885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:4e4618a328e5add07a00caabdb51565629cf629152dcecb709ab9e5000d04547

Observation 9a3e0a54-70b4-4614-b7d5-739b5fea1eb9 · outbound

This paper cites Maskgit: Masked generative image transformer.

Lance: Unified Multimodal Modeling by Multi-Task Synergy Maskgit: Masked generative image transformer

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T08:14:52.332264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:fe0a2718ff91862e1406cce09339cf4a182bc08392682e68b00ff4cbec07ae83

Observation 7e478804-4605-41bb-93b6-ca5866600e8a · outbound

This paper cites Videocrafter2: Overcoming data limitations for high-quality video diffusion models.

Lance: Unified Multimodal Modeling by Multi-Task Synergy Videocrafter2: Overcoming data limitations for high-quality video diffusion models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T08:14:52.339263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:b40d26d62d815e7243515c06291890a7afc6a62fa4aeb131f1b6bc6902beff60

Observation a7da9052-8c29-4148-8e9d-8da1f203e3c9 · outbound

This paper cites BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset.

Lance: Unified Multimodal Modeling by Multi-Task Synergy BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:59:50.580858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:42c27a8e625b282d66fd48cf0a30312b8b3cd5403a54acc1a8fb5e2eaca4c5b3

Observation b0bc4d67-3364-4bac-ad67-bd0102cb5a3c · outbound

This paper cites Pixart-σ: Weak-to-strong training of diffusion transformer for 4k text-to-image generation.

Lance: Unified Multimodal Modeling by Multi-Task Synergy Pixart-σ: Weak-to-strong training of diffusion transformer for 4k text-to-image generation

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T08:14:52.343643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:9b791aedc8e83d601591cd898761bcd154a6f61fb73c703ba36a1ca776ca5f90

Observation e81e7832-46d1-4918-8585-db1038c7d55e · outbound

This paper cites TimeMarker: A Versatile Video-LLM for Long and Short Video Understanding with Superior Temporal Localization Ability.

Lance: Unified Multimodal Modeling by Multi-Task Synergy TimeMarker: A Versatile Video-LLM for Long and Short Video Understanding with Superior Temporal Localization Ability

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:59:50.479363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:12bf359d01bb2e6a3cc8642996d6959bf3d6a82bde5dd33bb30b645d3090691b

Observation 611133b9-3f0a-41af-972c-c23a0b138a8b · outbound

This paper cites Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling.

Lance: Unified Multimodal Modeling by Multi-Task Synergy Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:59:50.488907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:99f3f4dc0d10100f241cbbcb3a0a5b8f695b0e1a7876f0c69d734d96da5d01ba

Observation 00123fe2-4cba-471b-ab1e-7b52c41676d9 · outbound

This paper cites How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites.

Lance: Unified Multimodal Modeling by Multi-Task Synergy How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T08:14:52.323032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:14e2f30b422e73f722ae060e76877243ada80057931674a67b2d1ab39d4f9776

Observation 3a05cbe2-d270-4427-8349-e48014c6e614 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

Lance: Unified Multimodal Modeling by Multi-Task Synergy Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T08:14:52.325322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:9e789beb286ffc9b238b13ea4062201d4c1f58465d79c814d4567c4de516104f

Observation 94e83882-5ab2-49c1-a065-6d4a07dc09ec · outbound

This paper cites UMO: Scaling Multi-Identity Consistency for Image Customization via Matching Reward.

Lance: Unified Multimodal Modeling by Multi-Task Synergy UMO: Scaling Multi-Identity Consistency for Image Customization via Matching Reward

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:59:50.448844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:3062f9877676427275107502830c2a358667a3209becdc3e5fd0ad797c384399

Observation a5d1dbaf-8baa-427d-895d-a97bef0a3d2a · outbound

This paper cites PaddleOCR 3.0 Technical Report.

Lance: Unified Multimodal Modeling by Multi-Task Synergy PaddleOCR 3.0 Technical Report

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:59:50.694759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:6b0fa89106522179badf50bab28f20a70373dd77634ae9cca61442e854b481e8

Observation 9813bfcc-dc4c-435b-92d1-de6b9fce24e2 · outbound

This paper cites Emu3.5: Native Multimodal Models are World Learners.

Lance: Unified Multimodal Modeling by Multi-Task Synergy Emu3.5: Native Multimodal Models are World Learners

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:59:50.460568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:225871ec987ebf5fd71c4c052701061d91219ddd22654c4ceb99cb47de513342

Observation 73f0a0e9-3b2c-4d3b-ab98-054ea216cbfa · outbound

This paper cites Instructblip: Towards general-purpose vision-language models with instruction tuning.

Lance: Unified Multimodal Modeling by Multi-Task Synergy Instructblip: Towards general-purpose vision-language models with instruction tuning

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T08:14:52.345884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:314a8e185df75768583c05953ef88d3aba4c2934d7b82d8e832d6bbcb30df70d

Observation 02508724-e637-43cf-b990-f1efea13a798 · outbound

This paper cites ChatUMM: Robust Context Tracking for Conversational Interleaved Generation.

Lance: Unified Multimodal Modeling by Multi-Task Synergy ChatUMM: Robust Context Tracking for Conversational Interleaved Generation

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-06-02T04:04:28.060339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:7b3f5143968d070bdc5439511777fa10ab71322f41f7b5b361803b997b3c994b

Observation 1fb140fb-cb35-449a-8f1d-615d3fd193ef · outbound

This paper cites Scaling vision transformers to 22 billion parameters.

Lance: Unified Multimodal Modeling by Multi-Task Synergy Scaling vision transformers to 22 billion parameters

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T08:14:52.348002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:b7c3dc28f78b4b412ddff804905cec4c0ff6c0829f8f4e987ad198a3de9032f8

Observation 48a508c7-b0f0-4546-aa2a-62fb452f7ce1 · outbound

This paper cites Emerging Properties in Unified Multimodal Pretraining.

Lance: Unified Multimodal Modeling by Multi-Task Synergy Emerging Properties in Unified Multimodal Pretraining

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:59:50.767254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:675658a7b5e2ec5827b43785f34fa33edf3f1c4f26cb62571dfaf757b369bcf2

Observation 7fd5e614-ce13-465f-85cd-b7f508f632e5 · outbound

This paper cites Cogview: Mastering text-to-image generation via transformers.NIPS, 34:19822–19835.

Lance: Unified Multimodal Modeling by Multi-Task Synergy Cogview: Mastering text-to-image generation via transformers.NIPS, 34:19822–19835

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T08:14:52.294713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:c22cda118296f6c700e153aeeeef698bb59a136b9a1444c28b307297512b51e6

Observation 02a9cb21-26e2-45a1-916c-69d3f140e8ce · outbound

This paper cites Taming transformers for high-resolution image synthesis.

Lance: Unified Multimodal Modeling by Multi-Task Synergy Taming transformers for high-resolution image synthesis

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T08:14:52.289983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:0775d6d8c4f84b0d1114b9b93cf736c06e0dfd8e9537d9a646d9411d6f230687

Observation ede6043b-ad4b-4bbd-9cba-91eedc485ce2 · outbound

This paper cites Scaling rectified flow transformers for high-resolution image synthesis.

Lance: Unified Multimodal Modeling by Multi-Task Synergy Scaling rectified flow transformers for high-resolution image synthesis

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T08:14:52.341546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:26a369fcfd202fb4912692b9bee9116a38eac3a37e946625997829205617836e

Observation d5172ce2-5713-4439-8ec6-0c928cda97c6 · outbound

This paper cites Unified Autoregressive Visual Generation and Understanding with Continuous Tokens.

Lance: Unified Multimodal Modeling by Multi-Task Synergy Unified Autoregressive Visual Generation and Understanding with Continuous Tokens

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:59:50.752143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:90cb984302430a8da379a1c3f4e2afe4ea71da1cd6e00b7b3c424137dafebfb8

Observation 60c46325-5641-4837-a59b-d9537c6f88d3 · outbound

This paper cites Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos.

Lance: Unified Multimodal Modeling by Multi-Task Synergy Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:59:50.748177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:277148ad10d69ed8c5bf382b44a931438f9dbac431511ec936908aa29ff0956b

Observation b5f3473d-104f-479a-a975-1291945aca4a · outbound

This paper cites Dreamlite: A lightweight on-device unified model for image generation and editing.

Lance: Unified Multimodal Modeling by Multi-Task Synergy Dreamlite: A lightweight on-device unified model for image generation and editing

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:59:50.738528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:a28be2a9c397dd640d1c4284c09525e3a4ed8893d3b0fff24b3e72b0bd3bcd54

Observation 292da351-5d4b-44ff-a72d-ee1556019aad · outbound

This paper cites Feededit: Text-based image editing with dynamic feedback regulation.

Lance: Unified Multimodal Modeling by Multi-Task Synergy Feededit: Text-based image editing with dynamic feedback regulation

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T08:14:52.275654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:6d782c176ac15df26906604504ea5e4f1369241b8672157db08bae872b340a1c

Observation faa5f6d2-ace7-41a3-bb0d-bfeefba198f0 · outbound

This paper cites Layeredit: Disentangled multi-object editing via conflict-aware multi-layer learning.

Lance: Unified Multimodal Modeling by Multi-Task Synergy Layeredit: Disentangled multi-object editing via conflict-aware multi-layer learning

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T08:14:52.268752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:979279ac152decfd1f061e0df12f902556a7b25f3e82576df02c3133bb6e556f

Observation 6f7d5ead-01ae-4faf-a593-de9c9a5619cd · outbound

This paper cites Mini-internvl: a flexible-transfer pocket multi-modal model with 5% parameters and 90% performance.

Lance: Unified Multimodal Modeling by Multi-Task Synergy Mini-internvl: a flexible-transfer pocket multi-modal model with 5% parameters and 90% performance

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T08:14:52.266580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:7c3622adcb6685a6193a603fa8f2c160deb648a3de59f2bd640596768138f567

Observation 1edc4828-c9a1-4244-acf0-7d70a78b3698 · outbound

This paper cites SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation.

Lance: Unified Multimodal Modeling by Multi-Task Synergy SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:59:50.734734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:d4fa5e20c14670835b545e1056c2d777b6971e3f87fb85f01b705096719859be

Observation 7fa33aea-fdaa-43b7-996a-1dfadf59551e · outbound

This paper cites Geneval: An object-focused framework for evaluating text-to-image alignment.Advancesin Neural Information Processing Systems, 36:52132–52152.

Lance: Unified Multimodal Modeling by Multi-Task Synergy Geneval: An object-focused framework for evaluating text-to-image alignment.Advancesin Neural Information Processing Systems, 36:52132–52152

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T08:14:52.216706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:06a621c0f1d1bac817c7436fb601ad00cf4b4229607cbab07c4f7baee02e059b

Observation d40669be-151e-4975-b21b-2fca56e80afd · outbound

This paper cites Gemini 3 Pro Image Model Card.

Lance: Unified Multimodal Modeling by Multi-Task Synergy Gemini 3 Pro Image Model Card

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T08:14:52.336859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:66ff975bdef0555e57afd1712bf5b2d3471261e8b51965e7a8f315e2cecd5aec

Observation 91d1bb12-0a51-463b-9cb1-d48933ed202a · outbound

This paper cites AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning.

Lance: Unified Multimodal Modeling by Multi-Task Synergy AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:59:50.731308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:91b8b57788b71cf1dd8b474cea0b7081d26fd517498c1a6e970dce7dfa565467

Observation be4499ca-34fa-4c9a-84c8-9a360a5777c7 · outbound

This paper cites Tv2tv: A unified framework for interleaved language and video generation.

Lance: Unified Multimodal Modeling by Multi-Task Synergy Tv2tv: A unified framework for interleaved language and video generation

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:59:50.743666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:a9f2813f4aaf83f6c31a2f822404faa020716d7b3710741824cb0593dd77052b

Observation 7d918d47-cfee-4b84-8f9e-cc94aa5403cd · outbound

This paper cites Emma: Efficient multimodal understanding, generation, and editing with a unified architecture.

Lance: Unified Multimodal Modeling by Multi-Task Synergy Emma: Efficient multimodal understanding, generation, and editing with a unified architecture

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:59:50.755818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:0f4ac4cc00e35187f8630671391b1586eb6a2fef3df2cd3da773aad14035a199

Observation 35cd92d7-e045-4e21-9fe5-a80380e5af89 · outbound

This paper cites Classifier-Free Diffusion Guidance.

Lance: Unified Multimodal Modeling by Multi-Task Synergy Classifier-Free Diffusion Guidance

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:59:50.610876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:e998dc88e5bbe5ba1db40ef06a06055482f87f6b252a1d5c102ca1cef9d9c732

Observation a6143f1d-c9f3-4681-b171-8e2b9e1a30b4 · outbound

This paper cites Denoising diffusion probabilistic models.NIPS, 33:6840–6851.

Lance: Unified Multimodal Modeling by Multi-Task Synergy Denoising diffusion probabilistic models.NIPS, 33:6840–6851

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T08:14:52.211787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:c43e317ef3dcca8e83eef6ce6bf7355334932e400972023e2f4c66ee548f1d2e

Observation f964116d-5729-4333-813b-ba0bf40e55ce · outbound

This paper cites CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers.

Lance: Unified Multimodal Modeling by Multi-Task Synergy CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:59:50.728150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:82cb5b02f38952a04733dc68427e2b3860bd47bafdada27afd2d85a3b1f9d6b4

Observation 663963de-e5fd-4d60-9076-f1f594dce0ce · outbound

This paper cites ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment.

Lance: Unified Multimodal Modeling by Multi-Task Synergy ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:59:50.717242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:e5bf203c5c030f0f4b6ade57e8ab39320eed9434fd216fa4a720dc7e65564124

Observation da9ec087-5f62-47b2-b537-7de764840ed1 · outbound

This paper cites Dse-gan: Dynamic semantic evolution generative adversarial network for text-to-image generation.

Lance: Unified Multimodal Modeling by Multi-Task Synergy Dse-gan: Dynamic semantic evolution generative adversarial network for text-to-image generation

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T08:14:52.214275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:054ae50b6f6a4d894f6ff77e95f2bbd44f268ee65f7dc9c190df274d92d36f1c

Observation eb69059c-3617-4839-ae24-2b70e7937980 · outbound

This paper cites Towards accurate image coding: Improved autoregressive image generation with dynamic vector quantization.

Lance: Unified Multimodal Modeling by Multi-Task Synergy Towards accurate image coding: Improved autoregressive image generation with dynamic vector quantization

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T08:14:52.254568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:b6a05c91b408bb8c2e257831b4d75e2496c6306c337d68b88c3110bccb9233f2

Observation a2d11953-8a1e-4a56-bffc-037cbb247d59 · outbound

This paper cites Realcustom: Narrowing real text word for real-time open-domain text-to-image customization.

Lance: Unified Multimodal Modeling by Multi-Task Synergy Realcustom: Narrowing real text word for real-time open-domain text-to-image customization

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T08:14:52.219234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:583f7428fd73a527f5209f50d697084fb6b90d12b7a485f9880b1558d03d3286

Observation 07d9c40a-6b4f-4272-8c65-171a4b918770 · outbound

This paper cites Self forcing: Bridging the train-test gap in autoregressive video diffusion.Advances in Neural Information Processing Systems, 38:167283–167308.

Lance: Unified Multimodal Modeling by Multi-Task Synergy Self forcing: Bridging the train-test gap in autoregressive video diffusion.Advances in Neural Information Processing Systems, 38:167283–167308

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T08:14:52.209335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:7aa8cb1d21f9b2c375114b34e87990d72af7e0ea1b825e6e5f7f1061b3f9d2db

Observation 386afc2f-968d-41e7-a1d1-794e78a54dd1 · outbound

This paper cites Vbench: Comprehensive benchmark suite for video generative models.

Lance: Unified Multimodal Modeling by Multi-Task Synergy Vbench: Comprehensive benchmark suite for video generative models

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T08:14:52.278099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:2f0f1e405e2deb1bdd2619e978db2d39c70ac6d81e30a49c530bc3337652ad3f

Observation a6fffe26-aa79-4e9d-8108-a06aaf7e0ad9 · outbound

This paper cites Vace: All-in-one video creation and editing.

Lance: Unified Multimodal Modeling by Multi-Task Synergy Vace: All-in-one video creation and editing

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T08:14:52.327433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:4b4d79d23bd6011dc3e8fda47baccf51bf7afd52ee60621e2e6ff8736305a734

Observation d9baa9e7-5671-43c7-bbaa-3d1cbcb8f471 · outbound

This paper cites EditVerse: Unifying Image and Video Editing and Generation with In-Context Learning.

Lance: Unified Multimodal Modeling by Multi-Task Synergy EditVerse: Unifying Image and Video Editing and Generation with In-Context Learning

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:59:50.706658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:5b550b8716222ae6fc028a08a78d5f4eff2aeabfb2db29f2c8bc821c6dc56458

Observation 3c03f3d8-85e4-441d-9848-c4352dbde82c · outbound

This paper cites FullDiT: Multi-Task Video Generative Foundation Model with Full Attention.

Lance: Unified Multimodal Modeling by Multi-Task Synergy FullDiT: Multi-Task Video Generative Foundation Model with Full Attention

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:59:50.710354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:139e207d7a4bf499937e18fdf90629783f9acd7bca6ef1ab17757b59d10848a1

Observation 45936d82-8f83-44b6-8311-49d512248324 · outbound

This paper cites Kling ai.https://klingai.kuaishou.com/.

Lance: Unified Multimodal Modeling by Multi-Task Synergy Kling ai.https://klingai.kuaishou.com/

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T08:14:52.235954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:8a1eabb0c69e2becd9f0e371b31543579eac90edf72b47b184942a109a4dddfb

Observation a5986faf-0494-466a-9eb0-a7170ce8fea6 · outbound

This paper cites VideoPoet: A Large Language Model for Zero-Shot Video Generation.

Lance: Unified Multimodal Modeling by Multi-Task Synergy VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:59:50.713727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:9f4d55239526ea4844c2afc9db58864bd9f26a64006a3340ce4059671725cc0d

Observation 5557a68d-39cc-45bc-8e98-f18b1d8cc77d · outbound

This paper cites AnyV2V: A Tuning-Free Framework For Any Video-to-Video Editing Tasks.

Lance: Unified Multimodal Modeling by Multi-Task Synergy AnyV2V: A Tuning-Free Framework For Any Video-to-Video Editing Tasks

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:59:50.724598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:7945c31d26b736e4b3626c2d6e134bed27ebb891d6f710ac005c395b11dbb4f3

Observation 93676173-f707-4a8c-8f4a-0de9b130506e · outbound

This paper cites Flux: Official inference repository for flux.1 models.

Lance: Unified Multimodal Modeling by Multi-Task Synergy Flux: Official inference repository for flux.1 models

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T08:14:52.238117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:a1b10ddcd33623d828cbc0c6b0806951399c4f80ebaf13213646b75f690c9dac

Observation ad3a9c14-1654-40d4-abda-6ef04d860b96 · outbound

This paper cites FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space.

Lance: Unified Multimodal Modeling by Multi-Task Synergy FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:59:50.762545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:0c4de1244c2f3a0dcf3db3ec0c3ca69d101c46d266f0ed727769b4a3cc28697b

Observation a509b74f-bebe-44cb-9bf6-eb63dcc5641f · outbound

This paper cites Obelics: An open web-scale filtered dataset of interleaved image-text documents.Advancesin Neural Information Processing Systems, 36:71683–71702.

Lance: Unified Multimodal Modeling by Multi-Task Synergy Obelics: An open web-scale filtered dataset of interleaved image-text documents.Advancesin Neural Information Processing Systems, 36:71683–71702

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T08:14:52.247695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:dd63c3898342d4f8d04bdbdf9230788a71b6c84f063705fbe05c04845868fbae

Observation bbd74472-1af9-475e-8d86-9ca007f2e93c · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Lance: Unified Multimodal Modeling by Multi-Task Synergy LLaVA-OneVision: Easy Visual Task Transfer

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:59:50.688079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:f3d70c4c32b9370c846c328ec6ed5932b4cf095d9b7b06c207c689ad97636ccc

Observation 29b32ebf-ba79-417c-88f8-bc803f973a36 · outbound

This paper cites Blip-diffusion: Pre-trained subject representation for controllable text-to-image generation and editing.

Lance: Unified Multimodal Modeling by Multi-Task Synergy Blip-diffusion: Pre-trained subject representation for controllable text-to-image generation and editing

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T08:14:52.231355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:c7e461be894bbc9b553bce844312fc4b271a7ff1b1eefba9ccd201d90d082281

Observation ae29a190-bdb5-4310-9852-4da3f427fe7c · outbound

This paper cites Onecat: Decoder-only auto-regressive model for unified understanding and generation.

Lance: Unified Multimodal Modeling by Multi-Task Synergy Onecat: Decoder-only auto-regressive model for unified understanding and generation

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:59:50.445066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:756ce2ebe17b28047714a2dad9f961ac9ec26bf914313128854e3604b9d84b97

Observation 5ae2595d-4e4f-4bdc-8a6e-1eb41e5b8841 · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding benchmark.

Lance: Unified Multimodal Modeling by Multi-Task Synergy Mvbench: A comprehensive multi-modal video understanding benchmark

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T08:14:52.233836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:0d6f69f80f66b6a6553f36f664549718bc4dbedcf6f6d81476691414d61056e3

Observation 2ec55ac7-abec-463a-862a-d06f8bc4793a · outbound

This paper cites Videochat: Chat-centric video understanding.Science China Information Sciences, 68(10):200102.

Lance: Unified Multimodal Modeling by Multi-Task Synergy Videochat: Chat-centric video understanding.Science China Information Sciences, 68(10):200102

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T08:14:52.240248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:4f1a8d8ce5ec6348746a9ec42c256224c8f5fd78d4748fee98c20642af56b076

Observation 9432ac7a-6425-4881-acf3-0704c5976927 · outbound

This paper cites Autoregressive image generation without vector quantization.Advancesin Neural Information Processing Systems, 37:56424–56445.

Lance: Unified Multimodal Modeling by Multi-Task Synergy Autoregressive image generation without vector quantization.Advancesin Neural Information Processing Systems, 37:56424–56445

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T08:14:52.243040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:64745bceb5fd8783ba7af62b344a5be37fe2db554652357490a796c47aca39fb

Observation 5af6ba44-4ffa-4bc4-afed-bcf0c6fd1524 · outbound

This paper cites Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding.

Lance: Unified Multimodal Modeling by Multi-Task Synergy Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 63

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:59:50.676924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:502e3f68f3295eedbe3d59c3f24e33e153713838338ccac9381de83b3735723c

Observation 9a2ba4bb-835b-4156-b3ec-e3e251c52eb0 · outbound

This paper cites Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation.

Lance: Unified Multimodal Modeling by Multi-Task Synergy Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation

Reference 64

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:59:50.684850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:fb4caf3b9b1047ae565a6415b1a67c7e0687ef57caef2ebba065e9519817d138

Observation 302463f0-8da9-4e90-a7d7-c0fd2eb37a62 · outbound

This paper cites Video-llava: Learning united visual representation by alignment before projection.

Lance: Unified Multimodal Modeling by Multi-Task Synergy Video-llava: Learning united visual representation by alignment before projection

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T08:14:52.259013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:b823cfe56253329b756dfe3c443249cccbde257f7616b98f8d1ccae31946fa83

Observation 4a622e07-4db7-4f79-a407-80f213d6b71b · outbound

This paper cites UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation.

Lance: Unified Multimodal Modeling by Multi-Task Synergy UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation

Reference 66

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:59:50.680120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:b8bab7a9a0fd0c6e438b00d5c85e03af87975448ecb0aae6a8ce33a21d6b2e37

Observation 26cb367b-8230-424d-a95f-e644d662139d · outbound

This paper cites Realgeneral: Unifying visual generation via temporal in-context learning with video models.

Lance: Unified Multimodal Modeling by Multi-Task Synergy Realgeneral: Unifying visual generation via temporal in-context learning with video models

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T08:14:52.261505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:0639f80ce45e930e4b60fa76cea6db532050aace6724a2bab87e95cbe7719e55

Observation c5ebd359-6ad9-450e-803b-c3b7c4e26826 · outbound

This paper cites Flow Matching Guide and Code.

Lance: Unified Multimodal Modeling by Multi-Task Synergy Flow Matching Guide and Code

Reference 68

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:59:50.660003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:b8186de9297fa59005b81a283e7015d56c4877b8d959d1b74343ba64c4af0acf

Observation d3083b19-dcad-4494-a176-49a73d4c1d9c · outbound

This paper cites DeepSeek-V3 Technical Report.

Lance: Unified Multimodal Modeling by Multi-Task Synergy DeepSeek-V3 Technical Report

Reference 69

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:59:50.655777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:d5f460f12884705573f3aafac45c0aa897df45bdf35d3d5878efa18f9a600835

Observation 94646241-b5af-4aa6-9c8c-8810edf43c51 · outbound

This paper cites World Model on Million-Length Video And Language With Blockwise RingAttention.

Lance: Unified Multimodal Modeling by Multi-Task Synergy World Model on Million-Length Video And Language With Blockwise RingAttention

Reference 70

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:59:50.648173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:f40f7602594647bfe14def52360ca26a17d7756e04ba37a3f2cfc531b36be9bb

Observation 78b334a4-5717-490c-a7d4-db9d56b1475a · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916.

Lance: Unified Multimodal Modeling by Multi-Task Synergy Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T08:14:52.280489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:01bf936c0456d1aaa9df02f22815f34d10a01ddf2aa9043dce6305b4b9b985ec

Observation 49cafb6d-6686-4c3d-b18b-ab9f62b6bb35 · outbound

This paper cites Improved baselines with visual instruction tuning.

Lance: Unified Multimodal Modeling by Multi-Task Synergy Improved baselines with visual instruction tuning

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T08:14:52.315655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:d0eb81c3f89a4bebd39681c9b5857ea4e12edf036008aa9381c91dcf87c9444f

Observation 0078a322-9d15-4eb2-826e-03597d4f9368 · outbound

This paper cites Llavanext: Improved reasoning, ocr, and world knowledge.

Lance: Unified Multimodal Modeling by Multi-Task Synergy Llavanext: Improved reasoning, ocr, and world knowledge

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T08:14:52.313485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:686ed7d1cea45764f16e37f88f87d279a0eedc4a5b740b3a2050f1d771d5399c

Observation 534419d3-5744-4d85-99f0-0f81aa9bbb25 · outbound

This paper cites MarDini: Masked Autoregressive Diffusion for Video Generation at Scale.

Lance: Unified Multimodal Modeling by Multi-Task Synergy MarDini: Masked Autoregressive Diffusion for Video Generation at Scale

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:59:50.644203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:4f3e7e39227c78cc022cc447fd6cf1983afe8054cc716ea4e148177bd1448714

Observation d6ffab13-b578-4a3f-a279-ded3546a1c17 · outbound

This paper cites Flow-grpo: Training flow matching models via online rl.Advances in neural information processing systems, 38:40783–40818.

Lance: Unified Multimodal Modeling by Multi-Task Synergy Flow-grpo: Training flow matching models via online rl.Advances in neural information processing systems, 38:40783–40818

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T08:14:52.306031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:53eb6e71c8ff1464097703565d1d5b4eb81485b9250cff1d1a27496f2b623c1d

Observation 06a61fe5-935b-4056-8de4-2b86806d98b2 · outbound

This paper cites St-llm: Large language models are effective temporal learners.

Lance: Unified Multimodal Modeling by Multi-Task Synergy St-llm: Large language models are effective temporal learners

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T08:14:52.270901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:0aee5a1c654cc1b2bf0c757080bc274610b30afab6534724e2a74e1e9c2bddc6

Observation e0652860-8976-42a3-9e0f-0419f411aa3f · outbound

This paper cites Step1X-Edit: A Practical Framework for General Image Editing.

Lance: Unified Multimodal Modeling by Multi-Task Synergy Step1X-Edit: A Practical Framework for General Image Editing

Reference 77

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:59:50.651883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:b9269cf904adc9e11dc627044ceb2881baa88e15c3b43d11c8df7048b455f53d

Observation 331d4a64-a009-48eb-b850-b00ba6c6420c · outbound

This paper cites Tuna: Taming unified visual representations for native unified multimodal models.

Lance: Unified Multimodal Modeling by Multi-Task Synergy Tuna: Taming unified visual representations for native unified multimodal models

Reference 78

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:59:50.668101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:0f75a672da78ea4768184c689b17378e6e3b4ce77f1d1ded8b168af01b903ba5

Observation 1841f1f9-3efe-45ae-ad77-4ed7a70d1b87 · outbound

This paper cites Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation.

Lance: Unified Multimodal Modeling by Multi-Task Synergy Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation

Reference 79

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:59:50.672370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:52c1910bc6b4588d8a60d888bffa87b7b6cd6c1c557f2c1b31179230b6d55022

Observation 8e42cb5d-e659-4374-9e79-958876ff0e86 · outbound

This paper cites Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model.

Lance: Unified Multimodal Modeling by Multi-Task Synergy Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model

Reference 80

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:59:50.631393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:7fd89bae5536ce1aaaf24e71be9cba12b345b7bbf0c9317e6db88b6f5567b162

Observation b002c5a8-d3da-461c-b753-7586e3b20a43 · outbound

This paper cites Janusflow: Harmonizing autoregression and rectified flow for unified multimodal understanding and generation.

Lance: Unified Multimodal Modeling by Multi-Task Synergy Janusflow: Harmonizing autoregression and rectified flow for unified multimodal understanding and generation

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T08:14:52.303758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:e6ca3602b525dfc23fc58f2d71cfd492a2236bae4ac1c35a98007d981720d993

Observation ef678626-0e70-4da0-9fe6-fec4d66b39f7 · outbound

This paper cites Video-chatgpt: Towards detailed video understanding via large vision and language models.

Lance: Unified Multimodal Modeling by Multi-Task Synergy Video-chatgpt: Towards detailed video understanding via large vision and language models

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T08:14:52.308326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:3ee602fe41037634e80f0c7646555550547f14c4162bb95debd2954e5259c0be

Observation 3c17b115-17fc-459a-a86b-ee2a4091c5c8 · outbound

This paper cites Realcustom++: Representing images as real-word for real-time customization.

Lance: Unified Multimodal Modeling by Multi-Task Synergy Realcustom++: Representing images as real-word for real-time customization

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:59:50.779597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:556ab17dcf1a7c7a1633ffb283e587818584c43a7a9aaafb624d76fab69bf39d

Observation a235a0e1-1aff-400d-b73e-b175301bd075 · outbound

This paper cites Realcustom++: Representing images as real textual word for real-time customization.IEEE Transactions on Pattern Analysis and Machine Intelligence.

Lance: Unified Multimodal Modeling by Multi-Task Synergy Realcustom++: Representing images as real textual word for real-time customization.IEEE Transactions on Pattern Analysis and Machine Intelligence

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T08:14:52.310998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:b6d761134f35816cb38c0b89928a976411ad52a9f433ef32f51817aadbfbff85

Observation eca066d7-229f-49b4-a675-14467d6013f3 · outbound

This paper cites Toward accurate image generation via dynamic generative image transformer.IEEE Transactions on Pattern Analysis and Machine Intelligence.

Lance: Unified Multimodal Modeling by Multi-Task Synergy Toward accurate image generation via dynamic generative image transformer.IEEE Transactions on Pattern Analysis and Machine Intelligence

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T08:14:52.318512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:cd957376192fa4f5233a0dd98ee46642a79fad48e4b3d7dec015e2ae2f332490

Observation 038933ab-0750-4bee-b5cc-cbb8e10582fe · outbound

This paper cites Dreamo: A unified framework for image customization.

Lance: Unified Multimodal Modeling by Multi-Task Synergy Dreamo: A unified framework for image customization

Reference 86

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:59:50.456285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:017e942df60a0933d098316b8bf79bd11d654e692f27ab125ed47140948310b9

Observation 1c1f9a08-ed54-4c58-8ee1-17c898347174 · outbound

This paper cites Gpt-4v(ision) system card.

Lance: Unified Multimodal Modeling by Multi-Task Synergy Gpt-4v(ision) system card

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T08:14:52.320972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:73f5f300b9d92c5ea9e020c1b14b25ae866366f2b9c30d439288dd72ae7e48cf

Observation 4e4dc0d2-be2d-407d-ba3a-aa35b69519c3 · outbound

This paper cites Introducing 4o image generation.https://openai.com/index/introducing-4o-image-generation/.

Lance: Unified Multimodal Modeling by Multi-Task Synergy Introducing 4o image generation.https://openai.com/index/introducing-4o-image-generation/

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T08:14:52.350436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:dd30b3185bad0b75cc48b2614644b5b7f95d8b0929f2eff6aee331ef542867de

Observation e1a31679-76f4-4a61-ad87-963730dc4738 · outbound

This paper cites an unresolved cited work.

Lance: Unified Multimodal Modeling by Multi-Task Synergy Unresolved cited work

Reference 89

Resolution
parse uncertain
raw_fallback, observed 2026-05-21T08:14:52.299175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:d6a5e6fa5a5cab0ce79c53ee737775ac3ea4c77563063ef4da595a7f3d475dbc

Observation b945f899-6701-408c-9add-6365eeb1cc75 · outbound

This paper cites Transfer between Modalities with MetaQueries.

Lance: Unified Multimodal Modeling by Multi-Task Synergy Transfer between Modalities with MetaQueries

Reference 90

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:59:50.627177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:f91aec72186cbb51b809a1c4ca1d16223b7870d85f34f6461cf3b4e17fb72f30

Observation f4c5776c-7830-42a5-bc9a-3b6f69971c12 · outbound

This paper cites Scalable diffusion models with transformers.

Lance: Unified Multimodal Modeling by Multi-Task Synergy Scalable diffusion models with transformers

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T08:14:52.301519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:d46a9060122a5cf955f1fe93910113163aa3c30fb27597a42db300117bd21f33

Observation cc3b0bd5-9e00-4cc4-b811-955a4a1ac514 · outbound

This paper cites Open-Sora 2.0: Training a Commercial-Level Video Generation Model in $200k.

Lance: Unified Multimodal Modeling by Multi-Task Synergy Open-Sora 2.0: Training a Commercial-Level Video Generation Model in $200k

Reference 92

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:59:50.623503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:f318de0815192c79e20ae2ae92eb4f805c79a069e0a2e40ad7c5d892f9bf0f22

Observation 3ff34f9f-16da-4616-95a8-b8c900f63293 · outbound

This paper cites SDXL: Improving latent diffusion models for high-resolution image synthesis.

Lance: Unified Multimodal Modeling by Multi-Task Synergy SDXL: Improving latent diffusion models for high-resolution image synthesis

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T08:14:52.334560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:205fea83114330e666dd67522c549590ca9c25182655482bef20e21b9cb008dc

Observation 2617c859-5f6d-400f-9b06-351b2420bfab · outbound

This paper cites Tokenflow: Unified image tokenizer for multimodal understanding and generation.

Lance: Unified Multimodal Modeling by Multi-Task Synergy Tokenflow: Unified image tokenizer for multimodal understanding and generation

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T08:14:52.282991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:9633148d51879412122b5d0575c5d99eef2b7e11e9c9ee92c7b96fd01680dfbc

Observation bcf12b67-f523-430b-8bcd-2f1497c1a7f7 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Lance: Unified Multimodal Modeling by Multi-Task Synergy Learning transferable visual models from natural language supervision

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T08:14:52.285355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:2f893f0cd2fc5cf3585384d89746cc978edbea3cb5856924f28e5e284244ac6f

Observation 83e70703-ff65-4600-80cd-f9200197a6a1 · outbound

This paper cites Zero-shot text-to-image generation.

Lance: Unified Multimodal Modeling by Multi-Task Synergy Zero-shot text-to-image generation

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T08:14:52.287492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:a2bdec6126729789881105e1cea0ed98568e3bafe6864d0a3d600ec2f6e37d7e

Observation 8e747262-4032-4002-872e-a828b659cdff · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Lance: Unified Multimodal Modeling by Multi-Task Synergy High-resolution image synthesis with latent diffusion models

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T08:14:52.297107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:69569d43e92ae9f857e76e62276c5af4fefa98fc4dd8052695ec9a0968f59ffb

Observation 7f56b735-27da-4771-a64a-4d282f3fcbab · outbound

This paper cites Introducing gen-3 alpha: A new frontier for video generation.https://runwayml.com/research/ introducing-gen-3-alpha, June 2024.

Lance: Unified Multimodal Modeling by Multi-Task Synergy Introducing gen-3 alpha: A new frontier for video generation.https://runwayml.com/research/ introducing-gen-3-alpha, June 2024

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T08:14:52.264193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:7d4aa5d26720c5bf65d855449aa830073e293acc4b48164f1dddb852598b7eca

Observation 62f27ca6-ab3f-4be2-b2d4-9c0efd6cab93 · outbound

This paper cites Seedance 2.0: Advancing Video Generation for World Complexity.

Lance: Unified Multimodal Modeling by Multi-Task Synergy Seedance 2.0: Advancing Video Generation for World Complexity

Reference 99

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:59:50.536388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:f9c1b830e9f26405337647a275312e4058ca0df4e5d05b684588e6cf9ac74084

Observation 2d519c2c-e30a-484f-9084-8cf46ab2c137 · outbound

This paper cites Seedream 4.0: Toward Next-generation Multimodal Image Generation.

Lance: Unified Multimodal Modeling by Multi-Task Synergy Seedream 4.0: Toward Next-generation Multimodal Image Generation

Reference 100

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:59:50.590246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:ea3f5b5dd78cf7ebd6df2e3a5df88ab10a0357c79d0444b539676b40092ec679

Pith citing papers

Observation cd6a6a09-d9f0-4481-adb5-d8c4ca9297f9 · inbound

Toward Native Multimodal Modeling: A Roadmap cites this paper.

Toward Native Multimodal Modeling: A Roadmap Lance: Unified Multimodal Modeling by Multi-Task Synergy

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-06-29T23:04:01.804804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T22:58:38.610609Z digest=sha256:637cf4379c49ec100e12baab9880ecddc4991c67256999f099da4062e3ee0737

Observation e4a15058-b41f-4fa0-9ef2-a7d55f221d97 · inbound

S1-Omni-Image: A Unified Model for Scientific Image Understanding, Generation, and Editing cites this paper.

S1-Omni-Image: A Unified Model for Scientific Image Understanding, Generation, and Editing Lance: Unified Multimodal Modeling by Multi-Task Synergy

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-04T16:39:58.068623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-26T00:20:42.801659Z digest=sha256:6a9ff041d478e499dbc645701c177ead01671abc413709ad7b038d3915e90e9f

Observation efb55953-d674-4c7c-92f9-2ca0dea6b9aa · inbound

HarmVideoBench: Benchmarking Harmful Video Understanding in Large Multimodal Models cites this paper.

HarmVideoBench: Benchmarking Harmful Video Understanding in Large Multimodal Models Lance: Unified Multimodal Modeling by Multi-Task Synergy

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-04T12:49:53.015032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-26T05:48:34.492354Z digest=sha256:6d775b53394f736d44bcf0e7798ae04bfd4a7887216415b2d43e354e5fa59b95

Observation 48f4cefe-1474-4347-8733-e1684dc7ccbc · inbound

HarmVideoBench: Benchmarking Harmful Video Understanding in Large Multimodal Models cites this paper.

HarmVideoBench: Benchmarking Harmful Video Understanding in Large Multimodal Models Lance: Unified Multimodal Modeling by Multi-Task Synergy

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-06-26T05:49:02.609960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-26T05:48:34.492354Z digest=sha256:5fad6075dd823955bb9b9eac164cd62b729954af27db6814e710e614f1146d12

Observation d2c65a08-fd6b-4f1f-bbf4-a6baa1149595 · inbound

MoRoute: Dynamic Routing for In-Context Multimodal Video Generation cites this paper.

MoRoute: Dynamic Routing for In-Context Multimodal Video Generation Lance: Unified Multimodal Modeling by Multi-Task Synergy

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T04:58:50.332245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:58:50.332245Z digest=sha256:d9e879efbe5df9147bfbf0783fd5998b805a0b2de53fa8413db93e5570a922b1

Observation 87316730-cb9d-464f-b5ad-91b868596498 · inbound

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes cites this paper.

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes Lance: Unified Multimodal Modeling by Multi-Task Synergy

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T11:55:27.862037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:55:27.862037Z digest=sha256:fe4a7290c888835913db104b7d7a34973cc027594e657136ef2e1585d7164318

Observation 8ed3445c-2933-41fe-91ec-e14877428f08 · inbound

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes cites this paper.

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes Lance: Unified Multimodal Modeling by Multi-Task Synergy

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T17:08:54.580802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:08:54.580802Z digest=sha256:cb50128a96a7a022248854c293809fb96504c1a6d2c3d6f3c4d23b99cb19ba3f

Observation 569f7263-26fc-45b3-8025-46d8129b94ac · inbound

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction cites this paper.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Lance: Unified Multimodal Modeling by Multi-Task Synergy

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.562911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.562911Z digest=sha256:4b1b3abebd465737fa36bc0934a3a67e39ac6b9dc930545a6724b7b0a55b6f88

Observation 6589330e-2668-4861-94ce-e2b11e0bead1 · inbound

HounsWorld: A Multimodal World Model for Hidden Patient-State Readout, Reconstruction, and Simulation cites this paper.

HounsWorld: A Multimodal World Model for Hidden Patient-State Readout, Reconstruction, and Simulation Lance: Unified Multimodal Modeling by Multi-Task Synergy

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-15T21:00:35.762221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:00:35.762221Z digest=sha256:00f55e08cbedd04db6980e4d07a80999799ba1e9068ad86ee362f654bd96f808