Pith. sign in

Paper Citation Record · LEDGER

LMFusion: Adapting Pretrained Language Models for Multimodal Generation

As of 17 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 53 inbound Pith citation observations for arXiv:2412.15188.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.15188 v4

Coverage vector

measured 24 of 24 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T11:37:11.041300Z

measured 77 of 77 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 53 of 53 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T05:18:15.233236Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

24 of 24 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved21
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 2357e8fc-cbdd-4e4d-9800-22e685329399 · outbound

This paper cites Jointly Training Large Autoregressive Multimodal Models.

LMFusion: Adapting Pretrained Language Models for Multimodal Generation Jointly Training Large Autoregressive Multimodal Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:10.936279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:10.936279Z digest=sha256:59fb82a0e0f3b6667db9a2fdd76a6b0afc50c3a7384f3c91863cfbc45b88534a

Observation 28098a00-1df7-4891-8855-94f5ea28dac9 · outbound

This paper cites The Llama 3 Herd of Models.

LMFusion: Adapting Pretrained Language Models for Multimodal Generation The Llama 3 Herd of Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:10.951261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:10.951261Z digest=sha256:e0c5a5e48f049e28dfd1e827f145d5c2e6d2a61bd2e34323cb04dad111ac5773

Observation e3684a2f-3ec4-403a-95b6-7fad20f51d8b · outbound

This paper cites SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation.

LMFusion: Adapting Pretrained Language Models for Multimodal Generation SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:10.960187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:10.960187Z digest=sha256:b58128c6dac41752c89ffc2e25cd72158e02eee78499ad21a9e3ceb8a082d770

Observation 942c7eae-ddf2-43b5-a237-6a86330fb483 · outbound

This paper cites Auto-Encoding Variational Bayes.

LMFusion: Adapting Pretrained Language Models for Multimodal Generation Auto-Encoding Variational Bayes

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:10.969318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:10.969318Z digest=sha256:a9f0f0f3eeba1b3a1a8786a4f5273ccc80dfae7d9b1a375fbc8b98e64202cb96

Observation b3e75af6-e46b-4bd4-965c-d3c1240024cd · outbound

This paper cites GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding.

LMFusion: Adapting Pretrained Language Models for Multimodal Generation GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:10.973378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:10.973378Z digest=sha256:c5975d341ea9baa50ea1b3843884ce79b5493b71a9751cf33f29255434deb0df

Observation e0417250-e716-417a-8e28-60b4761ef14a · outbound

This paper cites MoE-LLaVA: Mixture of Experts for Large Vision-Language Models.

LMFusion: Adapting Pretrained Language Models for Multimodal Generation MoE-LLaVA: Mixture of Experts for Large Vision-Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:10.977006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:10.977006Z digest=sha256:c777bedbb8bc0caf2451e8d9e44a06be4b6354e54fa14e182e70a92ba76d69f8

Observation 33177bf8-3934-4482-ac5f-5bcc1fedcd81 · outbound

This paper cites Improved baselines with visual instruction tuning.

LMFusion: Adapting Pretrained Language Models for Multimodal Generation Improved baselines with visual instruction tuning

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:37:11.345509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T11:37:10.986658Z digest=sha256:e89ddf0fa37c6b9f81de2b1212499f388c3eec16ec70ecfcaba503f30425c240

Observation fed01ad0-89a5-4f7b-a333-0bf396f2cffb · outbound

This paper cites ChartQA: A benchmark for question answering about charts with visual and logical reasoning.

LMFusion: Adapting Pretrained Language Models for Multimodal Generation ChartQA: A benchmark for question answering about charts with visual and logical reasoning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:10.991130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:10.991130Z digest=sha256:a1a0f09b963fb2aa6c596c5f6450e2ff9f66512d5f30679b9b44d6cecfb1db21

Observation 23855cfc-c50c-455d-818f-6ea21c9317fe · outbound

This paper cites OLMoE: Open Mixture-of-Experts Language Models.

LMFusion: Adapting Pretrained Language Models for Multimodal Generation OLMoE: Open Mixture-of-Experts Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:10.995277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:10.995277Z digest=sha256:20d441c56d117a4d94835688c0569f52a68b3841525af8f18e19ed7760e1b336

Observation f60c53a6-26f1-49d9-8de5-35c162883f37 · outbound

This paper cites Social iqa: Commonsense reasoning about social interactions.

LMFusion: Adapting Pretrained Language Models for Multimodal Generation Social iqa: Commonsense reasoning about social interactions

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:37:11.311650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T11:37:11.003822Z digest=sha256:ceabb79f73e82e4f4125d4d2ee8da1621356fef039212d30ee2a1f60712ed772

Observation 44a38423-3e55-448a-99fa-045b3e9612ad · outbound

This paper cites Emu: Generative Pretraining in Multimodality.

LMFusion: Adapting Pretrained Language Models for Multimodal Generation Emu: Generative Pretraining in Multimodality

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:11.018372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:11.018372Z digest=sha256:0379972885c3f6cd1b316ecc374f1a2ed0159417beb758e80d7f3a29a0c36a14

Observation 138570ee-9ce3-4384-8038-6d3b2b51818d · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

LMFusion: Adapting Pretrained Language Models for Multimodal Generation Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:11.023935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:11.023935Z digest=sha256:9a13ec9297f64e1a2d84ed6d84b3ce64e46ac49fe52ef6da3b219f63f85fe536

Observation 58c6be49-250f-4c1c-bb91-a452eda8afb6 · outbound

This paper cites MetaMorph: Multimodal Understanding and Generation via Instruction Tuning.

LMFusion: Adapting Pretrained Language Models for Multimodal Generation MetaMorph: Multimodal Understanding and Generation via Instruction Tuning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:11.028592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:11.028592Z digest=sha256:227d196ab3b7158e52b6444ac9f4f88c228176b84ec86f8c1af1240348f1a81c

Observation 4236d180-888b-408d-bf85-976a0421a142 · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

LMFusion: Adapting Pretrained Language Models for Multimodal Generation Emu3: Next-Token Prediction is All You Need

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:11.037017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:11.037017Z digest=sha256:396cc74c26eb0b0ee562357abbf0bd79c3affc7b7803277e33b0b14b12edf2c4

Observation 292f1f6e-ef4a-4b3a-a88b-4bf399200f4a · outbound

This paper cites Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model.

LMFusion: Adapting Pretrained Language Models for Multimodal Generation Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:11.041300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:11.041300Z digest=sha256:1cf6717954638b86bd366a8a8b2d9146288608200a2c6062cc86ede94dfafdb5

Observation 65164a7a-6899-474a-967b-4f3cf16dca0f · outbound

This paper cites MoMa: Efficient Early-Fusion Pre-training with Mixture of Modality-Aware Experts.

LMFusion: Adapting Pretrained Language Models for Multimodal Generation MoMa: Efficient Early-Fusion Pre-training with Mixture of Modality-Aware Experts

Reference 2014

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:10.981853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:10.981853Z digest=sha256:f3441f663aac5e4cbf90002c81d842fe0b3e6c5ea21f0c3b09cb8e8301b2489e

Observation cb3bc181-11ff-40b6-b434-7e5282954bce · outbound

This paper cites VLMo: Unified Vision-Language Pre-Training with Mixture-of-Modality-Experts.

LMFusion: Adapting Pretrained Language Models for Multimodal Generation VLMo: Unified Vision-Language Pre-Training with Mixture-of-Modality-Experts

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:11.032673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:11.032673Z digest=sha256:de007669bea368320a333cc005037936e1e57efc063192d7f956647a6eed2fb5

Observation f2447408-82ca-41d8-a704-11e2131eea0b · outbound

This paper cites Scaling Vision-Language Models with Sparse Mixture of Experts.

LMFusion: Adapting Pretrained Language Models for Multimodal Generation Scaling Vision-Language Models with Sparse Mixture of Experts

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:11.014189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:11.014189Z digest=sha256:71ddc47532a33ad843d2d62760356174d85ef6ea2d7624efd1c92adc76d772f1

Observation 676349cd-9188-40ac-badf-ff5c019aea5f · outbound

This paper cites Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer.

LMFusion: Adapting Pretrained Language Models for Multimodal Generation Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:11.008771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:11.008771Z digest=sha256:f53eb3644279c372ad10e1a9dedd48729bd152d833eb624aa26e83e3adf4bc78

Observation c6deebb6-705d-4903-bc10-e97a54c52f78 · outbound

This paper cites EVE: Efficient Vision-Language Pre-training with Masked Prediction and Modality-Aware MoE.

LMFusion: Adapting Pretrained Language Models for Multimodal Generation EVE: Efficient Vision-Language Pre-training with Masked Prediction and Modality-Aware MoE

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:10.941520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:10.941520Z digest=sha256:675f474903a23249f8d96ac3168dc02851b48ba6862e07327a374a961951b845

Observation e992d828-5c0b-4ca8-b2de-11554d5e041c · outbound

This paper cites U-net: Convolutional networks for biomedical image segmentation.

LMFusion: Adapting Pretrained Language Models for Multimodal Generation U-net: Convolutional networks for biomedical image segmentation

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:37:11.324797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T11:37:10.999642Z digest=sha256:e2a253d5eeedb62b4eee36744472c2fe40baefe732b22b7b6afbb541d226cc28

Observation eab11fe3-250f-456b-b8b5-3c52585ccb0f · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

LMFusion: Adapting Pretrained Language Models for Multimodal Generation MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:10.955844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:10.955844Z digest=sha256:ffe1c5c57313cb7fcd4658779638748cb02b9dffa7020ba7e29ffa1cddea0096

Observation f151846d-3be3-4942-b2b2-6a9a7eaed538 · outbound

This paper cites DreamLLM: Synergistic Multimodal Comprehension and Creation.

LMFusion: Adapting Pretrained Language Models for Multimodal Generation DreamLLM: Synergistic Multimodal Comprehension and Creation

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:10.946403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:10.946403Z digest=sha256:3ab261cd21d90ea83f1bcc9a6961dd3ad4ea07b3756b10b36e8f7fa2cc653666

Observation 80703a0e-fd71-4e46-9967-6a9dceae9847 · outbound

This paper cites MARS: Mixture of Auto-Regressive Models for Fine-grained Text-to-image Synthesis.

LMFusion: Adapting Pretrained Language Models for Multimodal Generation MARS: Mixture of Auto-Regressive Models for Fine-grained Text-to-image Synthesis

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:10.964681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:10.964681Z digest=sha256:a6fb367ab826861cfe9690264465a22aa1ebdbd7758459254cff14728079a055

Pith citing papers

Observation 4281b462-436e-40f7-899c-6eae8f1ce502 · inbound

UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths cites this paper.

UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T15:24:46.407707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:24:46.407707Z digest=sha256:35427cd08c532cda5510f71dee1a08ff163f563af24879b91e3417aac7b50b14

Observation ee30dcd7-c7ae-46b0-88c6-2244f47e3098 · inbound

Diffusion Instruction Tuning cites this paper.

Diffusion Instruction Tuning LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-09T11:20:02.915960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:20:02.915960Z digest=sha256:00503114029a038195e551dff4d1b82d20a21225ef3569b369e0fff6c397607c

Observation 82064ae3-1537-41f5-954b-ed80812f159e · inbound

I Think, Therefore I Diffuse: Enabling Multimodal In-Context Reasoning in Diffusion Models cites this paper.

I Think, Therefore I Diffuse: Enabling Multimodal In-Context Reasoning in Diffusion Models LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-08T10:24:41.573614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:24:41.573614Z digest=sha256:a37fcf3292f54f8943c762c3f7c04d404f4126a6331c5d248b90aeff138ef582

Observation df39cbdd-3381-441e-a123-c2e256285cdc · inbound

WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation cites this paper.

WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T16:24:27.719941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-15T16:24:27.407376Z digest=sha256:5f99bf2c26e79c5923906ffaf9d83567ea67534811dda8ae04eb816d05998ed3

Observation b00e0bf0-cad8-4fb7-9243-39b51f257aa3 · inbound

Transfer between Modalities with MetaQueries cites this paper.

Transfer between Modalities with MetaQueries LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T22:49:23.256298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-14T22:49:23.074271Z digest=sha256:c4ac75bf5034824db61ddaa13c33d37684dd383ae1eefb501694909bd705eaf0

Observation 3a95b799-ed57-4421-bde8-6ab16e94e4f6 · inbound

X-Fusion: Introducing New Modality to Frozen Large Language Models cites this paper.

X-Fusion: Introducing New Modality to Frozen Large Language Models LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-16T05:18:15.233236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:18:15.233236Z digest=sha256:3475add60707c7be1fa9f074bcd4a2e38005c28141d8f1f9942b86ca3b12c0e4

Observation 7ee21f9f-92f9-4b36-930e-fb64704c0f09 · inbound

Nexus-Gen: Unified Image Understanding, Generation, and Editing via Prefilled Autoregression in Shared Embedding Space cites this paper.

Nexus-Gen: Unified Image Understanding, Generation, and Editing via Prefilled Autoregression in Shared Embedding Space LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T05:10:53.459495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:10:53.459495Z digest=sha256:f148a7942bd02f5f3041a1bed397c1b043f7d3ade77e64fac1cb61f42f81cbbb

Observation fe53db0e-c508-4791-9d1c-6ea7850c202a · inbound

TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation cites this paper.

TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-15T23:09:10.901259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:09:10.901259Z digest=sha256:0fd5be9c6f43838ff345172055aadcc528fe5936f8384427f3a8afda680357dc

Observation f1969a78-60b6-4416-b4b6-16c8b833dcbf · inbound

BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset cites this paper.

BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T23:34:27.059549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-11T23:34:26.878354Z digest=sha256:9b82ba7e8165b13801fcd232ae5547bc6484c1bc0d764fed9c07834c8e3f8236

Observation 04671a1f-5f16-4c92-8f25-e16a2ac660be · inbound

Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis cites this paper.

Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T21:21:22.481319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:21:22.481319Z digest=sha256:28d58f9e619b83dcdd7d769072323b7f0cdab877930b5314406c63f1988a18c5

Observation 51b2df35-9b9c-4a5f-92ae-eaf04c263e21 · inbound

Emerging Properties in Unified Multimodal Pretraining cites this paper.

Emerging Properties in Unified Multimodal Pretraining LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 66

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T16:23:41.942313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T16:23:41.854132Z digest=sha256:5e6800127be182b8e1d0a4333efdbea83ed16c011b5635ccaf68e3d8211ab13b

Observation 49da6413-21c2-4efa-906c-0995830cdede · inbound

Embedding-to-Prefix: Parameter-Efficient Personalization for Pre-Trained Large Language Models cites this paper.

Embedding-to-Prefix: Parameter-Efficient Personalization for Pre-Trained Large Language Models LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T21:02:03.202470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:02:03.202470Z digest=sha256:b54590067c3f328cc76d19647ae96973f1eadc6a83851b94ce70bc961c1a630a

Observation 05dac12b-2ec1-4114-b49e-baa9bdda25cf · inbound

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning cites this paper.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:25.463377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:25.463377Z digest=sha256:deb2747fec4ac560a7787ef650c0947f5f700f42f9c04f5911b0e2977bc082d0

Observation 4e7650c6-7a75-4f80-a7ee-22d476024b72 · inbound

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation cites this paper.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:17.119373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:17.119373Z digest=sha256:e32ed3c67d41d20962f9a4e20a9b55d655da36f63aa0caab31a4c2173500cbc3

Observation 31f5f6e2-7a61-446a-af72-fef9c92bf846 · inbound

Knowledge Insulating Vision-Language-Action Models: Train Fast, Run Fast, Generalize Better cites this paper.

Knowledge Insulating Vision-Language-Action Models: Train Fast, Run Fast, Generalize Better LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:23.846395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:23.846395Z digest=sha256:329759ade64a16473ac90aacbae27b6ee254345eb815fc11a5f77a0d08ff5757

Observation a98f8182-ac5b-424c-b2ac-401da3970f5a · inbound

LaTtE-Flow: Layerwise Timestep-Expert Flow-based Transformer cites this paper.

LaTtE-Flow: Layerwise Timestep-Expert Flow-based Transformer LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T05:51:42.915375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:51:42.915375Z digest=sha256:406c3cc2bb397df0cb4988467d01c7f846fc39231f39233403c56c8db3acc2be

Observation 58e5e446-5684-49ed-a9fb-a57566674917 · inbound

MADFormer: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation cites this paper.

MADFormer: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:02.245217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:26:02.245217Z digest=sha256:116d2ed4646adc012427a8be47a6cf5695332f3ae0bf5207a5a36a4ef63c30ba

Observation 4a9b2bc1-764f-4d42-96fb-7ae08c69dbe3 · inbound

Dreamland: Controllable World Creation with Simulator and Generative Models cites this paper.

Dreamland: Controllable World Creation with Simulator and Generative Models LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:16.360980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:26:16.360980Z digest=sha256:013ad7744a9d9b71fbb0b2ed92f0fffd387c73471cb984dd503740e6686f96ec

Observation 73a3cffc-4bfb-4dea-88c1-15afa9fa889f · inbound

Show-o2: Improved Native Unified Multimodal Models cites this paper.

Show-o2: Improved Native Unified Multimodal Models LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 95

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T18:51:15.590013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-12T18:51:15.428692Z digest=sha256:748ad3ebe0be81ec15a9cdf849bb411e81d9009dc6cfd5b995bb128ddd60b765

Observation 301bff5f-135a-4629-bd06-220f1bcd1fea · inbound

OmniGen2: Towards Instruction-Aligned Multimodal Generation cites this paper.

OmniGen2: Towards Instruction-Aligned Multimodal Generation LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 68

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T07:52:10.858881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:c3af7ea8d08bef8a1a101a81d8677387047b4f825f8e789dd907605d0c1f240c

Observation cd09d3a2-d46b-49ef-be2a-5baa8f857fa3 · inbound

WordCon: Word-level Typography Control in Scene Text Rendering cites this paper.

WordCon: Word-level Typography Control in Scene Text Rendering LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T22:34:16.184346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:34:16.184346Z digest=sha256:d8f481039098eaa032001ad8b23cbc005f9120e5aa0dd426b880d8718e2babe3

Observation 72d32ee3-09c9-486d-a573-65be95ab85cf · inbound

Rethinking Discrete Tokens: Treating Them as Conditions for Continuous Autoregressive Image Synthesis cites this paper.

Rethinking Discrete Tokens: Treating Them as Conditions for Continuous Autoregressive Image Synthesis LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T20:49:42.028479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:49:42.028479Z digest=sha256:18bdd06e55e041ed6d34c7e4456752229f271cd0541d24ec853f9efa3ee7599b

Observation d2c473d0-10b9-4e03-938d-35c50e31c5ef · inbound

FreeLoRA: Enabling Training-Free LoRA Fusion for Autoregressive Multi-Subject Personalization cites this paper.

FreeLoRA: Enabling Training-Free LoRA Fusion for Autoregressive Multi-Subject Personalization LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T20:46:25.122139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:46:25.122139Z digest=sha256:6c2a355cc1161da9ff4c5e5ff124ad69ebd267a8d02c552dbc798a96a73a8973

Observation 594a0dbc-5e29-48b0-89c5-62155ea44e51 · inbound

X-Omni: Reinforcement Learning Makes Discrete Autoregressive Image Generative Models Great Again cites this paper.

X-Omni: Reinforcement Learning Makes Discrete Autoregressive Image Generative Models Great Again LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T12:10:08.289255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:10:08.289255Z digest=sha256:630e587100e150af376ae78ece5e439620489615e817888757481dc0ffc676a8

Observation c0b5789b-2dbd-4c63-8b2f-c2f0f5f01f4b · inbound

OmniHuman-1.5: Instilling an Active Mind in Avatars via Cognitive Simulation cites this paper.

OmniHuman-1.5: Instilling an Active Mind in Avatars via Cognitive Simulation LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-15T16:57:29.686053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:57:29.686053Z digest=sha256:f9895fe5952d7fbe65d595c68b4d25415f6204b3cdc75581492226fcf18ae1db

Observation b17a1063-74ee-4200-8211-e01266c29ff6 · inbound

Galaxea Open-World Dataset and G0 Dual-System VLA Model cites this paper.

Galaxea Open-World Dataset and G0 Dual-System VLA Model LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T13:31:09.919265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:31:09.919265Z digest=sha256:6f849a87b442d296f87ec6368d9eeb9740dcb41cf605f535a955af809c2f19a7

Observation fc05f3a1-bb1c-445f-8eda-3a47119e9884 · inbound

OccVLA: Vision-Language-Action Model with Implicit 3D Occupancy Supervision cites this paper.

OccVLA: Vision-Language-Action Model with Implicit 3D Occupancy Supervision LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-15T16:26:15.644541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:26:15.644541Z digest=sha256:4e3bad4ee4b46a4a94a983531c855117bea37bd5140bd5eae66421c9ca8fdada

Observation 93875204-2588-4310-b563-7bab02756f4c · inbound

Interleaving Reasoning for Better Text-to-Image Generation cites this paper.

Interleaving Reasoning for Better Text-to-Image Generation LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:44.950531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:44.950531Z digest=sha256:9886ca0b01388db84f6f089458e2a822c218a7e5d4170d688c36d34113814679

Observation b562f445-e176-474e-9c2d-9b13fbd5371d · inbound

Reconstruction Alignment Improves Unified Multimodal Models cites this paper.

Reconstruction Alignment Improves Unified Multimodal Models LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-04T22:36:08.140951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T22:36:08.140951Z digest=sha256:3e8a4d0312a0434198c5bc83f017bfb0c48e7a74c225e701458dce888eeb3f9b

Observation ef364fbd-f8fd-4b0c-bf37-103f5d74a52b · inbound

UniVideo: Unified Understanding, Generation, and Editing for Videos cites this paper.

UniVideo: Unified Understanding, Generation, and Editing for Videos LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T10:49:49.019944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:49:49.019944Z digest=sha256:82a57a82503cd6ef6fc1902799b30236eac04b9fe36c56a0045767c9a67c7302

Observation 527a0908-9afd-40e7-91e4-9cdb4c2ca112 · inbound

SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models cites this paper.

SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T09:54:26.449060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:54:26.449060Z digest=sha256:925af82f9b8104ab80049fc70faaa30f146e66ab316002b069ca054245566207

Observation 70e91f46-7163-4878-a42c-6e83fdf87912 · inbound

FasterVAR: Plug-and-Play Acceleration for Visual Autoregressive Models cites this paper.

FasterVAR: Plug-and-Play Acceleration for Visual Autoregressive Models LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T15:36:51.177913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T15:36:51.177913Z digest=sha256:e7f82f761bdab5bef3d8c797943a927c1f95ddf49435d968f44b089e0e705e3f

Observation 71837f0d-f365-4e9a-814e-146d94d08b85 · inbound

CG-MLLM: Captioning and Generating 3D content via Multi-modal Large Language Models cites this paper.

CG-MLLM: Captioning and Generating 3D content via Multi-modal Large Language Models LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-21T14:50:14.718128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-21T14:48:21.787919Z digest=sha256:b6012fc6f410dceb7aac551d4fb4aa47423748d381a496984418e1a0302e7c72

Observation 8854c87b-93c1-45fc-926c-f848d8caded8 · inbound

ChatUMM: Robust Context Tracking for Conversational Interleaved Generation cites this paper.

ChatUMM: Robust Context Tracking for Conversational Interleaved Generation LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T03:57:32.062558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:57:32.062558Z digest=sha256:b3f0396bd6d497b75a08819edf266fd8cdde61f4d8987570da613b31bea9348e

Observation c89b769c-3ee7-42e6-8247-0d57e1dcebd4 · inbound

LLaMo: Scaling Pretrained Language Models for Unified Motion Understanding and Generation with Continuous Autoregressive Tokens cites this paper.

LLaMo: Scaling Pretrained Language Models for Unified Motion Understanding and Generation with Continuous Autoregressive Tokens LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 58

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T05:02:19.618705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-16T05:01:11.880003Z digest=sha256:edcba5010a1d3666ac4850a78f82b28cb974f2a7aef367f4a0ddad986b206dfe

Observation aa3453fd-dc74-4795-95b3-f947d0d1f384 · inbound

Demystifying Video Reasoning cites this paper.

Demystifying Video Reasoning LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 50

Resolution
unresolved
no resolver link, observed 2026-07-13T23:27:11.006580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:27:11.006580Z digest=sha256:62d90c802da7137638950b2f6a79b0feeb45583483d8c776391050dccad5d39b

Observation 8f54f030-7751-43f7-9d1a-645e1db41827 · inbound

Demystifying Video Reasoning cites this paper.

Demystifying Video Reasoning LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-03T02:33:57.876583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:33:57.876583Z digest=sha256:28390687d59bd160ffbb51ebe454e6f739677c7e8ba5368fe3bf1df869534899

Observation 10617339-23d3-4317-89c2-12d670fc9617 · inbound

LMGenDrive: Bridging Multimodal Understanding and Generative World Modeling for End-to-End Driving cites this paper.

LMGenDrive: Bridging Multimodal Understanding and Generative World Modeling for End-to-End Driving LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:31:01.398245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T17:36:38.627415Z digest=sha256:9c152e8c72fdc845b517f833cc74acb7de32bf98d5439b95fb8cad15f9ce0f35

Observation b0800302-88df-4745-bf7b-dd720fc3f116 · inbound

Free Lunch for Unified Multimodal Models: Enhancing Generation via Reflective Rectification with Inherent Understanding cites this paper.

Free Lunch for Unified Multimodal Models: Enhancing Generation via Reflective Rectification with Inherent Understanding LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T14:15:28.953805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T14:15:02.723774Z digest=sha256:2cc1efec6c4e7f0868bf660e824cd987f048c9d089ec972429a767e70b975105

Observation 06ef9691-e5ef-45bd-8e01-14279bf16a8f · inbound

Co-generation of Layout and Shape from Text via Autoregressive 3D Diffusion cites this paper.

Co-generation of Layout and Shape from Text via Autoregressive 3D Diffusion LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T09:23:37.044329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T09:22:10.651922Z digest=sha256:a730af3835c8b5b60290ffae0b2c3b974e3b3218c9f14129c8cda1fddc954e93

Observation c235ae80-05e8-4e10-8822-a65b8dec7d74 · inbound

MMCORE: MultiModal COnnection with Representation Aligned Latent Embeddings cites this paper.

MMCORE: MultiModal COnnection with Representation Aligned Latent Embeddings LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T03:29:21.669427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T03:27:50.144706Z digest=sha256:d9337da0817ff5ce596db629b7d1a50c7241b69a36dadb8b82759398945b5fef

Observation 66ee8cee-fdc4-41e5-ab3a-66281e56cb1d · inbound

Meta-CoT: Enhancing Granularity and Generalization in Image Editing cites this paper.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 52

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T21:46:14.691378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:aa57e8b23b38c82a8583c656e5db51b00209ae55cd8c165ee78ab45ac65d38d9

Observation 51bd0ddd-055f-465b-a524-64859693240b · inbound

SpatialFusion: Endowing Unified Image Generation with Intrinsic 3D Geometric Awareness cites this paper.

SpatialFusion: Endowing Unified Image Generation with Intrinsic 3D Geometric Awareness LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:46:26.816740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-07T13:45:53.346402Z digest=sha256:41cfbb9c87685399922037e0493595ed12af4292f69c0096422c0493f3d8ffee

Observation 73068788-af9b-4a9b-89b8-001858cd0320 · inbound

STARFlow2: Bridging Language Models and Normalizing Flows for Unified Multimodal Generation cites this paper.

STARFlow2: Bridging Language Models and Normalizing Flows for Unified Multimodal Generation LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T02:45:57.156891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-11T02:44:59.644215Z digest=sha256:d09421999375fec3788067ea052798301cfc0a74edc50ca554a7853ce6013d70

Observation c63185a6-3916-499f-8a62-05a744587d3b · inbound

Reversing the Flow: Generation-to-Understanding Synergy in Large Multimodal Models cites this paper.

Reversing the Flow: Generation-to-Understanding Synergy in Large Multimodal Models LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T18:48:53.450985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T18:44:54.835575Z digest=sha256:ab50d16c260bd4223f0d738b1e39baea08850db5a1178098aa5a46d3d516e4f6

Observation 8fe1fa5e-1efa-4d0a-a3b9-8971c93eae2c · inbound

Semantic Generative Tuning for Unified Multimodal Models cites this paper.

Semantic Generative Tuning for Unified Multimodal Models LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:33:14.202728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T11:32:24.007847Z digest=sha256:fef1b93c51997c2f13dd867d90812e408cffd341c9f4c870cd7b66764313fe5b

Observation e554ff05-2944-404f-b193-242c8dc4740e · inbound

Semantic Generative Tuning for Unified Multimodal Models cites this paper.

Semantic Generative Tuning for Unified Multimodal Models LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:35:00.282413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:21d9d6200c6dd019dbf6759d354a509371c251c2cfcd83c12326dfd090c867a1

Observation 86284e46-aff6-4509-b9a9-68049f183b84 · inbound

Where to Refine, When to Stop: Rethinking Redundancy via Latent Discrepancy for Efficient Visual Autoregressive Generation cites this paper.

Where to Refine, When to Stop: Rethinking Redundancy via Latent Discrepancy for Efficient Visual Autoregressive Generation LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T22:42:46.709383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-28T22:38:26.147074Z digest=sha256:183d039d8f9201d3450547b9fd668385a72cecd7be1173903f0a63f6df58151b

Observation a6e04f3e-7839-49c8-b13a-3a891a308702 · inbound

Polaris: Scaling Up Instruction-Guided Image Generation Towards Millions of Personalized Style Needs cites this paper.

Polaris: Scaling Up Instruction-Guided Image Generation Towards Millions of Personalized Style Needs LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 97

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T22:46:18.726597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-06-28T15:07:03.439928Z digest=sha256:1ae0d386821abbb862889948772666035a05d8f8de7fe252c9ba52d5dcf53672

Observation 74f42b3d-7ebf-4053-b30b-be90316c1160 · inbound

UniTac: A Unified Multimodal Model for Cross-Sensor Tactile Understanding and Generation cites this paper.

UniTac: A Unified Multimodal Model for Cross-Sensor Tactile Understanding and Generation LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-01T10:05:40.624385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-07-01T05:56:03.597839Z digest=sha256:fda282a1d54ba10fd648c64c233867b90590edb89d1c0badfb07ed0e88c5018f

Observation 583ff840-3bba-428a-894e-007a3b67114d · inbound

Streaming Multi-Agent Autoregressive Diffusion Model with World State Registers cites this paper.

Streaming Multi-Agent Autoregressive Diffusion Model with World State Registers LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 126

Resolution
unresolved
no resolver link, observed 2026-08-01T07:02:51.493005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:02:51.493005Z digest=sha256:b08c557d5d2cdbe8cea93dc0a8618c1137d509c17569cfe5bb932f6697355b82

Observation 9c8c05c8-dcfb-46f2-961d-2d9147766e61 · inbound

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes cites this paper.

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-06T11:55:28.069606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:55:28.069606Z digest=sha256:459db9198dfbe029b5fdfb22bec21f665e1c94c63b383c73ade64c5d17704df9

Observation 1b47279d-1775-4e41-80f8-d3ed0ac8850b · inbound

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes cites this paper.

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-08T17:08:55.803846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:08:55.803846Z digest=sha256:1007ec4f5cf64e83a90f5c98b0c4e88ef66f6edc086229a63caac5a0e7997e11