Pith. sign in

Paper Citation Record · LEDGER

Liquid: Language Models are Scalable and Unified Multi-modal Generators

As of 6 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 15 inbound Pith citation observations for arXiv:2412.04332.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.04332 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T11:55:28.167608Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T00:04:22.401157Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 9bbaf9aa-4ec2-4012-b95c-27c78871a41e · inbound

WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation cites this paper.

WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation Liquid: Language Models are Scalable and Unified Multi-modal Generators

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-15T16:24:27.491335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T16:24:27.407376Z digest=sha256:f90967f61d8bfdb2f02373fd2184f33ee3c0ba8b119b5d3bab8cf13ce205fd1a

Observation ce410087-09ce-4ded-974b-f0731aa4b796 · inbound

Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation cites this paper.

Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation Liquid: Language Models are Scalable and Unified Multi-modal Generators

Reference 85

Resolution
verified exact
arxiv_id, observed 2026-05-17T07:24:04.698301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T07:24:04.460276Z digest=sha256:e12ce9ce94a987c26e0529ce0dca7c8994847ccfa95534b2d17db0fe263c5da0

Observation 67e2ff16-9551-4c35-a2cc-7fc008b4854b · inbound

Emerging Properties in Unified Multimodal Pretraining cites this paper.

Emerging Properties in Unified Multimodal Pretraining Liquid: Language Models are Scalable and Unified Multi-modal Generators

Reference 85

Resolution
verified exact
arxiv_id, observed 2026-05-10T16:23:42.041939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:23:41.854132Z digest=sha256:f88d4f2916cf75e9b06f10c5dc921f83530d77a76ddd58e6ad6958abf671938c

Observation 09c1505d-c155-403f-a089-bbac8913b2f5 · inbound

Show-o2: Improved Native Unified Multimodal Models cites this paper.

Show-o2: Improved Native Unified Multimodal Models Liquid: Language Models are Scalable and Unified Multi-modal Generators

Reference 118

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:51:15.779964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T18:51:15.428692Z digest=sha256:cb07cf898113f516f07bca399d467f56de74f5b9309ec924a2a474e28e22a3b0

Observation 5afcd07c-6d9f-464d-8890-d84d5122f57b · inbound

Reconstruction Alignment Improves Unified Multimodal Models cites this paper.

Reconstruction Alignment Improves Unified Multimodal Models Liquid: Language Models are Scalable and Unified Multi-modal Generators

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-04T22:36:08.238282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T22:36:08.238282Z digest=sha256:63ed7954f31f30485a07d3be145c40582bb3172977b1ffe07ab88392acd8daad

Observation b540e1e8-a5c9-400e-a37f-a14975165cb7 · inbound

AutoNeural: Co-Designing Vision-Language Models for NPU Inference cites this paper.

AutoNeural: Co-Designing Vision-Language Models for NPU Inference Liquid: Language Models are Scalable and Unified Multi-modal Generators

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-03T18:57:00.890684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:57:00.890684Z digest=sha256:fb8acbef30d31bc7978a685753ed6b0aece491cfc2c91191f6e7275a5a0aff99

Observation ed2cb239-2d97-4609-b8e9-b737b8539c91 · inbound

Pseudo-Unification: Entropy Probing Reveals Divergent Information Patterns in Unified Multimodal Models cites this paper.

Pseudo-Unification: Entropy Probing Reveals Divergent Information Patterns in Unified Multimodal Models Liquid: Language Models are Scalable and Unified Multi-modal Generators

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:50:59.197308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:27:14.491492Z digest=sha256:168401a5e11460ad860ebe84d1dffad1dae7dfecb9c5b892bfbfc158308a9ffe

Observation 45accae2-06e0-437a-a0c2-3dc9140cdc29 · inbound

WinTok: A Win-Win Hybrid Tokenizer via Decomposing Visual Understanding and Generation with Transferable Tokens cites this paper.

WinTok: A Win-Win Hybrid Tokenizer via Decomposing Visual Understanding and Generation with Transferable Tokens Liquid: Language Models are Scalable and Unified Multi-modal Generators

Reference 91

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:08:15.782804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T12:04:19.761430Z digest=sha256:bb7d24a35f1e399a55be46cce86892e75595d9ce484ca4249bb85c5056bb271b

Observation 510d55b9-c96c-44f8-8248-8422ac28ebf1 · inbound

On the Limits of Token Reduction for Efficient Unified Vision Language Training cites this paper.

On the Limits of Token Reduction for Efficient Unified Vision Language Training Liquid: Language Models are Scalable and Unified Multi-modal Generators

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-06-28T17:02:24.295561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:6d76db39410ebc001952a51fa2aa3b53b36adfb3bbbac2207238fbbcb9ba514d

Observation 0c738ec6-7764-4a1d-a5fd-22cfe2530c27 · inbound

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers cites this paper.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers Liquid: Language Models are Scalable and Unified Multi-modal Generators

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:38:28.909508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:cfd7fabb8ad57e512efb0f013145ad0c33dd61a1ec5f02826dbafd9cbd33a254

Observation 8f2c5e0c-af6d-4e79-b120-7b19fba7d402 · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence Liquid: Language Models are Scalable and Unified Multi-modal Generators

Reference 137

Resolution
verified exact
local_arxiv, observed 2026-07-08T00:04:22.402454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-07-07T23:59:38.702609Z digest=sha256:5dccf2bdd5829714eab3a081d44ce62cae0e971153a65609d24c5ae7babe1497

Observation 4b2703cb-d4f9-48e6-a5af-b518d68a27a4 · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence Liquid: Language Models are Scalable and Unified Multi-modal Generators

Reference 137

Resolution
unresolved
no resolver link, observed 2026-07-11T07:46:49.059192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T07:46:49.059192Z digest=sha256:f18e0c8d6f879bb0c3f98d4e41d24f31980ee868da0106f2735a1570f781014c

Observation 385d2136-e0a2-4319-94e1-3412150f6711 · inbound

Twins: Learn to Predict Unified Representations with Focal Loss cites this paper.

Twins: Learn to Predict Unified Representations with Focal Loss Liquid: Language Models are Scalable and Unified Multi-modal Generators

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-01T04:29:51.029704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T04:29:51.029704Z digest=sha256:6ef862937e1c85bea204bb23dcb4cf5194474c39ac480f0913654bb70ed12fca

Observation 2ef15d77-cc04-4d18-80d0-a35ec502ee27 · inbound

Argus-Unified: Towards A Compact and Economical Unified Model for Image Understanding and Generation cites this paper.

Argus-Unified: Towards A Compact and Economical Unified Model for Image Understanding and Generation Liquid: Language Models are Scalable and Unified Multi-modal Generators

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T02:14:09.687868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:14:09.687868Z digest=sha256:cae6a7ed1178b3327daee398d44a9b28c62f49162ef987c8744b9f5ce2337e3f

Observation 06fe37aa-195e-4ebb-95c9-68e02a64175f · inbound

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes cites this paper.

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes Liquid: Language Models are Scalable and Unified Multi-modal Generators

Reference 138

Resolution
unresolved
no resolver link, observed 2026-08-06T11:55:28.167608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:55:28.167608Z digest=sha256:8f5ab4c8968d52de1f78ed112466ef37edfd3d2af9d20daf0c2725d0c6a38956