Pith. sign in

Paper Citation Record · LEDGER

Representation Forcing for Bottleneck-Free Unified Multimodal Models

As of 9 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 3 inbound Pith citation observations for arXiv:2605.31604.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.31604 v4

Coverage vector

measured 65 of 65 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-12T15:31:57.426559Z

measured 68 of 68 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T05:45:51.915796Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-03T23:59:06.123206Z

Reference resolution

65 of 65 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved65
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ee98f9b5-ddb5-4680-8fd8-3fb7b7d0803a · outbound

This paper cites Latent forcing: Reordering the diffusion trajectory for pixel-space image generation.arXiv preprint arXiv:2602.11401, 2026.

Representation Forcing for Bottleneck-Free Unified Multimodal Models Latent forcing: Reordering the diffusion trajectory for pixel-space image generation.arXiv preprint arXiv:2602.11401, 2026

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-12T15:31:57.426559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T15:31:57.426559Z digest=sha256:81af0d661ede331bd0d2196d9236658da4eb2382931edaac643e3acd453ad9df

Observation cf14e559-e5a4-4189-b7be-946ff2e5f16e · outbound

This paper cites Improving image generation with better captions.OpenAI Technical Report,https: // cdn.

Representation Forcing for Bottleneck-Free Unified Multimodal Models Improving image generation with better captions.OpenAI Technical Report,https: // cdn

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-12T15:31:57.426559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T15:31:57.426559Z digest=sha256:255c20e512f1d4b982b8355ea01a20cd91189d757d1b9a0773a351ac33b08f86

Observation cf610772-11ae-44d0-bbc5-116cb59b4169 · outbound

This paper cites FLUX.https://github.com/black-forest-labs/flux, 2024.

Representation Forcing for Bottleneck-Free Unified Multimodal Models FLUX.https://github.com/black-forest-labs/flux, 2024

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-12T15:31:57.426559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T15:31:57.426559Z digest=sha256:058bf8a2a9a20233b7ed4af85dd6046f68a7202b96266db21853c225bee3b9db

Observation 79f71b40-2ea1-4a39-8f84-a480ba44c8ef · outbound

This paper cites Unsupervised learning of visual features by contrasting cluster assignments.

Representation Forcing for Bottleneck-Free Unified Multimodal Models Unsupervised learning of visual features by contrasting cluster assignments

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-12T15:31:57.426559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T15:31:57.426559Z digest=sha256:00c883ee87c90aa58581d410b4be2589343f642d762c554a1e0cb7fda7c5e110

Observation 56d802b2-668a-405c-ad12-93d97afc94b1 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

Representation Forcing for Bottleneck-Free Unified Multimodal Models Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-12T15:31:57.426559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T15:31:57.426559Z digest=sha256:a8670b3d26e519743adb02165691c885d19135347b4eb7f039ff16a2535618ce

Observation 25fa5aff-909c-428d-b74a-cb1e89301b66 · outbound

This paper cites BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset.

Representation Forcing for Bottleneck-Free Unified Multimodal Models BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-12T15:31:57.426559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T15:31:57.426559Z digest=sha256:2aee6f48c234d8925b284102b8181c4686964b68bd20c4e5d184796eaf2229ad

Observation 447b46f2-0f61-4bf4-98bd-b11066588b35 · outbound

This paper cites PixArt-α: Fast training of diffusion transformer for photorealistic text-to-image synthesis.

Representation Forcing for Bottleneck-Free Unified Multimodal Models PixArt-α: Fast training of diffusion transformer for photorealistic text-to-image synthesis

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-12T15:31:57.426559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T15:31:57.426559Z digest=sha256:a2f517ae50b0ec39ba202a2a5ac8ad18d113b33770f902f1ca49c27a9028de78

Observation 5d53b2b8-b652-48ba-9a3d-45cb3819a60d · outbound

This paper cites PixelFlow: Pixel-Space Generative Models with Flow.

Representation Forcing for Bottleneck-Free Unified Multimodal Models PixelFlow: Pixel-Space Generative Models with Flow

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-12T15:31:57.426559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T15:31:57.426559Z digest=sha256:1139048cb3520c9c4d6b56daa1228de640ebbe0a31178446152e91f2e916afc4

Observation 0db96706-1a79-4764-93cd-539fa656f5d0 · outbound

This paper cites Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling.

Representation Forcing for Bottleneck-Free Unified Multimodal Models Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-12T15:31:57.426559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T15:31:57.426559Z digest=sha256:34728a7fd333b31cfe4fec9657dddf482c47beb3d3822e6fe6b1c0a646ab0add

Observation d335398d-8371-4596-bf76-40b3d80154dd · outbound

This paper cites Patch n’ Pack: NaViT, a vision transformer for any aspect ratio and resolution.

Representation Forcing for Bottleneck-Free Unified Multimodal Models Patch n’ Pack: NaViT, a vision transformer for any aspect ratio and resolution

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-12T15:31:57.426559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T15:31:57.426559Z digest=sha256:3a595c2bfe3496e7b387155ade5d975465a46d72b11ecfb20a2a16e5401aa3b1

Observation 30a5e4ab-0b1a-46a2-a73a-bc2569e16f6c · outbound

This paper cites Emerging Properties in Unified Multimodal Pretraining.

Representation Forcing for Bottleneck-Free Unified Multimodal Models Emerging Properties in Unified Multimodal Pretraining

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-12T15:31:57.426559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T15:31:57.426559Z digest=sha256:6348959cdc2623806bcc048159f6e8db6a69f13b8ee0e1c69fe4b1ca8a25a76a

Observation b4f94780-4672-4075-936e-3be7658574c4 · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

Representation Forcing for Bottleneck-Free Unified Multimodal Models Imagenet: A large-scale hierarchical image database

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-12T15:31:57.426559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T15:31:57.426559Z digest=sha256:9b58ee133641dbc50ed2486b91550615419aadd432dd8b703837e1088a4bb195

Observation d8b9f36a-4054-4cf4-b61f-51706c7d3227 · outbound

This paper cites Diffusion models beat gans on image synthesis.

Representation Forcing for Bottleneck-Free Unified Multimodal Models Diffusion models beat gans on image synthesis

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-12T15:31:57.426559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T15:31:57.426559Z digest=sha256:57b26265b020337beb9a840c0f3ff715df8a2928736301ae9154753c37bca735

Observation eec1c095-5b55-49ab-8e4e-e111cafe4944 · outbound

This paper cites SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture.

Representation Forcing for Bottleneck-Free Unified Multimodal Models SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-12T15:31:57.426559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T15:31:57.426559Z digest=sha256:c00bd90f314eed6d2b9cce0edf33193392b962717193ebdb17663b70d9330676

Observation a21bc18c-463d-4e49-bcd4-b73badce8fc5 · outbound

This paper cites Taming transformers for high-resolution image synthesis.

Representation Forcing for Bottleneck-Free Unified Multimodal Models Taming transformers for high-resolution image synthesis

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-12T15:31:57.426559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T15:31:57.426559Z digest=sha256:5d1601fb492656391677a268d37959f76cfe4ba918898c31163bb5611aa70580

Observation de88a2e8-26ff-40b1-9ccf-c70a500b27e8 · outbound

This paper cites Scaling rectified flow transformers for high-resolution image synthesis.

Representation Forcing for Bottleneck-Free Unified Multimodal Models Scaling rectified flow transformers for high-resolution image synthesis

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-12T15:31:57.426559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T15:31:57.426559Z digest=sha256:a6cb76eb9d1a1d7d729499986e2f91bbf928980cc652d9ace507762e7a4f77d4

Observation e14f7bf3-8f66-4a5b-85b2-2ce0bb697474 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

Representation Forcing for Bottleneck-Free Unified Multimodal Models MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-12T15:31:57.426559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T15:31:57.426559Z digest=sha256:36b73742081e6d8d3e946a0fee071b0475c1c17a08aad1a60c9e65d595d8e3b4

Observation eb484af2-cfc6-42d6-9414-5fafe928d9c4 · outbound

This paper cites Smith, Wei-Chiu Ma, and Ranjay Krishna.

Representation Forcing for Bottleneck-Free Unified Multimodal Models Smith, Wei-Chiu Ma, and Ranjay Krishna

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-12T15:31:57.426559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T15:31:57.426559Z digest=sha256:6dc4480fa7bc343787c465ceab0220725426121df7bf44cd8975075b0b52173a

Observation 718affd4-fede-46e9-bdcc-0e395725a9a3 · outbound

This paper cites Seedream 3.0 Technical Report.

Representation Forcing for Bottleneck-Free Unified Multimodal Models Seedream 3.0 Technical Report

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-12T15:31:57.426559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T15:31:57.426559Z digest=sha256:3f030971f09e2a9472e9f06a673a83481d9a551dee6a2be0ad99c3d1f1b356fb

Observation ac8fcb53-ab2b-476c-bab2-ffa7d099b020 · outbound

This paper cites SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation.

Representation Forcing for Bottleneck-Free Unified Multimodal Models SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-12T15:31:57.426559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T15:31:57.426559Z digest=sha256:d093b7d468437dad89181fe0a4a5fec82896f9217576634662e41cafd7a8fd1d

Observation 6a601bf0-b870-4cfc-b0e4-8e260a6ddac7 · outbound

This paper cites GenEval: An object-focused framework for evaluating text-to-image alignment.

Representation Forcing for Bottleneck-Free Unified Multimodal Models GenEval: An object-focused framework for evaluating text-to-image alignment

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-12T15:31:57.426559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T15:31:57.426559Z digest=sha256:6ace14a16ac1dc9cf63c29601995bb7c5d4c9e8175deba57b5d179866211f96c

Observation a814cb94-4f9f-4d2f-b4ba-5aeac8e4843d · outbound

This paper cites HallusionBench: An advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models.

Representation Forcing for Bottleneck-Free Unified Multimodal Models HallusionBench: An advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-12T15:31:57.426559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T15:31:57.426559Z digest=sha256:1287716f236fe4af65c39c4af3951b50c25f6a2a27de865619be141e3a04c87b

Observation 6a7d2f00-d00b-47bb-8e01-b8369cc96e3a · outbound

This paper cites Denoising diffusion probabilistic models.

Representation Forcing for Bottleneck-Free Unified Multimodal Models Denoising diffusion probabilistic models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-12T15:31:57.426559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T15:31:57.426559Z digest=sha256:bca9addb6c2ad64f023c4e0f24c715ea0a6642343f22f66de6e6cf8f1288f72d

Observation a4636529-06f0-4b30-82be-b768c435c7ca · outbound

This paper cites Simpler diffusion (SiD2): 1.5 FID on ImageNet512 with pixel-space diffusion.

Representation Forcing for Bottleneck-Free Unified Multimodal Models Simpler diffusion (SiD2): 1.5 FID on ImageNet512 with pixel-space diffusion

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-12T15:31:57.426559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T15:31:57.426559Z digest=sha256:80b728c79f44d0dc95f6c4c6d745fce8cae837b28b6d2e4e9a6c191348bbdac5

Observation c01d60d0-3499-421c-a9b4-401464a2dd2d · outbound

This paper cites ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment.

Representation Forcing for Bottleneck-Free Unified Multimodal Models ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-12T15:31:57.426559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T15:31:57.426559Z digest=sha256:bc0fff969c2c70d6d0606c8fa9332b3adee492c3526c34e9531cb6a2c2605d8a

Observation 9fb6db5a-f2d2-4a8f-89a8-d34761787717 · outbound

This paper cites A diagram is worth a dozen images.

Representation Forcing for Bottleneck-Free Unified Multimodal Models A diagram is worth a dozen images

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-12T15:31:57.426559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T15:31:57.426559Z digest=sha256:e636105d0a72e1d528d4cf129e4047cd6482970e9ce27174573b1782aef0e5fc

Observation 4811a82a-5758-450f-a856-b6b93d408c76 · outbound

This paper cites Auto-Encoding Variational Bayes.

Representation Forcing for Bottleneck-Free Unified Multimodal Models Auto-Encoding Variational Bayes

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-12T15:31:57.426559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T15:31:57.426559Z digest=sha256:7fd317d6316c50dbe68047427509dde2ca51bef8be87ed194cda17178bb8d56b

Observation 67c35343-616c-4e20-9071-2d5de573cb16 · outbound

This paper cites Back to Basics: Let Denoising Generative Models Denoise.

Representation Forcing for Bottleneck-Free Unified Multimodal Models Back to Basics: Let Denoising Generative Models Denoise

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-12T15:31:57.426559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T15:31:57.426559Z digest=sha256:32c5289ce77469d901ed159df5cef5ac56d1c59c190ed758c242b9d2c98f50c7

Observation 9a6fff01-7a55-423e-ac93-0548897990aa · outbound

This paper cites UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation.

Representation Forcing for Bottleneck-Free Unified Multimodal Models UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-12T15:31:57.426559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T15:31:57.426559Z digest=sha256:e132c448e60bdcda14e2a46ba30f2235140c3fdcad85e74a06138a803cd61986

Observation 9ba1cf97-7cc6-4897-9298-0e9016bd9db9 · outbound

This paper cites World Model on Million-Length Video And Language With Blockwise RingAttention.

Representation Forcing for Bottleneck-Free Unified Multimodal Models World Model on Million-Length Video And Language With Blockwise RingAttention

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-12T15:31:57.426559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T15:31:57.426559Z digest=sha256:606c731c5c756377e27cd90f00926e10b611696be9a4518a0d77765babeaf837

Observation 3fd2ec9a-f921-45db-a224-2f68be5cd6e7 · outbound

This paper cites Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation.

Representation Forcing for Bottleneck-Free Unified Multimodal Models Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-12T15:31:57.426559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T15:31:57.426559Z digest=sha256:c949addd07df26d8f17676c088b3389a38c84a1f034cccff2c122714cf9a1794

Observation 0e76227a-78c1-416b-82bc-04b4366fe459 · outbound

This paper cites Decoupled weight decay regularization.

Representation Forcing for Bottleneck-Free Unified Multimodal Models Decoupled weight decay regularization

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-12T15:31:57.426559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T15:31:57.426559Z digest=sha256:1740cb56fee9522a263ea28d525d8febfff2516b69ca350ca7fb34212064ca87

Observation eeeff4de-3610-4171-b3a1-0c16591adb5a · outbound

This paper cites JanusFlow: Harmonizing autoregression and rectified flow for unified multimodal understanding and generation.

Representation Forcing for Bottleneck-Free Unified Multimodal Models JanusFlow: Harmonizing autoregression and rectified flow for unified multimodal understanding and generation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-12T15:31:57.426559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T15:31:57.426559Z digest=sha256:0a92e38c4074eceee9b5e564dfb6e1bc99d59742700d2a80659682bdfd8b13e5

Observation 91b641b2-8468-4a4b-a317-9af22353db9e · outbound

This paper cites ChartQA: A benchmark for question answering about charts with visual and logical reasoning.

Representation Forcing for Bottleneck-Free Unified Multimodal Models ChartQA: A benchmark for question answering about charts with visual and logical reasoning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-12T15:31:57.426559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T15:31:57.426559Z digest=sha256:c051a19e64fe706a4a26e2ca31b4b124beb53f750832d1a7c6bf3ad56f85da11

Observation 4c9a5c00-a365-4c3e-89fc-85f4d9e504a0 · outbound

This paper cites DocVQA: A Dataset for VQA on Document Images.

Representation Forcing for Bottleneck-Free Unified Multimodal Models DocVQA: A Dataset for VQA on Document Images

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-12T15:31:57.426559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T15:31:57.426559Z digest=sha256:366721961fc1a3573b5bb0b8d08c63ef339824d2943e858d8ca64e2ae204238f

Observation 37c5734d-1b23-4b88-aa57-a2b6e51cc4ac · outbound

This paper cites an unresolved cited work.

Representation Forcing for Bottleneck-Free Unified Multimodal Models Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-12T15:31:57.426559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T15:31:57.426559Z digest=sha256:72ea62be6b694dae981b805b1ccead616dd4a5d8d894ce9bb90fb8624bf672f9

Observation 7ea7e89a-6461-42b2-a2ef-8686c26b6ca3 · outbound

This paper cites Transfer between Modalities with MetaQueries.

Representation Forcing for Bottleneck-Free Unified Multimodal Models Transfer between Modalities with MetaQueries

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-12T15:31:57.426559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T15:31:57.426559Z digest=sha256:8f91bda22094723caff959808cac84f55107822772da6f8ad447624fd2d60f09

Observation 67dcfd7d-4c5d-4302-a93e-d42d9bc1a14d · outbound

This paper cites SDXL: Improving latent diffusion models for high-resolution image synthesis.

Representation Forcing for Bottleneck-Free Unified Multimodal Models SDXL: Improving latent diffusion models for high-resolution image synthesis

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-12T15:31:57.426559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T15:31:57.426559Z digest=sha256:c2808f6dfa2ba1d481017c7dcbaa4cbc318a2d6a3f0556e4046004cbbbadb5dc

Observation b2952c8b-a3dd-46e8-8661-f7d6ecac24fd · outbound

This paper cites Du, Zehuan Yuan, and Xinglong Wu.

Representation Forcing for Bottleneck-Free Unified Multimodal Models Du, Zehuan Yuan, and Xinglong Wu

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-12T15:31:57.426559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T15:31:57.426559Z digest=sha256:aa742c94e5e65a01849ab9bf7a6f95d09d8be5ff1c8bda683767eb35bf88bb14

Observation 7c2ba7e0-ee72-4f80-b79f-6154a7ac1ace · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

Representation Forcing for Bottleneck-Free Unified Multimodal Models Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-12T15:31:57.426559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T15:31:57.426559Z digest=sha256:b6d03528c5890836315ac5932c2bf35471bc650eb626e988481efdf8f9c3c189

Observation 3a187d1e-824f-4c04-b8ac-fee29126b089 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Representation Forcing for Bottleneck-Free Unified Multimodal Models High-resolution image synthesis with latent diffusion models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-12T15:31:57.426559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T15:31:57.426559Z digest=sha256:ea7e5eeedc3a2379c02bc743e4f0862faa996b795e716a137a359824a51de60c

Observation d6f43eeb-99ce-4e5a-8ad5-2266efe89ae6 · outbound

This paper cites Latent diffusion model without variational autoencoder.arxiv: 2510.15301, 2025.

Representation Forcing for Bottleneck-Free Unified Multimodal Models Latent diffusion model without variational autoencoder.arxiv: 2510.15301, 2025

Reference 42

Resolution
unresolved
no resolver link, observed 2026-07-12T15:31:57.426559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T15:31:57.426559Z digest=sha256:21855df2c26d3ed1945d6d9fef6c7c3072d9fb5a4e2ebc13951cb331ccae0b3d

Observation b414f980-526f-4589-83dc-c9fd71210dbb · outbound

This paper cites DINOv3.

Representation Forcing for Bottleneck-Free Unified Multimodal Models DINOv3

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-12T15:31:57.426559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T15:31:57.426559Z digest=sha256:eb307071e66e5401a3928978f0cadb6de2c5d7bb8f18b176f5492e07050ab74f

Observation 0c96917a-63ba-45b4-b2bb-433680bbc4f1 · outbound

This paper cites Generative multimodal models are in-context learners.

Representation Forcing for Bottleneck-Free Unified Multimodal Models Generative multimodal models are in-context learners

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-12T15:31:57.426559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T15:31:57.426559Z digest=sha256:630ee01158d11d2c8a68700d2170a92d608b5d5aa4e4de708e4c30f9137e551c

Observation 5109de35-9c1f-482e-88e6-7ff8e1a1b3da · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

Representation Forcing for Bottleneck-Free Unified Multimodal Models SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 45

Resolution
unresolved
no resolver link, observed 2026-07-12T15:31:57.426559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T15:31:57.426559Z digest=sha256:f2d6acd89accfa9360f4896cc8d248b2665a1da6e07753e0b80bd3df50c757a5

Observation 6bf52a3e-9c07-4463-ad59-885d32286853 · outbound

This paper cites Neural discrete representation learning.

Representation Forcing for Bottleneck-Free Unified Multimodal Models Neural discrete representation learning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-07-12T15:31:57.426559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T15:31:57.426559Z digest=sha256:41c394244c343beff418885575979e91e788bec13f679f1cdfa63b60d1a6e28c

Observation b5cc0082-fe85-46ef-9135-196aa48d17aa · outbound

This paper cites ILLUME: Illuminating your LLMs to see, draw, and self-enhance.

Representation Forcing for Bottleneck-Free Unified Multimodal Models ILLUME: Illuminating your LLMs to see, draw, and self-enhance

Reference 47

Resolution
unresolved
no resolver link, observed 2026-07-12T15:31:57.426559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T15:31:57.426559Z digest=sha256:7aa477bb5e8f7eea76eea60842e41423915bf0bd0d2aa2edf7c65f4462ca572f

Observation 3963e05a-db2e-43bd-9c65-33860daaead3 · outbound

This paper cites PixNerd: Pixel Neural Field Diffusion.

Representation Forcing for Bottleneck-Free Unified Multimodal Models PixNerd: Pixel Neural Field Diffusion

Reference 48

Resolution
unresolved
no resolver link, observed 2026-07-12T15:31:57.426559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T15:31:57.426559Z digest=sha256:34202a29495c9c3952926c494f33a471b4accfa961d6736bcda5f6c2b40c83cd

Observation b1cadfe5-22f6-4d5e-bd6b-b46cf8b75fdd · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

Representation Forcing for Bottleneck-Free Unified Multimodal Models Emu3: Next-Token Prediction is All You Need

Reference 49

Resolution
unresolved
no resolver link, observed 2026-07-12T15:31:57.426559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T15:31:57.426559Z digest=sha256:94d9d8baf0fef9cc482552334b01fa8e81a40b2cb6358446833def37679ce370

Observation 1f76f380-4e7a-4895-b126-660d9eb3c36a · outbound

This paper cites Cubic discrete diffusion: Discrete visual generation on high-dimensional representation tokens.

Representation Forcing for Bottleneck-Free Unified Multimodal Models Cubic discrete diffusion: Discrete visual generation on high-dimensional representation tokens

Reference 50

Resolution
unresolved
no resolver link, observed 2026-07-12T15:31:57.426559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T15:31:57.426559Z digest=sha256:c5ceab3a5790e4e4dacee4c005aa3c9f59d705159b79c603a9c0f458edc7bacd

Observation 1a48579d-3f14-4131-9451-fec5fa3acfdc · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

Representation Forcing for Bottleneck-Free Unified Multimodal Models Wan: Open and Advanced Large-Scale Video Generative Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-07-12T15:31:57.426559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T15:31:57.426559Z digest=sha256:23d5e9cb4f614ff9b2c417f4593e20c7fd6cef7d308f6c2ee6a12480fca8960a

Observation df4b6c76-3515-4b4c-9fa5-74e03582d2c2 · outbound

This paper cites Qwen-Image Technical Report.

Representation Forcing for Bottleneck-Free Unified Multimodal Models Qwen-Image Technical Report

Reference 52

Resolution
unresolved
no resolver link, observed 2026-07-12T15:31:57.426559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T15:31:57.426559Z digest=sha256:1a775e7791e9a367745c70dcbaca4646c2748f9dd8d9d6828aab0a68e37e55cc

Observation ae0397fa-b5f3-465a-a96c-f989e40ee16d · outbound

This paper cites Janus: Decoupling visual encoding for unified multimodal understanding and generation.

Representation Forcing for Bottleneck-Free Unified Multimodal Models Janus: Decoupling visual encoding for unified multimodal understanding and generation

Reference 53

Resolution
unresolved
no resolver link, observed 2026-07-12T15:31:57.426559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T15:31:57.426559Z digest=sha256:baff90ecd641933f2135ddb0b115df91d642dc6e8ab61d2e74375a97429b562a

Observation bead864d-e8cf-4f43-97e8-96de1d0b857a · outbound

This paper cites OmniGen2: Towards Instruction-Aligned Multimodal Generation.

Representation Forcing for Bottleneck-Free Unified Multimodal Models OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 54

Resolution
unresolved
no resolver link, observed 2026-07-12T15:31:57.426559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T15:31:57.426559Z digest=sha256:9d013e0759eb474b116e7763f2fd5459419dd9739e11893d3688549dbd0e4502

Observation a8d0f5fe-c403-4420-bd67-7b0e4a4102be · outbound

This paper cites RealWorldQA.https://huggingface.co/datasets/xai-org/RealworldQA, 2024.

Representation Forcing for Bottleneck-Free Unified Multimodal Models RealWorldQA.https://huggingface.co/datasets/xai-org/RealworldQA, 2024

Reference 55

Resolution
unresolved
no resolver link, observed 2026-07-12T15:31:57.426559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T15:31:57.426559Z digest=sha256:95a668cf7e66587b4f97d42ff6ec019a6b86e4ffc86db15eaf123a3ef953940c

Observation 1302f11a-3c2e-4776-9e08-48ca82a1c58b · outbound

This paper cites Show-o: One single transformer to unify multimodal understanding and generation.

Representation Forcing for Bottleneck-Free Unified Multimodal Models Show-o: One single transformer to unify multimodal understanding and generation

Reference 56

Resolution
unresolved
no resolver link, observed 2026-07-12T15:31:57.426559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T15:31:57.426559Z digest=sha256:b29f50f0b0f2b9162a647137f42d3521e0a0c818a4ec5988d122e501420e0906

Observation d952adf4-9e74-4cae-865f-abd790f17f38 · outbound

This paper cites Show-o2: Improved native unified multimodal models.

Representation Forcing for Bottleneck-Free Unified Multimodal Models Show-o2: Improved native unified multimodal models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-07-12T15:31:57.426559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T15:31:57.426559Z digest=sha256:420d6e173e846e71af19ca5d71361b25a9eb6289a2cc886a7b3806c664aea9da

Observation 6da44f9a-1181-4959-9684-c3f56eb86c31 · outbound

This paper cites Qwen3 Technical Report.

Representation Forcing for Bottleneck-Free Unified Multimodal Models Qwen3 Technical Report

Reference 58

Resolution
unresolved
no resolver link, observed 2026-07-12T15:31:57.426559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T15:31:57.426559Z digest=sha256:bae2de1ecf1f81fb2a230cad37c6af102735623b177aacef3afe00c291e2bf33

Observation 24966641-45da-4f6d-8c10-5b702c88dc4f · outbound

This paper cites Context Unrolling in Omni Models.

Representation Forcing for Bottleneck-Free Unified Multimodal Models Context Unrolling in Omni Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-07-12T15:31:57.426559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T15:31:57.426559Z digest=sha256:ddd276856d06e2745b4ba84a0478f8f5b242791a9b571eaae09dbe90faa24c2b

Observation e7509791-14da-459b-a708-57f8acfdcbd4 · outbound

This paper cites Representation alignment for generation: Training diffusion transformers is easier than you think.

Representation Forcing for Bottleneck-Free Unified Multimodal Models Representation alignment for generation: Training diffusion transformers is easier than you think

Reference 60

Resolution
unresolved
no resolver link, observed 2026-07-12T15:31:57.426559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T15:31:57.426559Z digest=sha256:9c7e5c1957c8ed17e2856123fd02e6e535502170333148ff87999d2d27e6b11f

Observation 2bdf6708-a946-430a-ab16-e1b28be01d33 · outbound

This paper cites MMMU: A massive multi-discipline multimodal understanding and reasoning benchmark for expert AGI.

Representation Forcing for Bottleneck-Free Unified Multimodal Models MMMU: A massive multi-discipline multimodal understanding and reasoning benchmark for expert AGI

Reference 61

Resolution
unresolved
no resolver link, observed 2026-07-12T15:31:57.426559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T15:31:57.426559Z digest=sha256:641b79d14af174eff1198ff05bcefb88f2cb7dfd73176731029e73a681b76560

Observation 1ea2d78c-ba30-4fbc-9249-11a3c2607b33 · outbound

This paper cites Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer.

Representation Forcing for Bottleneck-Free Unified Multimodal Models Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 62

Resolution
unresolved
no resolver link, observed 2026-07-12T15:31:57.426559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T15:31:57.426559Z digest=sha256:17c43d8ba55bda2c4070f5a4cb38e1e7605949805db6be844924d0bc8ef1fe5a

Observation 46e97e8e-bb93-4246-86e7-a13d4615fb6e · outbound

This paper cites Sigmoid loss for language image pre-training.

Representation Forcing for Bottleneck-Free Unified Multimodal Models Sigmoid loss for language image pre-training

Reference 63

Resolution
unresolved
no resolver link, observed 2026-07-12T15:31:57.426559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T15:31:57.426559Z digest=sha256:0e160226cca7eca189b2c2a07a76209d96e016249a61493e5daa8475283f6214

Observation 0336a190-c260-48d3-ad0b-83799df53e93 · outbound

This paper cites Diffusion Transformers with Representation Autoencoders.

Representation Forcing for Bottleneck-Free Unified Multimodal Models Diffusion Transformers with Representation Autoencoders

Reference 64

Resolution
unresolved
no resolver link, observed 2026-07-12T15:31:57.426559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T15:31:57.426559Z digest=sha256:ee588d235e9f850784deec195a40f9e9b679d1c6e355d1adfe5f738a40fa443d

Observation c6409544-364c-48c5-bb68-87be1b19612b · outbound

This paper cites Transfusion: Predict the next token and diffuse images with one multi-modal model.

Representation Forcing for Bottleneck-Free Unified Multimodal Models Transfusion: Predict the next token and diffuse images with one multi-modal model

Reference 65

Resolution
unresolved
no resolver link, observed 2026-07-12T15:31:57.426559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T15:31:57.426559Z digest=sha256:b6bf74f3bdb508b29faa083c9d5c09de9c4aafc81577548902c6ed6f91409e24

Pith citing papers

Observation b7656482-5415-4733-9a24-0c270ba24f6f · inbound

LLM can Read Spectrogram: Encoder-free Speech-Language Modeling cites this paper.

LLM can Read Spectrogram: Encoder-free Speech-Language Modeling Representation Forcing for Bottleneck-Free Unified Multimodal Models

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-07-03T03:47:35.731856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T14:40:40.673091Z digest=sha256:a44a5d5bde07a390e61af13af2815f6752fc0fc5655097b28d834af25420858e

Observation 921868c2-f87e-478e-92ee-4b377756fa90 · inbound

LLM can Read Spectrogram: Encoder-free Speech-Language Modeling cites this paper.

LLM can Read Spectrogram: Encoder-free Speech-Language Modeling Representation Forcing for Bottleneck-Free Unified Multimodal Models

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-07-03T23:59:06.125682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-03T23:55:11.698728Z digest=sha256:bb6dff5a7342c0312a564e5ef3952d3260a7bb7544da0c74b94b23f25406abed

Observation b199c5fd-609f-490d-93d7-a54484d3769d · inbound

dRAE: Representation Autoencoder with Hyper-Spherical Codes cites this paper.

dRAE: Representation Autoencoder with Hyper-Spherical Codes Representation Forcing for Bottleneck-Free Unified Multimodal Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-01T05:45:51.915796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T05:45:51.915796Z digest=sha256:3b6dc052644dcf3206f5dd9eb8bb01d25ecae97ba152f57384b38e693b0c8f7f