Pith. sign in

Paper Citation Record · LEDGER

UniSpace: Unified Visual Representation and Scalable Multimodal Modeling

As of 19 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 0 inbound Pith citation observations for arXiv:2608.08676.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.08676 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-14T04:35:00.082460Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

29 of 29 outbound references displayed

  • verified exact0
  • verified fuzzy2
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e621f310-3270-469f-89f5-129f7d825fbf · outbound

This paper cites Qwen3-VL Technical Report.

UniSpace: Unified Visual Representation and Scalable Multimodal Modeling Qwen3-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-14T04:34:59.973147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:34:59.973147Z digest=sha256:cefc92d78e730c7735722607d981ec3b61b44a77e69f191d07cbd369d8d3b6f5

Observation d72c16c2-eb37-42a8-87f1-73260244fc5b · outbound

This paper cites BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset.

UniSpace: Unified Visual Representation and Scalable Multimodal Modeling BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-14T04:34:59.985870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:34:59.985870Z digest=sha256:66b7f332ecf6a47e362a1d1ebf1a4d239748de784e3dda5da45a0702c8b0e683

Observation bd069a96-b81a-43a6-8f97-f25542075de6 · outbound

This paper cites Emerging Properties in Unified Multimodal Pretraining.

UniSpace: Unified Visual Representation and Scalable Multimodal Modeling Emerging Properties in Unified Multimodal Pretraining

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-14T04:34:59.994218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:34:59.994218Z digest=sha256:85f0050b0d769336f2d2429740c9ff0e70c21565ae65a3a33ae03b75e4ff423f

Observation 53a6db67-010e-423c-9a81-3eb2e55ddbf6 · outbound

This paper cites FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space.

UniSpace: Unified Visual Representation and Scalable Multimodal Modeling FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-14T04:35:00.002489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:35:00.002489Z digest=sha256:c0cae575638fa77b3e29b39542c139e1b4e5798cac01ec93d8b8e6ecb50542c1

Observation 4b5ffe82-787e-4fe4-a740-f8db8b3152b2 · outbound

This paper cites UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation.

UniSpace: Unified Visual Representation and Scalable Multimodal Modeling UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-14T04:35:00.006708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:35:00.006708Z digest=sha256:ba9bab9b94297408d16fd9b309a7dd8c5b4f4552c72e891152bddce997f2fd47

Observation 8496c9a4-e396-40e2-aeac-5f7f016a4c8c · outbound

This paper cites SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture.

UniSpace: Unified Visual Representation and Scalable Multimodal Modeling SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-14T04:35:00.014679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:35:00.014679Z digest=sha256:7e8645f7aae15932629b77ea0b7302b7152ca2743222472909b080e342f1551a

Observation add6a57d-7e87-4fbf-9cd8-d308a1d22fa1 · outbound

This paper cites Improved Baselines with Representation Autoencoders.

UniSpace: Unified Visual Representation and Scalable Multimodal Modeling Improved Baselines with Representation Autoencoders

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-14T04:35:00.018811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:35:00.018811Z digest=sha256:916b263440c283f92dac56fdf5d11b8bcfa7725b992b276bf50acece1becc8a6

Observation 381f21c1-0d50-4c37-9cca-de69f1eea056 · outbound

This paper cites Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation.

UniSpace: Unified Visual Representation and Scalable Multimodal Modeling Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-14T04:35:00.022696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:35:00.022696Z digest=sha256:702ec41a8eff064852fb29a37553fa52b57e4effd74a444d474a13e963977913

Observation 2ec9e8b9-182e-4218-a2d5-89e091b772c0 · outbound

This paper cites Unilip: Adapting clip for unified multimodal understanding, generation and editing.arXiv preprint arXiv:2507.23278,.

UniSpace: Unified Visual Representation and Scalable Multimodal Modeling Unilip: Adapting clip for unified multimodal understanding, generation and editing.arXiv preprint arXiv:2507.23278,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-14T04:35:00.026740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:35:00.026740Z digest=sha256:56a430f7b0a39c8e6da41a630d289bfb5f97bcb66e082c03121266f9bba66322

Observation aa365a61-c9ca-4868-b640-c32eab5a5bbc · outbound

This paper cites LongCat-Image Technical Report.

UniSpace: Unified Visual Representation and Scalable Multimodal Modeling LongCat-Image Technical Report

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-14T04:35:00.031628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:35:00.031628Z digest=sha256:a75bab6da7ec0447f445777fa1aa327ae97b452e1e931dffa5bf6f05567158a5

Observation 456da09c-2e1b-470d-bc74-8214fa49e118 · outbound

This paper cites Internvl-u: Democratizing unified multimodal models for understanding, reason- ing, generation and editing.arXiv preprint arXiv:2603.09877,.

UniSpace: Unified Visual Representation and Scalable Multimodal Modeling Internvl-u: Democratizing unified multimodal models for understanding, reason- ing, generation and editing.arXiv preprint arXiv:2603.09877,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-14T04:35:00.035768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:35:00.035768Z digest=sha256:e97bd85e8528bb6295861f5cb515dbe0bfd56c7b7ea75ab014731ca1da5e4434

Observation dbbbce84-ff93-4d0f-882d-d088928faa1e · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

UniSpace: Unified Visual Representation and Scalable Multimodal Modeling SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-14T04:35:00.039219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:35:00.039219Z digest=sha256:1970438621e8b96392f6bdf4602cd05d4164dabe72ae1d8c215df5dd7b68da6a

Observation 7a48512a-b97c-49d7-bdb8-14fd04e6e34a · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

UniSpace: Unified Visual Representation and Scalable Multimodal Modeling Wan: Open and Advanced Large-Scale Video Generative Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-14T04:35:00.042931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:35:00.042931Z digest=sha256:97c562f94ca28c074b0e1e4a92c6f6e56078f8dbda927918ff7bfe06bf6acb02

Observation 23ad3e4c-6cc5-435c-ad12-8a2e9c55ece3 · outbound

This paper cites Ovis-U1 Technical Report.

UniSpace: Unified Visual Representation and Scalable Multimodal Modeling Ovis-U1 Technical Report

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-14T04:35:00.046640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:35:00.046640Z digest=sha256:60e5fdc7b19dedb8e980d7d3d36b2356c06af51629e9a8942a9cb6571d5241fd

Observation 3039595b-5f4e-4713-b413-1aba85e15f3c · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

UniSpace: Unified Visual Representation and Scalable Multimodal Modeling Emu3: Next-Token Prediction is All You Need

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-14T04:35:00.050798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:35:00.050798Z digest=sha256:a9f812769d845200c1eb876dfbb3d5934c9a369c61c267abd6878a3fcf219671

Observation 5cae6dc2-f84e-42ad-ac26-b0dcec11157e · outbound

This paper cites Qwen-Image Technical Report.

UniSpace: Unified Visual Representation and Scalable Multimodal Modeling Qwen-Image Technical Report

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-14T04:35:00.054543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:35:00.054543Z digest=sha256:73d5f5fb2d049613d92d0cb4735386eba8c3f1a6a9ee2d17faee0e0f1c54c837

Observation 67e3f80d-8d05-48eb-bc73-2e5ef6b31a02 · outbound

This paper cites Show-o: One single transformer to unify multimodal understanding and generation.

UniSpace: Unified Visual Representation and Scalable Multimodal Modeling Show-o: One single transformer to unify multimodal understanding and generation

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:35:00.580074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-14T04:35:00.058497Z digest=sha256:de4df3ce94ebdc0a81fd7b410671ea697ae2ed21d8e460b8a11ec11a6e2d4f8f

Observation 8be55882-6aff-4a27-94bc-0c52087fc11d · outbound

This paper cites Qwen3 Technical Report.

UniSpace: Unified Visual Representation and Scalable Multimodal Modeling Qwen3 Technical Report

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-14T04:35:00.062047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:35:00.062047Z digest=sha256:6fb046db001f8396d5bc325097e88bd6a44816993cbeb71f1cee14e4d4970122

Observation 0043ccc4-1615-4a77-b51f-456fe860ed6c · outbound

This paper cites Towards scalable pre-training of visual tokenizers for generation.arXiv preprint arXiv:2512.13687, 2025a.

UniSpace: Unified Visual Representation and Scalable Multimodal Modeling Towards scalable pre-training of visual tokenizers for generation.arXiv preprint arXiv:2512.13687, 2025a

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-14T04:35:00.065208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:35:00.065208Z digest=sha256:88f486778205c874d2b60115578b88740ec5ed8b46fdff6f14af409fe57e53ef

Observation d49702a4-d7c7-4551-ae81-78e028b02717 · outbound

This paper cites Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think.

UniSpace: Unified Visual Representation and Scalable Multimodal Modeling Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-14T04:35:00.068270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:35:00.068270Z digest=sha256:8d931dc05a8f9ff9844dc670fe9db498202b2057b22acbcc188845e63efb87b7

Observation ec95d48c-a851-4217-b034-c6b8e5d509e4 · outbound

This paper cites Uniflow: A unified pixel flow tokenizer for visual understanding and generation.

UniSpace: Unified Visual Representation and Scalable Multimodal Modeling Uniflow: A unified pixel flow tokenizer for visual understanding and generation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-14T04:35:00.071818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:35:00.071818Z digest=sha256:ce7912a5805c646f954ab0cda6bdbb8825e82dec62c5d413b6914842cc80fd6c

Observation d10f0e56-a29e-428a-8be7-940bff895fdf · outbound

This paper cites QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation.

UniSpace: Unified Visual Representation and Scalable Multimodal Modeling QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-14T04:35:00.075374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:35:00.075374Z digest=sha256:ae1b2cc0de9fc62623134fcf1547ffb4f8a8807d15672a2635e7ba4d1d3b7039

Observation 5eb9ceb8-263f-4320-8ce5-5b32c9aaacfe · outbound

This paper cites Diffusion Transformers with Representation Autoencoders.

UniSpace: Unified Visual Representation and Scalable Multimodal Modeling Diffusion Transformers with Representation Autoencoders

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-14T04:35:00.079028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:35:00.079028Z digest=sha256:6b2b0b1abbe7bc20488eb17d50f54957d27a85f9111cb371716d38edad451e07

Observation 5f755917-476c-4da8-8b5f-18dd6f89baac · outbound

This paper cites All variants share the same encoder, decoder, training schedule, and sampling setting, isolating the effect of the semantic–reconstruction balance in the flow-matching objective.

UniSpace: Unified Visual Representation and Scalable Multimodal Modeling All variants share the same encoder, decoder, training schedule, and sampling setting, isolating the effect of the semantic–reconstruction balance in the flow-matching objective

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:35:00.568663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-14T04:35:00.082460Z digest=sha256:32f79c74c00b8afc3f8dfd663b84245b4d55d31bf4348a905678691707872a7b

Observation 5cf963ea-b8d0-4ac6-86bd-d6810d7a49e3 · outbound

This paper cites OneIG-Bench: Omni-dimensional Nuanced Evaluation for Image Generation.

UniSpace: Unified Visual Representation and Scalable Multimodal Modeling OneIG-Bench: Omni-dimensional Nuanced Evaluation for Image Generation

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-14T04:34:59.981913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:34:59.981913Z digest=sha256:9910c3203dad2c5706075610346efb33590c3d73c2179b6fa48437b54ee646b0

Observation 487af554-5389-45fc-8fb4-da075f47c1b3 · outbound

This paper cites ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment.

UniSpace: Unified Visual Representation and Scalable Multimodal Modeling ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-14T04:34:59.998626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:34:59.998626Z digest=sha256:002863180e8238470aabcb4d1e4760d578b0f979bc7c62e2fa0ca7adb11801cd

Observation d8a3a56c-eb6e-4729-a48a-1510864321de · outbound

This paper cites Emu3.5: Native Multimodal Models are World Learners.

UniSpace: Unified Visual Representation and Scalable Multimodal Modeling Emu3.5: Native Multimodal Models are World Learners

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-14T04:34:59.990040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:34:59.990040Z digest=sha256:cb5d72cea149bf2454ef2327c0c31b5b4b5d2a9b72f089807c8bf6331648c682

Observation efdd9383-919a-4126-aadf-0fd5219db74e · outbound

This paper cites Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer.

UniSpace: Unified Visual Representation and Scalable Multimodal Modeling Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-14T04:34:59.977929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:34:59.977929Z digest=sha256:9558fba270f0fc348be231583c3c10876760ed630d8e47000dfed4630d0a40d9

Observation a0e391c7-a656-417e-a3c7-b90697520443 · outbound

This paper cites Open-MAGVIT2: An Open-Source Project Toward Democratizing Auto-regressive Visual Generation.

UniSpace: Unified Visual Representation and Scalable Multimodal Modeling Open-MAGVIT2: An Open-Source Project Toward Democratizing Auto-regressive Visual Generation

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-14T04:35:00.010792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:35:00.010792Z digest=sha256:54dc7d0b97ec37d7a745a016fcdac123ea84d558a52a2bf36fadbe66347592b8

Pith citing papers

No inbound Pith citation observations are available.